Image processing method and device, electronic equipment, computer readable storage medium and computer program product
By decoding the control signals and viewing angle vectors of the target object model, dynamic neural texture maps and viewing angle neural texture maps are generated and fused with diffuse texture maps, solving the problem of insufficient expressive power of head texture maps and achieving more detailed texture rendering effects.
Patent Information
- Application Number
- CN202410517868.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-04-26
- Publication Date
- 2025-10-28
AI Technical Summary
The expressive power of rendering head texture maps in related technologies is limited, and they cannot fully express the detailed textures of different areas of the head, resulting in poor rendering results.
By acquiring the control signals and viewing angle vectors of the target object model, dynamic neural texture maps and viewing angle neural texture maps are decoded respectively, and then fused with pre-trained diffuse neural texture maps to generate more detailed texture maps.
It improves the detail representation of texture maps, depicting the texture of target objects more precisely from both appearance and perspective, thus enhancing the rendering effect.
Smart Images

Figure CN120852609A_ABST
Abstract
Description
Technical Field
[0001] This application relates to artificial intelligence technology, and more particularly to an image processing method, apparatus, electronic device, computer-readable storage medium, and computer program product. Background Technology
[0002] In image rendering technology, the expressive power of texture maps has a significant impact on the image rendering effect. Therefore, choosing an appropriate method to obtain texture maps with good expressive power is an important research direction in image rendering technology.
[0003] Taking the head as an example, the facial texture map rendered by related technologies has limited expressive power and cannot fully express the detailed textures of different areas of the head, such as extremely fine hair strands, which leads to poor rendering results of the details of the target object. Summary of the Invention
[0004] This application provides an image processing method, apparatus, electronic device, computer-readable storage medium, and computer program product that can improve the ability to express the detail of textures in rendered images.
[0005] The technical solution of this application embodiment is implemented as follows:
[0006] This application provides an image processing method, the method comprising:
[0007] Obtain control signals and viewing angle vectors for the target object model;
[0008] Based on the control signal, a dynamic neural texture map of the first region in the appearance dimension of the target object model is obtained by decoding.
[0009] Decoding is performed based on the observation view vector to obtain the view neural texture map of the first region in the view dimension;
[0010] The dynamic neural texture map, the visual neural texture map, and the pre-trained diffuse neural texture map are fused to obtain the texture map of the first region.
[0011] This application provides an image processing apparatus, the apparatus comprising:
[0012] The acquisition module is used to acquire control signals and viewing angle vectors for the target object model.
[0013] The decoding module is used to decode based on the control signal to obtain a dynamic neural texture map of the first region in the appearance dimension of the target object model; and to decode based on the viewing angle vector to obtain a viewing angle neural texture map of the first region in the viewing angle dimension.
[0014] The fusion module is used to fuse the dynamic neural texture map, the visual neural texture map, and the pre-trained diffuse neural texture map to obtain the texture map of the first region.
[0015] This application provides an electronic device, the electronic device comprising:
[0016] Memory is used to store executable instructions for a computer;
[0017] The processor, when executing computer-executable instructions stored in the memory, implements the image processing method provided in the embodiments of this application.
[0018] This application provides a computer-readable storage medium storing a computer program or computer-executable instructions for implementing the image processing method provided in this application when executed by a processor.
[0019] This application provides a computer program product, including a computer program or computer executable instructions, which, when executed by a processor, implement the image processing method provided in this application.
[0020] The embodiments of this application have the following beneficial effects:
[0021] By decoding the control signals and viewing angle vectors of the target object model, the dynamic neural texture map of the first region in the target object model in the appearance dimension and the viewing angle neural texture map in the viewing angle dimension are obtained respectively. The dynamic neural texture map, the viewing angle neural texture map and the pre-trained diffuse neural texture map are then fused to obtain the texture map of the first region. Compared with the related technologies that directly obtain the overall texture map of the target object, this method can depict the texture of the target object in more detail from both the appearance and viewing angle dimensions, thereby improving the detail expression capability of the texture map used for image rendering of the target object. Attached Figure Description
[0022] Figure 1 This is a schematic diagram of the architecture of the image processing system 100 provided in an embodiment of this application;
[0023] Figure 2 This is a schematic diagram of the structure of the electronic device 500 provided in the embodiments of this application;
[0024] Figure 3A This is a schematic flowchart of the image processing method provided in the embodiments of this application;
[0025] Figure 3B This is a schematic diagram of the process for obtaining dynamic neural texture maps provided in an embodiment of this application;
[0026] Figure 3CThis is a schematic diagram of the process for obtaining a viewpoint neural texture map provided in an embodiment of this application;
[0027] Figure 3D This is a schematic diagram of the process for obtaining a diffuse neural texture map provided in an embodiment of this application;
[0028] Figure 3E This is a schematic diagram of the process for obtaining the first offset information provided in an embodiment of this application;
[0029] Figure 3F This is a schematic diagram of the process for generating the first image to be rendered provided in an embodiment of this application;
[0030] Figure 3G This is a schematic diagram of the pixel-by-pixel decoding process provided in an embodiment of this application;
[0031] Figure 3H This is a schematic diagram of the first process for generating a target rendered image provided in an embodiment of this application;
[0032] Figure 3I This is a schematic diagram of the process for obtaining the second offset information provided in an embodiment of this application;
[0033] Figure 3J This is a schematic diagram of the second process for generating a target rendered image provided in an embodiment of this application;
[0034] Figure 4A This is a schematic diagram of the structure of the deep learning model provided in the embodiments of this application;
[0035] Figure 4B This is a decoding diagram of the deep learning model provided in the embodiments of this application;
[0036] Figure 5A This is a schematic diagram of the first region offset provided in the embodiments of this application;
[0037] Figure 5B This is a schematic diagram of the three-dimensional Gaussian representation provided in the embodiments of this application;
[0038] Figure 6 This is a schematic diagram of the facial representation principle provided in the embodiments of this application;
[0039] Figure 7 This is a schematic diagram of the pixel-by-pixel decoder provided in an embodiment of this application;
[0040] Figure 8 This is a schematic diagram of the hair representation principle provided in an embodiment of this application;
[0041] Figure 9 This is a schematic diagram illustrating the principle of the image processing method provided in the embodiments of this application;
[0042] Figure 10This application provides known occlusion blending strategies in its embodiments.
[0043] Figure 11 This is a schematic diagram of the rendering effect provided in the embodiment of this application.
[0044] It should be noted that the terms "first" and "second" mentioned above are only used to distinguish between different options and do not represent the degree of superiority or inferiority of the options or their priority in the implementation process. Detailed Implementation
[0045] To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limitations on this application. All other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0046] In the following description, references are made to “some embodiments,” which describe a subset of all possible embodiments. However, it is understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.
[0047] In the following description, the terms "first / second / third" are used merely to distinguish similar objects and do not represent a specific ordering of objects. It is understood that "first / second / third" may be interchanged in a specific order or sequence where permitted, so that the embodiments of this application described herein can be implemented in an order other than that illustrated or described herein.
[0048] In this application embodiment, the terms "module" or "unit" refer to a computer program or part of a computer program that has a predetermined function and works with other related parts to achieve a predetermined goal, and can be implemented wholly or partially using software, hardware (such as processing circuitry or memory), or a combination thereof. Similarly, a processor (or multiple processors or memory) can be used to implement one or more modules or units. Furthermore, each module or unit can be part of an overall module or unit that includes the functionality of that module or unit.
[0049] Unless otherwise specified, "at least one" as used below refers to one or more cases, and "multiple" can refer to two or more cases.
[0050] Unless otherwise defined, all technical and scientific terms used in the embodiments of this application have the same meaning as commonly understood by one of ordinary skill in the art. The terminology used in the embodiments of this application is for the purpose of describing the embodiments of this application only and is not intended to limit this application.
[0051] In the implementation of this application, the collection and processing of relevant data should strictly comply with the requirements of relevant national laws and regulations, obtain the informed consent or separate consent of the personal information subject, and carry out subsequent data use and processing within the scope of laws and regulations and the authorization of the personal information subject.
[0052] Before providing a further detailed description of the embodiments of this application, the nouns and terms involved in the embodiments of this application will be explained, and the nouns and terms involved in the embodiments of this application shall be interpreted as follows.
[0053] 1) Object Model: A model built for the object to be rendered, used to model different regions of the object separately in order to render a rendered image of the object. For example, a head model includes the side head region and the top hair region. Another example is a flower pot model, which includes the side flower pot body region and the top plant region.
[0054] 2) First region: The surface-like region in the object model, such as the head region (including the face region and scalp region) in the head model, or the flower pot body region in the flower pot model.
[0055] 3) Second region: The volumetric region in the object model, such as the hair region in head modeling, or the flower and stem region in flowerpot modeling.
[0056] 4) Mesh: Also known as a triangular mesh, it is a polygonal mesh composed of triangles, used to simulate the surface of complex objects, such as the first region of an object model. It can be a parametric linear model, such as the Flexible and Lightweight Analysis-by-synthesis Head Model (FLAME) mesh, also known as the FLAME triangular mesh. Taking the head region of a head model as an example, it can be constructed with three vector parameters—shape, expression, and pose—to obtain the same topological triangular mesh of the target expression for the corresponding head region.
[0057] 5) 3D Gaussian Representation: A multidimensional Gaussian distribution model for the second region of an object model, including the mean and covariance matrix. The mean represents the center point of the multidimensional Gaussian distribution model, and the covariance matrix describes the shape and orientation of the multidimensional Gaussian distribution model. For example, when modeling the hair region of a human head, it can be represented by 3D Gaussian Splatting (3D GS).
[0058] 6) Control Signals: Input signals used to control and influence the mesh shape changes of the object model, which may include appearance parameters, shape parameters, and pose parameters. Appearance parameters describe and control the appearance features of the object model, such as color, material, and texture; shape parameters describe and control the shape features of the object model, such as spheres and cylinders; pose parameters describe the object model's position (translation), orientation (rotation), and scale (scaling) in 3D space.
[0059] For example, for a head model, the shape of the face, facial expression, and head pose can be combined with control signals to achieve dynamic changes in the head representation. For a flowerpot model, the shape of the flowerpot, the pattern of the flowerpot, and the pose of the flower can be combined with control signals to achieve dynamic changes in the flowerpot representation.
[0060] 7) The Iterative Closest Point (ICP) algorithm is a classic data registration algorithm that iteratively optimizes the matching error between the sampling points in the static space and the corresponding matching points in the dynamic space of the overlapping area of the first and second regions until the error is minimized.
[0061] 8) Multilayer Perceptron (MLP): A basic feedforward neural network structure used to solve machine learning tasks such as classification and regression.
[0062] 9) Dynamic texture maps, in the fields of computer graphics and computer vision, refer to natural textures generated through image synthesis or enhancement that reflect the dynamic changes of an object. These textures change over time, capturing motion, deformation, or environmental changes on the object's surface. Dynamic texture maps are often combined with deep learning techniques, especially neural network architectures such as Generative Adversarial Networks (GANs) or Recurrent Neural Networks (RNNs), to generate high-quality dynamic texture effects. The applications of dynamic texture maps are very broad, including but not limited to film special effects, video games, virtual reality (VR), and augmented reality (AR). They can be used to add realistic dynamic effects to objects in virtual environments or to create complex visual effects in video editing.
[0063] 10) View Texture: Also known as view-dependent texture mapping, it is a technique used for texturing in 3D computer graphics. Typically, in a 3D model, a texture might be applied to its surface to make it look more realistic and detailed. The texture map contains view-related information and therefore varies depending on the observer's viewpoint.
[0064] 11) Diffuse Texture: Also known as a diffuse texture map, it's a texture map used in computer graphics to simulate diffuse lighting effects. It contains information about the texture, color, and reflection properties of a target object's surface. When rendering 3D images, diffuse texture maps are used to simulate the lighting and color changes on an object's surface to make it look more realistic and vivid. Diffuse refers to the uniform reflection of light in all directions after it passes through an object's surface. Diffuse texture maps simulate this diffusion, helping to enhance the lighting and shadow effects of objects and providing a more realistic visual experience.
[0065] Taking head modeling as an example, the expressive power of explicit UV textures in related technologies is limited. High-definition rendering results are usually required with very accurate facial geometry. However, accurate facial geometry is difficult to obtain in practice, so explicit UV textures usually produce blurry rendering results. Neural UV textures usually cannot express facial detail textures. Volume rendering is also difficult to express facial detail textures.
[0066] Based on the above analysis, the applicant found that the texture maps obtained by the relevant technologies have limited expressive power and cannot fully express the detailed textures of different areas of the target object. In response to the above technical problems, this application provides an image processing method that can improve the ability to express the details of textures used for rendering images.
[0067] This application provides an image processing method, apparatus, electronic device, computer-readable storage medium, and computer program product, which can improve the detail expression capability of textures used for rendering images. The following describes exemplary applications of the electronic device provided in this application. The electronic device provided in this application can be implemented as various types of terminals such as laptops, tablets, desktop computers, set-top boxes, smartphones, smart speakers, smartwatches, smart TVs, and in-vehicle terminals, or it can be implemented as a server. The following will describe exemplary applications when the electronic device is implemented as a terminal or server.
[0068] See Figure 1 , Figure 1This is a schematic diagram of the architecture of the image processing system 100 provided in the embodiments of this application. In order to support an image processing application, the terminal 400 (exemplarily showing a graphical interface 410) connects to the server 200 through the network 300. The network 300 can be a wide area network or a local area network, or a combination of both.
[0069] The image processing method provided in this application can be applied to any scenario that requires image rendering based on texture maps. For example, it can be used to render virtual characters in a game scene based on texture maps, or applied to instant messaging programs (APPs) or operating systems to render the image of a smart assistant based on texture maps.
[0070] Terminal 400 is used to send the control signals and viewing angle vectors for the target object model submitted by the client to server 200 via network 300. Server 200 decodes the received control signals and viewing angle vectors respectively to obtain the neural texture maps of the target object model in the first region in appearance dimension and viewing angle dimension. Then, it fuses the neural texture maps in appearance dimension and viewing angle dimension with the pre-trained diffuse neural texture map to obtain the texture map of the first region. The fused texture map is then sent to terminal 400 via network 300 for display on graphical interface 410.
[0071] In some embodiments, the terminal 400 is used to obtain the control signal and viewing angle vector of the target object model submitted by the client, decode them based on the control signal and viewing angle vector respectively, obtain the neural texture map of the target object model in the first region in the appearance dimension and viewing angle dimension, fuse the neural texture map in the appearance dimension and viewing angle dimension with the pre-trained diffuse neural texture map to obtain the texture map of the first region, and display the fused texture map on the graphical interface 410.
[0072] In some embodiments, server 200 may be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms. Terminals and servers can be connected directly or indirectly via wired or wireless communication, which is not limited in this embodiment.
[0073] The embodiments of this application can be implemented using artificial intelligence (AI) technology. AI is the theory, methods, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results. In other words, AI is a comprehensive technology within computer science that attempts to understand the essence of intelligence and produce a new type of intelligent machine that can react in a way similar to human intelligence. AI also studies the design principles and implementation methods of various intelligent machines, enabling them to possess perception, reasoning, and decision-making capabilities.
[0074] Artificial intelligence (AI) is a comprehensive discipline encompassing a wide range of fields, including both hardware and software technologies. Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, pre-trained model technology, operating / interactive systems, and mechatronics. Among these, pre-trained models, also known as large-scale models or foundational models, can be widely applied to downstream tasks across various AI fields after fine-tuning. AI software technologies primarily include computer vision, speech processing, natural language processing, and machine learning / deep learning.
[0075] The electronic device that implements the image processing method provided in the embodiments of this application may be Figure 1 Terminal 400 or server 200. See also Figure 2 , Figure 2 This is a schematic diagram of the structure of the electronic device 500 provided in the embodiments of this application. Figure 2 The illustrated electronic device 500 includes at least one processor 510, a memory 550, at least one network interface 520, and a user interface 530. The various components in the electronic device 500 are coupled together via a bus system 540. It is understood that the bus system 540 is used to implement communication between these components. In addition to a data bus, the bus system 540 also includes a power bus, a control bus, and a status signal bus. However, for clarity, ... Figure 2 The general labeled all buses as Bus System 540.
[0076] The processor 510 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor, etc.
[0077] User interface 530 includes one or more output devices 531 that enable the presentation of media content, including one or more speakers and / or one or more visual displays. User interface 530 also includes one or more input devices 532, including user interface components that facilitate user input, such as a keyboard, mouse, microphone, touch screen display, camera, other input buttons and controls.
[0078] The memory 550 may be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state storage, hard disk drives, optical disk drives, etc. The memory 550 may optionally include one or more storage devices physically located away from the processor 510.
[0079] The memory 550 may include volatile memory or non-volatile memory, or both. The non-volatile memory may be read-only memory (ROM), and the volatile memory may be random access memory (RAM). The memory 550 described in this application embodiment is intended to include any suitable type of memory.
[0080] In some embodiments, memory 550 is capable of storing data to support various operations, examples of which include programs, modules, and data structures or subsets or supersets thereof, as illustrated below.
[0081] Operating system 551 includes system programs for handling various basic system services and performing hardware-related tasks, such as the framework layer, core library layer, driver layer, etc., for implementing various basic business functions and handling hardware-based tasks;
[0082] The network communication module 552 is used to reach other electronic devices via one or more (wired or wireless) network interfaces 520, exemplary network interfaces 520 including: Bluetooth, WiFi, and Universal Serial Bus (USB), etc.
[0083] Presentation module 553 is configured to enable the presentation of information (e.g., a user interface for operating peripheral devices and displaying content and information) via one or more output devices 531 (e.g., a display screen, a speaker, etc.) associated with user interface 530;
[0084] The input processing module 554 is used to detect and translate one or more user inputs or interactions from one or more input devices 532.
[0085] In some embodiments, the apparatus provided in this application can be implemented in software. Figure 2 An image processing apparatus 555 stored in memory 550 is shown. This apparatus can be software in the form of programs and plug-ins, and includes the following software modules: an acquisition module 5551, a decoding module 5552, and a fusion module. These modules are logically linked and can therefore be arbitrarily combined or further separated according to their implemented functions. The functions of each module will be described below.
[0086] In some embodiments, the terminal or server can implement the image processing method provided in the embodiments of this application by running various computer-executable instructions or computer programs. For example, computer-executable instructions may be microprogram-level commands, machine instructions, or software instructions. Computer programs may be native programs or software modules in an operating system; in short, the aforementioned computer-executable instructions may be any form of instruction, and the aforementioned computer programs may be any form of application program, module, or plug-in.
[0087] The image processing method provided in this application will be described in conjunction with exemplary applications and implementations of the electronic devices provided in the embodiments of this application.
[0088] The image processing method provided in the embodiments of this application will be described below. As mentioned above, the electronic device implementing the image processing method of the embodiments of this application can be a terminal or a server, or a combination of both. Therefore, the executing entity of each step will not be described again below.
[0089] See Figure 3A , Figure 3A This is a flowchart illustrating the image processing method provided in the embodiments of this application, with an electronic device as the execution subject, combining... Figure 3A The steps shown are explained.
[0090] In step 101, control signals and viewing angle vectors for the target object model are obtained.
[0091] In some embodiments, the control signals for the target object model can be manually set to control the shape, pose, and appearance of the target object model, and the viewing angle vector represents the vector of the viewing angle for the target object model.
[0092] As an example, for a head model, the shape control signal could be a head scaling signal to control the head size, the pose control signal could be a head turning signal to control the head's turning movement, and the appearance control signal could be an eye closing signal to control the blinking expression. For a flowerpot model, the shape control signal could be a flowerpot height adjustment signal, such as stretching or compressing the flowerpot's height, the pose control signal could be a signal to move the flowerpot left, right, or forward and backward, and the appearance control signal could be a signal to modify the pattern on the outside of the flowerpot.
[0093] In step 102, the control signal is decoded to obtain the dynamic neural texture map of the first region in the appearance dimension of the target object model.
[0094] In some embodiments, the first region in the target object model refers to the side region of the target object model. For example, the first region of a head model is the side head region; the first region of a flowerpot model is the side flowerpot body region. The dynamic neural texture map of the first region in the appearance dimension refers to the dynamic neural texture map of the first region changing synchronously with the appearance of the first region. For example, in the case of a head model, the dynamic neural texture map of the first region in the appearance dimension means that the dynamic neural texture map of the head region changes synchronously with the expression (appearance) of the head region.
[0095] In some embodiments, dynamic neural texture maps can be generated as follows: First, a neural network is trained to understand and capture the spatial structure of an image, including visual elements such as color, shape, and texture. Then, a model is trained to predict changes in the image under different random motion conditions, such as translation, rotation, and scaling, and the effects of these motions are simulated to enhance the image's dynamism. Finally, an infinitely looping video is generated by combining user input or specific scene requirements. For example, users can influence and generate dynamic effects through simple drag-and-drop operations.
[0096] In some embodiments, see Figure 3B , Figure 3B This is a schematic diagram of the process for obtaining dynamic neural texture maps provided in the embodiments of this application. Figure 3A Step 102 can be achieved through Figure 3B Steps 1021 to 1022 are implemented, and the details are explained below.
[0097] In step 1021, feature extraction is performed on the appearance parameters and pose parameters in the control signal to obtain feature vectors for the appearance parameters and pose parameters, respectively.
[0098] In some embodiments, feature extraction of appearance parameters and pose parameters can be performed using a pre-trained deep learning model to obtain feature vectors for appearance parameters and pose parameters, respectively. See also Figure 4A , Figure 4A This is a schematic diagram of the structure of the deep learning model provided in an embodiment of this application. Figure 4A The convolutional layer of the deep learning model with the structure shown extracts features from the appearance and pose parameters of the input control signal to obtain feature maps corresponding to the appearance and pose parameters. Then, the feature maps are pooled by the pooling layer to obtain feature vectors corresponding to the appearance and pose parameters of the control signal.
[0099] As an example, the pre-trained deep learning model can be a Convolutional Neural Network (CNN) or an autoencoder. Pooling layers are used to process the feature map; this can be max pooling, which preserves the main features in the feature map, or average pooling, which reduces the dimensionality of the features in the feature map.
[0100] In step 1022, the feature vectors of the appearance parameters and the feature vectors of the pose parameters are decoded to obtain a dynamic neural texture map that changes synchronously with the changes in the appearance dimension of the first region in the target object model.
[0101] In some embodiments, a pre-trained deep learning model can be used to decode the feature vectors of appearance parameters and pose parameters to obtain a dynamic neural texture map that changes synchronously with the changes in appearance dimension of the first region in the target object model.
[0102] See Figure 4B , Figure 4B This is a decoding diagram of the deep learning model provided in an embodiment of this application. Figure 4B The deep learning model shown performs deconvolution processing on the feature vectors corresponding to the appearance parameters and pose parameters of the input control signal to obtain the mapping features of the first region relative to the grid; the mapping features are upsampled to obtain a dynamic neural texture map that changes synchronously with the changes in the appearance dimension of the first region in the target object model.
[0103] Deconvolution is the process of mapping feature vectors extracted through operations such as convolution and pooling back to the original image space. By increasing the size and dimension of the feature map, the details and structure of the image are restored in reverse. For example, after feature extraction from the input image of a head model, feature vectors representing head information are obtained. By gradually increasing the size and depth of the head feature map through deconvolution operations, the feature vectors are mapped back to the image space, which can restore the original appearance and features of the head model, such as the head contour and facial features.
[0104] Upsampling is the process of increasing the size and dimensions of an image or feature map through interpolation and padding. By increasing the number of pixels and the level of detail, it improves the image's detail and resolution, thereby enhancing visual effects and visualization quality. For example, upsampling the mapped features of a head model can increase the detail and resolution of the generated image, making the head model more realistic and lifelike.
[0105] This application embodiment decodes the parameters of the control signal by calling a pre-trained machine learning model, and obtains a dynamic neural texture map that changes synchronously with the changes in the appearance dimension of the first region in the target object model. It can fully explore the correlation between the parameters in the control signal and the dynamic neural texture map of the first region, thereby achieving effective decoding of the control signal. It has advantages such as high-precision decoding, improved efficiency, and strong technical reliability, and has broad application prospects and practical value.
[0106] See also Figure 3A The following will be an explanation following step 102 above.
[0107] In step 103, decoding is performed based on the observation view vector to obtain the view neural texture map of the first region in the view dimension.
[0108] In some embodiments, the viewpoint neural texture map of the first region in the target object model in the viewpoint dimension refers to the viewpoint neural texture map of the first region changing synchronously according to the observation viewpoint of the first region.
[0109] In some embodiments, a viewpoint neural texture map can be generated in the following manner: First, the viewpoint is calculated in a first grid, and a corresponding texture map is selected according to the viewpoint; then, the pixel value of the corresponding texture is calculated based on the calculated viewpoint texture coordinates, and a viewpoint neural texture map is generated based on the pixel value of the texture.
[0110] In some embodiments, a pre-trained deep learning model can be used to decode the viewing angle vector to obtain the viewpoint neural texture map of the first region in the viewpoint dimension. The training process of the deep learning model can be implemented as follows: First, a training dataset is obtained, including the viewing angle vector and the corresponding reference viewpoint neural texture map in the viewpoint dimension; second, the data in the training dataset is preprocessed, including normalization, noise reduction, and other operations; then, the parameters of the deep learning model are initialized, taking the viewing angle vector in the training dataset as input, and the initialized deep learning model outputs the predicted viewpoint neural texture map; finally, based on the difference between the predicted viewpoint neural texture map and the reference viewpoint neural texture map, the parameters of the initialized deep learning model are updated using the backpropagation algorithm to obtain the pre-trained deep learning model.
[0111] As an example, a pre-trained deep learning model can be a convolutional neural network or an autoencoder.
[0112] In some embodiments, see Figure 3C , Figure 3C This is a schematic diagram of the process for obtaining a viewpoint neural texture map provided in an embodiment of this application. Figure 3A Step 103 can be achieved through Figure 3C Steps 1031 to 1032 are implemented, and the details are explained below.
[0113] In step 1031, feature extraction is performed on the observation view vector to obtain the feature vector of the observation view vector.
[0114] In some embodiments, features can be extracted from the viewpoint vector using the convolutional layer of a pre-trained deep learning model to obtain a feature map corresponding to the viewpoint vector; then, the feature map can be pooled using a pooling layer to obtain a feature vector of the viewpoint vector.
[0115] In step 1032, the feature vector of the observation view vector is decoded to obtain a view neural texture map that changes synchronously with the changes in the view dimension of the first region.
[0116] In some embodiments, a pre-trained deep learning model can be used to deconvolve the feature vector of the input viewpoint vector to obtain the mapping features of the first region relative to the grid; the mapping features are then upsampled to obtain a viewpoint neural texture map that changes synchronously with the changes in the viewpoint dimension of the first region.
[0117] This application embodiment can learn the mapping relationship between the observation view vector and the view dimension neural texture map by calling a pre-trained deep learning model, thereby decoding the observation view and generating the view dimension neural texture map, thus improving the accuracy and effect of image generation.
[0118] See also Figure 3A The following will be an explanation following step 103 above.
[0119] In step 104, the dynamic neural texture map, the visual neural texture map, and the pre-trained diffuse neural texture map are fused to obtain the texture map of the first region.
[0120] In some embodiments, weights can be set for the dynamic neural texture map, the visual neural texture map, and the pre-trained diffuse neural texture map, and the dynamic neural texture map, the visual neural texture map, and the pre-trained diffuse neural texture map can be weighted and fused to obtain the texture map of the first region.
[0121] As an example, if the pixel value of the dynamic neural texture map is 0.8 and the weight is 0.4, the pixel value of the view neural texture map is 0.6 and the weight is 0.3, and the pixel value of the pre-trained diffuse neural texture map is 0.5 and the weight is 0.3, then the three texture maps are weighted and fused to obtain the weighted fused pixel value as the sum of the product of the pixel value and the weight of each texture map, which is 0.65.
[0122] In some embodiments, the weights of the dynamic neural texture map, the viewpoint neural texture map, and the pre-trained diffuse neural texture map can be obtained through an attention encoding model. First, pre-configured weight parameters are set for the dynamic neural texture map, the viewpoint neural texture map, and the pre-trained diffuse neural texture map respectively. Then, for each texture map, weighting is performed according to the pre-configured weight parameters to obtain weighted texture data. Finally, the weighted texture data is input into the attention encoding model to learn the feature information of the input texture map and output the attention weight of each texture map.
[0123] As examples, attention encoding models can be attention mechanism neural networks, recurrent neural networks (RNNs), deep attention neural networks, etc.
[0124] This embodiment decodes the control signal and the viewing angle vector to obtain the dynamic neural texture map of the first region in the target object model in the appearance dimension and the viewing angle neural texture map in the viewing angle dimension. The dynamic neural texture map, the viewing angle neural texture map and the pre-trained diffuse neural texture map are then fused to obtain the texture map of the first region. Compared with the related technology of directly obtaining the overall texture map of the target object, this embodiment depicts the texture of the target object in more detail from the two dimensions of appearance and viewing angle, thus improving the expressive power of the texture map.
[0125] In some embodiments, see Figure 3D , Figure 3D This is a schematic diagram of the process for obtaining a diffuse neural texture map provided in an embodiment of this application. During execution... Figure 3A Before step 104, it can be done through Figure 3D Steps 201 to 208 are used to obtain diffuse neural texture maps, which are explained in detail below.
[0126] In step 201, the initialized diffuse neural texture map is obtained.
[0127] In some embodiments, obtaining an initialized diffuse neural texture map can be achieved as follows: First, a noise matrix is randomly generated using a uniform or normal distribution random number generator, where each element takes a value between -1 and 1; second, the generated random noise map is smoothed (e.g., Gaussian blur) or features (e.g., texture blocks) are added; then, the processed noise map is applied to the texture coordinates of the first region to obtain the initialized diffuse neural texture map.
[0128] In step 202, the dynamic neural texture map, the visual neural texture map, and the initialized diffuse neural texture map are fused to obtain the predicted texture map of the first region.
[0129] In some embodiments, weights are first set for the dynamic neural texture map, the visual neural texture map, and the initialized diffuse neural texture map, and then the dynamic neural texture map, the visual neural texture map, and the initialized diffuse neural texture map are fused by weighted fusion to obtain the predicted texture map of the first region.
[0130] As an example, if the pixel value of the dynamic neural texture map is 0.7 and the weight is 0.5, the pixel value of the visual neural texture map is 0.6 and the weight is 0.4, and the pixel value of the initialized diffuse neural texture map is 0.4 and the weight is 0.5, then the three texture maps are weighted and fused to obtain the weighted fused pixel value as the sum of the product of the pixel value and the weight of each texture map, which is 0.79.
[0131] In some embodiments, the weights of the dynamic neural texture map, the viewpoint neural texture map, and the initialized diffuse neural texture map can be obtained through an attention encoding model. First, pre-configured weight parameters are set for the dynamic neural texture map, the viewpoint neural texture map, and the initialized diffuse neural texture map respectively. Then, for each texture map, weighting is performed according to the pre-configured weight parameters to obtain weighted texture data. Finally, the weighted texture data is input into the attention encoding model to learn the feature information of the input texture map and output the attention weight of each texture map.
[0132] As an example, attention encoding models can be attention mechanism neural networks, recurrent neural networks, attention deep neural networks, etc.
[0133] In step 203, a first mesh of the first region is constructed based on the shape parameters, appearance parameters and pose parameters in the control signal, and the first offset information of the first region relative to the first mesh is obtained.
[0134] In some embodiments, the first region of the target object can be constructed into a representation of a first mesh based on the shape parameters, appearance parameters, and pose parameters in the control signal. For the vertices in the first mesh of the first region of the target object model, a positional offset will occur under the drive of the parameters included in the control signal. Based on the parameters of the control signal, the first offset information of the first region relative to the first mesh can be obtained.
[0135] As an example, the first mesh constructed based on the shape parameters, appearance parameters, and pose parameters in the control signal can be a FLAME triangular mesh.
[0136] In some embodiments, see Figure 3E , Figure 3E This is a schematic diagram of the process for obtaining the first offset information provided in an embodiment of this application. Figure 3D Step 203, "obtaining the first offset information of the first region relative to the first grid," can be achieved through... Figure 3E Steps 2031 to 2032 are implemented, and the details are explained below.
[0137] In step 2031, feature extraction is performed on the appearance parameters and pose parameters to obtain feature vectors for the appearance parameters and pose parameters.
[0138] In some embodiments, feature vectors are obtained by extracting features from the appearance and pose parameters of the control signal by invoking a pre-trained deep learning model. See also... Figure 4A It can be done Figure 4A The convolutional layer of the deep learning model with the structure shown extracts features from the appearance and pose parameters of the input control signal to obtain feature maps corresponding to the appearance and pose parameters. Then, the feature maps are pooled by the pooling layer to obtain feature vectors corresponding to the appearance and pose parameters of the control signal.
[0139] As an example, the pre-trained deep learning model can be a convolutional neural network or an autoencoder. Pooling layers are used to process the feature map; this can be max pooling, which preserves the main features in the feature map, or average pooling, which reduces the dimensionality of the features in the feature map.
[0140] Max pooling of feature maps can reduce their spatial size, the number of parameters, and computational complexity, while extracting important features, which helps improve the accuracy of modeling. Average pooling can reduce the size of feature maps, which helps reduce the spatial dimension of feature maps and reduce the risk of overfitting.
[0141] In step 2032, the feature vectors of the appearance parameters and the feature vectors of the pose parameters are decoded to obtain the first offset information of the first region relative to the first grid.
[0142] In some embodiments, the feature vectors of appearance parameters and pose parameters can be decoded using a pre-trained deep learning model to obtain the first offset information of the first region relative to the first grid.
[0143] As an example, when the first region is the head region of the head model, the first offset information could be to move each vertex in the triangular mesh of the head region 3 centimeters to the right. The movement distance of each vertex in the triangular mesh can also be calculated at the pixel level, such as moving a vertex in the triangular mesh upwards by 3 pixels.
[0144] See also Figure 4B ,pass Figure 4B The deep learning model with the structure shown performs deconvolution processing on the feature vectors corresponding to the appearance parameters and pose parameters of the input control signal to obtain the mapping features of the first region relative to the grid; the mapping features are then upsampled to obtain the first offset information of the first region relative to the first grid.
[0145] This application embodiment obtains the first offset information of the first region by calling a pre-trained deep learning model to decode the appearance parameters and pose parameters of the control signal. This enables the deep learning model to combine the appearance parameters and pose parameters in different control signals to achieve fine adjustment of the mesh modeling of the first region, thereby making the mesh modeling adaptable to different scene requirements.
[0146] See also Figure 3D The following will be an explanation following step 203 above.
[0147] In step 204, each vertex in the first grid is offset based on the first offset information to obtain the second grid of the first region.
[0148] In some embodiments, firstly, the first grid data of the first region of the target object model is loaded, and the coordinates of each vertex in the first grid are obtained. Then, for each vertex in the first grid, an offset operation is performed on the vertex according to the corresponding offset value in the first offset information. The offset operation can be implemented in a way such as along the normal vector direction, at a fixed distance, or proportionally. Finally, the vertices that have undergone the offset operation are updated to obtain the second grid of the first region.
[0149] As an example, see Figure 5A , Figure 5A This is a schematic diagram of the first region offset provided in an embodiment of this application. For example... Figure 5AAs shown, each vertex in the first grid 501A is offset 3 centimeters to the right according to the first offset information in the example above, to obtain the second grid 502A.
[0150] In step 205, according to a preset mapping relationship, the texture at each position in the predicted texture map is mapped to the corresponding vertex in the second grid to obtain the third grid. The mapping relationship is used to characterize the correspondence between the position of the vertex in the second grid and the position of the texture in the predicted texture map.
[0151] In some embodiments, firstly, for each vertex in the second grid, according to the correspondence between the position of the vertex in the second grid and the texture coordinates of the predicted texture map in the mapping relationship, the corresponding texture coordinates in the predicted texture map are obtained; then, based on the texture coordinates, sampling is performed in the predicted texture map to obtain the texture information (such as texture, color, etc.) corresponding to each vertex; finally, the obtained texture information is mapped onto the second grid so that each vertex contains the texture information of the corresponding position, thus obtaining the texture information of the third grid.
[0152] In step 206, the third grid is rasterized to obtain the coordinates of each pixel in the first region and the corresponding predicted latent vector.
[0153] In some embodiments, taking a triangular mesh to be rendered as an example, the triangular mesh includes multiple triangular faces. First, a projection transformation is performed on each vertex in the face, mapping the vertices in three-dimensional space (i.e., the vertices of the triangular face) to a two-dimensional screen coordinate system, obtaining the position of each face vertex in the two-dimensional screen coordinate system. Then, based on the position and color value of each face vertex in the two-dimensional screen coordinate system, color value interpolation is performed on the positions between vertices. For example, the color value of the pixels between vertices can be obtained by linear interpolation based on the centroid coordinates of the pixels within the face. Finally, the latent vector of each pixel is calculated based on the color value of the pixels (including the pixels between vertices).
[0154] In step 207, pixel-level decoding is performed based on the coordinates of each pixel in the first region and the corresponding predicted latent vector to obtain the predicted rendered image of the first region.
[0155] In some embodiments, a pre-trained deep learning model is used to extract features from the coordinates of each pixel in the first region and the corresponding predicted latent vector to obtain the feature vector and corresponding position coordinates of each pixel in the first region. Then, the feature vector of each pixel is linearly transformed to obtain a linear feature vector. Finally, the linear feature vector is non-linearly mapped to obtain the color value of the corresponding position of each pixel. The predicted rendering image of the first region is generated based on the color value of each pixel.
[0156] In step 208, based on the difference between the predicted rendered image and the preset reference image, the initialized diffuse neural texture map is updated to obtain the pre-trained diffuse neural texture map of the first region.
[0157] In some embodiments, the difference between the predicted rendered image and a preset reference image can be calculated using the mean squared error (MSE). The smaller the MSE, the more similar the predicted rendered image is to the preset reference image. Based on the pixel-level difference between the predicted rendered image and the real image, the gradient descent algorithm can be used to update the initialized diffuse neural texture map to obtain a pre-trained diffuse neural texture map for the first region.
[0158] This application embodiment uses an initialized diffuse neural texture map, which is fused with the dynamic neural texture map and the viewpoint neural texture map of the first region of the target object to generate a predicted texture map of the first region. The predicted texture map is then mapped to a second grid after an offset operation based on a preset mapping relationship to obtain a third grid. The third grid is then rasterized and decoded pixel by pixel to obtain a predicted rendered image of the first region. Finally, the loss between the predicted rendered image and a preset reference image is calculated, the gradient is calculated using the backpropagation algorithm, and the initialized diffuse neural texture map is updated to obtain a diffuse neural texture map of the first region. This allows the mapping relationship between the predicted texture map and the second grid during the forward propagation process to be learned, thereby optimizing the diffuse neural texture map and learning the texture map generation process. This enables fine control over the texture changes in the diffuse neural texture map and improves the expressive power of the texture map.
[0159] In some embodiments, during execution Figure 3A After step 104, the first image to be rendered in the first region can be generated in the first mesh of the first region based on the first offset information and the texture map of the first region. See also Figure 3F , Figure 3F This is a schematic diagram of the process for generating a first image to be rendered, provided in an embodiment of this application. Figure 3F Steps 301 to 303 are used to obtain and generate the first image to be rendered, which are explained in detail below.
[0160] In step 301, according to the mapping relationship, the texture at each position in the texture map is mapped to the vertex corresponding to the position in the second mesh to obtain the mesh to be rendered.
[0161] As an example, the preset mapping relationship is as follows: texture coordinate 1 in the texture map corresponds to vertex coordinate 1 in the second grid, texture coordinate 2 in the texture map corresponds to vertex coordinate 2 in the second grid, and texture sampling is performed based on the texture coordinates to obtain texture A corresponding to texture coordinate 1 and texture B corresponding to texture coordinate 2. Then, A is mapped to the position of vertex coordinate 1 in the second grid, and B is mapped to the position of vertex coordinate 2 in the second grid to obtain the grid to be rendered.
[0162] In step 302, the mesh to be rendered is rasterized to obtain the coordinates of each pixel in the first region and the corresponding latent vector.
[0163] In some embodiments, taking a triangular network as an example, the mesh to be rendered comprises multiple triangular facets. During rasterization of the mesh, firstly, a projection transformation is performed on each vertex of the mesh to map the vertices in three-dimensional space (i.e., the vertices of the triangular facets) to a two-dimensional screen coordinate system, determining the positions of the multiple triangular facets in the two-dimensional screen coordinate system; then, for each pixel in each triangle, color value interpolation is performed based on its position within the triangle. The pixel's color value can be determined through linear interpolation based on the pixel's centroid coordinates within the triangle; finally, the latent vector of each pixel is calculated based on its color value.
[0164] In step 303, pixel-level decoding is performed based on the coordinates of each pixel in the first region and the corresponding latent vector to obtain the first image to be rendered in the first region.
[0165] In some embodiments, a pre-trained deep learning model is used to extract features from the coordinates and corresponding latent vectors of each pixel in the first region to obtain the feature vector and corresponding position coordinates of each pixel in the first region. Then, the feature vector of each pixel is linearly transformed to obtain a linear feature vector. Finally, the linear feature vector is non-linearly mapped to obtain the color value of the corresponding position of each pixel. The first image to be rendered in the first region is generated based on the color value of each pixel.
[0166] In some embodiments, see Figure 3G , Figure 3G This is a schematic diagram of the pixel-by-pixel decoding process provided in an embodiment of this application. Figure 3F Step 303 can be achieved through Figure 3G Steps 3031 to 3034 are implemented, and the details are explained below.
[0167] In step 3031, feature extraction is performed on the coordinates of each pixel and its corresponding latent vector to obtain the feature vector of each pixel.
[0168] In some embodiments, a pre-trained deep learning model is used to extract features from the pixel coordinates and corresponding latent vectors to obtain the pixel's feature vector and corresponding position coordinates. The pre-trained deep learning model can be obtained as follows: First, a training dataset is acquired, in which each training sample includes the pixel coordinate position information, latent vector representation, and corresponding reference feature vector; then, the training samples are input into an initialized deep learning model, which outputs the pixel's predicted feature vector; finally, based on the difference between the predicted feature vector and the reference feature vector, the parameters of the initialized deep learning model are updated using the backpropagation algorithm to obtain the pre-trained deep learning model.
[0169] In step 3032, a linear transformation is performed on the feature vector of each pixel to obtain a linear feature vector.
[0170] In some embodiments, a linear feature vector is obtained by performing a linear transformation on the feature vector of a pixel through a fully connected layer of a multilayer perceptron.
[0171] In step 3033, a non-linear mapping is performed on the linear feature vector to obtain the color value of each pixel.
[0172] In some embodiments, the color value of a pixel is obtained by performing a non-linear mapping on a linear feature vector using an activation function.
[0173] As an example, the activation function can be the Sigmoid function, the Softmax function, or the Rectified Linear Unit (ReLU) function.
[0174] In step 3034, a predicted rendering image of the first region is generated based on the color value of each pixel.
[0175] This embodiment of the application maps the texture map of the first region to a second mesh after offset operation to obtain a mesh to be rendered. Then, the mesh to be rendered is rasterized and decoded pixel by pixel to obtain the first image to be rendered of the first region. This realizes the generation of the first image to be rendered of the first region through texture mapping and pixel-by-pixel decoding operations. It can learn the image texture details and structure of the first region, making the rendered image clearer and more accurate.
[0176] In some embodiments, after generating a first image to be rendered in a first grid of the first region based on the first offset information and the texture map of the first region, see [link to previous section]. Figure 3H , Figure 3H This is a schematic diagram of the first process for generating a target rendered image provided in an embodiment of this application, which can be achieved through... Figure 3H Steps 401 to 404 generate the target rendering image, which are explained in detail below.
[0177] In step 401, the static three-dimensional Gaussian representation of the second region of the target object model is obtained, and the second offset information of the static three-dimensional Gaussian representation is obtained.
[0178] In some embodiments, a second region of the target object model can be modeled using a three-dimensional Gaussian representation. The static three-dimensional Gaussian representation of the second region is the three-dimensional Gaussian representation of the second region in static space. The second offset information of the static three-dimensional Gaussian representation refers to information about transformation operations such as rigid transformation or non-rigid offset performed on the static three-dimensional Gaussian representation of the second region.
[0179] In some embodiments, see Figure 3I , Figure 3I This is a schematic diagram of the process for obtaining the second offset information provided in an embodiment of this application. Figure 3H Step 401, "obtaining the second offset information of the static three-dimensional Gaussian representation," can be achieved through... Figure 3I Steps 4011 to 4015 are implemented, and the details are explained below.
[0180] In step 4011, the overlapping area between the first region and the second region is determined.
[0181] As an example, the first region of the head model is the head region (including the face region and the scalp region), and the second region is the hair region. Since the hair region is attached to the scalp region, the overlapping area of the head region and the hair region is the scalp region.
[0182] In step 4012, the rigid transformation information between the sampling points of the overlapping region in the static space and the matching points corresponding to the sampling points in the dynamic space is determined.
[0183] In some embodiments, the matching point in the dynamic space corresponding to the sampling point is the point closest to the sampling point. The corresponding sampling point and matching point in the static space and the dynamic space can be found by using a feature point matching algorithm.
[0184] As an example, feature point matching algorithms can be Scale-Invariant Feature Transform (SIFT) algorithms or Speeded-Up Robust Features (SURF) algorithms. These algorithms can determine the rigid transformation information between the sampling points of the overlapping region in static space and the matching points corresponding to the sampling points in dynamic space through iterative nearest-point algorithms.
[0185] A rigid transformation is a geometric transformation that preserves distance and angle between two points in a plane or space. In two-dimensional space, rigid transformations include three basic transformations: translation, rotation, and mirroring, which maintain the distance and angular relationship between two points. See also: Figure 5B , Figure 5B This is a schematic diagram of the three-dimensional Gaussian representation provided in the embodiments of this application, as shown below. Figure 5B As shown, the static three-dimensional Gaussian representation 501B is rotated 90 degrees counterclockwise to obtain the initialized dynamic three-dimensional Gaussian representation 502B. This transformation is called a rigid transformation.
[0186] In step 4013, feature extraction is performed on the appearance parameters in the control signal to obtain the feature vector of the appearance parameters.
[0187] In some embodiments, the appearance parameters in the control signal are convolved using a pre-trained deep learning model to obtain a feature map corresponding to the appearance parameters in the control signal; the feature map is then pooled to obtain a feature vector of the appearance parameters in the control signal.
[0188] In step 4014, the feature vectors of the appearance parameters are nonlinearly mapped to obtain non-rigid offset information in static three-dimensional Gaussian representation.
[0189] In some embodiments, the feature vectors of the appearance can be nonlinearly mapped using the activation function of a multilayer perceptron to obtain non-rigid offset information in a static three-dimensional Gaussian representation.
[0190] Non-rigid offsets refer to geometric transformations that preserve the overall shape of an object but allow for localized deformation. Unlike rigid transformations, non-rigid offsets allow deformation of localized areas rather than maintaining the size or shape of the entire object. See also: Figure 5B The width of the initialized dynamic three-dimensional Gaussian representation 502B is reduced to three-quarters of its original width, resulting in the dynamic three-dimensional Gaussian representation 503B. This transformation is called non-rigid offset.
[0191] In step 4015, the rigid transformation information and the non-rigid offset information are added together to obtain the second offset information in static three-dimensional Gaussian representation.
[0192] In some embodiments, the rigid transformation information calculated based on the sampling points and matching degree of the second region is summed with the non-rigid offset information represented by the static three-dimensional Gaussian representation, and this sum is used as the second offset information of the static three-dimensional Gaussian representation of the second region.
[0193] As an example, if the rigid transformation information is a 90-degree counterclockwise rotation and the non-rigid offset information is to reduce the width to three-quarters of its original value, then the second offset information in the static 3D Gaussian representation is a 90-degree counterclockwise rotation followed by a reduction in width to three-quarters of its original value.
[0194] See also Figure 3H The following will be an explanation following step 401 above.
[0195] In step 402, based on the second offset information, the static three-dimensional Gaussian representation is transformed from the static space to the dynamic space to obtain the dynamic three-dimensional Gaussian representation.
[0196] In some embodiments, a rigid transformation is performed on the static three-dimensional Gaussian representation in the static space based on the rigid transformation information in the second offset information to obtain an initialized dynamic three-dimensional Gaussian representation in the dynamic space. An offset operation is then performed on the initialized dynamic three-dimensional Gaussian representation based on the non-rigid offset information in the second offset information to obtain the dynamic three-dimensional Gaussian representation.
[0197] As an example, see Figure 5B ,like Figure 5B As shown, the static three-dimensional Gaussian representation 501B undergoes a rigid transformation to obtain the initialized dynamic three-dimensional Gaussian representation 502B. Based on the non-rigid offset information, the initialized dynamic three-dimensional Gaussian representation 502B is offset to obtain the dynamic three-dimensional Gaussian representation 503B.
[0198] In step 403, a second image to be rendered is generated based on a dynamic three-dimensional Gaussian representation.
[0199] In some embodiments, a rendering engine can be used to project a dynamic 3D Gaussian representation onto a 2D image plane, convert the information of the dynamic 3D Gaussian representation into pixel values or texture information of the image, and generate a second image to be rendered.
[0200] In step 404, the first image to be rendered and the second image to be rendered are blended to obtain the target rendered image of the target object model.
[0201] In some embodiments, a first image to be rendered corresponding to a first region of the target object model and a second image to be rendered corresponding to a second region can be blended to obtain a target rendered image of the target object model. The blending can employ different blending modes, such as overlay, transparency, additive blending, etc., and the blending ratio of the first and second images to be rendered can be adjusted as needed to achieve different rendering effects.
[0202] In some embodiments, see Figure 3J , Figure 3J This is a schematic diagram of the second process for generating a target rendered image provided in an embodiment of this application. Figure 3H Step 404 can be achieved through Figure 3I Steps 4041 to 4044 are implemented, and the details are explained below.
[0203] In step 4041, a first rendering depth corresponding to each pixel in the first image to be rendered and a second rendering depth corresponding to each pixel in the second image to be rendered are determined.
[0204] In some embodiments, the first rendering depth corresponding to each pixel in the first image to be rendered is the depth obtained after rasterization processing, and the second rendering depth corresponding to each pixel in the second image to be rendered is the rendering depth value of the Gaussian hair among all three-dimensional Gaussian distributions covering the pixel that is closest to the pixel.
[0205] In step 4042, in response to the second rendering depth of a pixel at the same position being greater than or equal to the first rendering depth, the color value at the position corresponding to the first rendering depth is used as the color value of the pixel at the corresponding position in the target rendering image of the target object model.
[0206] As an example, when creating a head model, if the hair rendering depth of a pixel at the same location is greater than or equal to the face rendering depth, the color of the location corresponding to the face rendering depth is used as the color of the corresponding pixel in the target rendering image of the target object model.
[0207] In step 4043, in response to the fact that the second rendering depth of a pixel at the same location is less than the first rendering depth, the color value of the location corresponding to the second rendering depth is used as the color value of the pixel at the corresponding location in the target rendering image of the target object model.
[0208] As an example, when creating a head model, if the hair rendering depth of a pixel at the same location is less than the face rendering depth, the color of the location corresponding to the hair rendering depth is used as the color of the corresponding pixel in the target rendering image of the target object model.
[0209] In step 4044, for pixels at different locations, the color of the corresponding location is used as the color of the corresponding pixel in the target rendering image of the target object model.
[0210] As an example, for pixels with different positions in the first and second regions, the colors of the first and second regions at the corresponding positions are used as the colors of the pixels at the corresponding positions in the target rendered image of the target object model.
[0211] In step 4045, a target rendering image of the target object model is generated based on the color value of the pixel at each location of the target rendering image.
[0212] As an example, when creating a head model, the pixels at each location in the head region and hair region are combined to obtain the target rendered image of the target object.
[0213] In this embodiment, different methods are used to model the first and second regions of the target object model. A first image to be rendered is generated based on the first offset information and the texture map of the first region. Based on the second offset information, the three-dimensional Gaussian representation of the second region is transformed from static space to dynamic space, thereby rendering the second image to be rendered. Finally, the first image to be rendered and the second image to be rendered are mixed to obtain the target rendered image, which can improve the accuracy of image rendering.
[0214] This embodiment decodes the control signals and viewing angle vectors of the target object model to obtain dynamic neural texture maps of the first region in the target object model in both appearance and viewing angle dimensions. The dynamic neural texture map, the viewing angle neural texture map, and a pre-trained diffuse neural texture map are then fused to obtain the texture map of the first region. Compared to directly obtaining the overall texture map of the target object in related technologies, this method can more meticulously depict the texture of the target object from both appearance and viewing angle dimensions, improving the expressive power of the texture map used for image rendering of the target object. A first image to be rendered is generated based on the first offset information and the texture map of the first region. Based on the second offset information, the three-dimensional Gaussian representation of the second region is transformed from static space to dynamic space, thereby rendering the second image to be rendered. Finally, the first image to be rendered is blended with the first image to be rendered to obtain the target rendered image, which improves the accuracy of image rendering.
[0215] The image processing method provided in this application can be applied to different scenarios that require the creation of high-definition virtual human avatars, such as games, augmented reality (AR) / virtual reality (VR) remote social networking, remote collaboration, etc.
[0216] The following will describe an exemplary application of the embodiments of this application in the scenario of rendering a smart assistant image in a virtual scene.
[0217] In the scenario of rendering images of intelligent assistants, the expressive power of texture maps has a significant impact on the rendering effect. Therefore, choosing an appropriate method to obtain texture maps with good expressive power is an important research direction in image rendering technology. The expressive power of facial texture maps obtained by related technologies is limited, and they cannot express the detailed textures of the face, such as facial wrinkles and extremely fine hair strands. Using facial texture maps obtained by related technologies for head modeling will result in poor rendering results of the facial details of the target object.
[0218] The image processing method provided in this application decodes the control signals of the human head (i.e., the shape, expression, and pose parameters of the FLAME) and the viewing angle vector respectively to obtain texture maps of different facial effects in different dimensions (expression dimension and viewing angle dimension) of the head region (i.e., viewing angle-dependent view texture map and expression-dependent dynamic texture map). Combined with the basic neural texture map (diffuse neural texture map), a facial neural texture map is obtained, which supports more detailed skin texture rendering. Combined with 3D Gaussian modeled hair, high-quality rendering of any expression on the whole head can be achieved.
[0219] See Figure 6 , Figure 6 This is a schematic diagram of the facial representation principle provided in the embodiments of this application; as follows: Figure 6 As shown, an enhanced FLAME triangular mesh is used as the facial geometry representation of the human head to ensure accurate expression driving, while relying on predicted neural texture maps and pixel-by-pixel decoders to render detailed facial textures.
[0220] First, the control signals of the human head (i.e., the shape β, expression ψ, and pose φ of the FLAME) are obtained. Based on the shape β, expression ψ, and pose φ parameters in the control signals, an enhanced FLAME triangular mesh is generated, as shown in formula (1), where, This represents the generated enhanced FLAME triangular mesh. This indicates that the triangular mesh has 16428 vertices, and each vertex contains three positional information, namely the coordinates of the x-axis, y-axis, and z-axis; Represents the triangular facets of the FLAME triangular mesh. This indicates that there are 40212 triangular faces in the triangular mesh, and each triangular face has 3 vertices.
[0221]
[0222] Secondly, the UV offset map is obtained by decoding based on the facial expression ψ and pose φ. This is used to apply a per-vertex offset to the FLAME triangular mesh, making the mesh more closely resemble the real facial geometry. As shown in Equation (2), the offset operation is based on the shape β, expression ψ, and pose φ of the FLAME. To indicate, triangular mesh The refined triangular mesh is obtained by applying per-vertex offset.
[0223]
[0224] The UV offset map is obtained by decoding the facial expression ψ and pose φ based on the control signal. Simultaneously, the decoding yielded a dynamic neural texture map of the head region in the facial expression dimension. Furthermore, the view neural texture map of the head region in the view dimension was obtained by decoding from the view vector. Dynamic Neural Texture Map and visual neural texture map These are used to render facial textures in the expression dimension and the view dimension, respectively. Finally, the dynamic neural texture map is... and visual neural texture map With diffuse neural texture map The summation yields a facial neural texture map. As shown in formula (3).
[0225]
[0226] Facial nerve texture map The texture in the image is mapped to the corresponding position of the triangular mesh according to the preset UV mapping relationship to obtain the mesh to be rendered. After the mesh to be rendered is rasterized, the rendering image of the head area is obtained by using a pixel-by-pixel decoder.
[0227] See Figure 7 , Figure 7 This is a schematic diagram of the pixel-by-pixel decoder provided in the embodiments of this application; as shown Figure 7 As shown, the uv coordinates and latent vector z obtained by rasterization are used as inputs to the pixel-by-pixel decoder. After passing through the fully connected layer of MLPs, a linear transformation is performed to obtain a linear feature vector. Then, a linear transformation is performed through a linear layer, and finally, a non-linear mapping is performed through an activation function to obtain the color value at the corresponding position of the pixel. The color value at the corresponding pixel position is then output to obtain the face rendering image.
[0228] The following section continues by explaining how the hair region of a maneuverable human head is represented. See also... Figure 8 , Figure 8 This is a schematic diagram of the hair representation principle provided in the embodiments of this application; as follows: Figure 8 As shown, this embodiment of the application uses 3D Gaussian as the representation of hair.
[0229] First, to obtain a high-fidelity dynamic hair rendering result, a 3D Gaussian model of hair was trained using a multi-view image of a frame from the training data as the standard space representation of hair, denoted as... Where i represents the i-th Gaussian distribution among N Gaussian distributions, This indicates the center position of the i-th Gaussian distribution. Indicates the direction of the i-th Gaussian distribution. This represents the magnitude of the i-th Gaussian distribution. Let represent the opacity of the i-th Gaussian distribution. This represents the color of the i-th Gaussian distribution.
[0230] Then, in order to obtain the dynamic hair in different frames, a rigid transformation from standard space scalp to dynamic space scalp was first estimated using ICP. After the rigid transformation, the 3D Gaussian representation of the hair was transformed into... The calculation method is shown in formula (4).
[0231]
[0232] Among them, (R) i ,t i (Standard space scalp) To the dynamic space scalp Rigid transformation.
[0233] Then, using the facial expression parameters in the control signal as input, the non-rigid offset of each 3D Gaussian is predicted using MLP, as shown in Equation (5).
[0234]
[0235] in, ψ represents the calculation of the non-rigid offset of each 3D Gaussian using the expression parameter ψ as input, resulting in the non-rigid offset (δx,δr,δs,δo,δc).
[0236] Finally, based on rigid transformation and non-rigid offset, the 3D Gaussian representation of the hair in dynamic space is obtained as shown in Equation (6), and the hair is rendered using Gaussian splashing to obtain the rendered image of the hair.
[0237]
[0238] in, This represents the 3D Gaussian representation of hair in dynamic space.
[0239] After obtaining the rendered images of the head region and the hair region using the method described above, see [link to documentation]. Figure 9 , Figure 9This is a schematic diagram illustrating the principle of the image processing method provided in this application embodiment. It blends the rendered images of the head region and the hair region to obtain the final rendered image. The color of each pixel in the final rendered image is determined by comparing the rendering depth of the 3D Gaussian hair with the rendering depth of the facial triangular mesh. If the rendering depth of the 3D Gaussian hair is less than the rendering depth of the facial triangular mesh, the pixel location in the final image uses the rendering color of the 3D Gaussian hair; otherwise, the rendering color of the facial mesh is used. It should be noted that the rendering depth of the 3D Gaussian hair is directly taken as the closest depth among all 3D Gaussian meshes covering the pixel to ensure stability during optimization. The rendering depth of the facial mesh is directly obtained using the depth obtained through rasterization.
[0240] See Figure 10 , Figure 10 This application provides known occlusion blending strategies in its embodiments; such as... Figure 10 As shown, by comparing the "near z" depth map D of the hair region nz The rendering depth at the same location as the head depth map Dh is used to determine the color value corresponding to the rendering depth to be used in the final rendered image. Figure 10 The color value of the rendering location corresponding to the white area in the binary mask M0 is shown, and then combined with the mesh occlusion-aware opacity map A. g Get the hair blending texture This results in the head blending texture as follows: Finally, the final rendered image is obtained using formula (7).
[0241]
[0242] in, The image to be rendered is the hair area. The image to be rendered is the head region. Render the image for the final target.
[0243] In this application embodiment, the peak signal-to-noise ratio (PSNR), structural similarity (SSIM), and perceptual similarity (LPIPS) of the effective pixel region after rasterization were used as evaluation metrics to evaluate the rendering results.
[0244] The mean variance of the effective pixel region after rasterization is calculated as shown in formula (8), and the peak signal-to-noise ratio is calculated as shown in formula (9). M and N are the width and height of the image, respectively, MAX is the maximum pixel value of the image (usually 1.0 or 255), and Q is the mask that marks the effective pixel region (effective pixels are 1, and invalid pixels are 0).
[0245]
[0246]
[0247] The calculation of the structural similarity of the effective pixel region after rasterization is based on the sliding window. That is, each calculation takes an N×N window from the image and calculates the SSIM index based on the window. After traversing the entire image, the average value of all windows is taken as the SSIM index of the entire image. For each window, the structural similarity calculation method is shown in formula (10).
[0248]
[0249] Where x and y represent the windows corresponding to the two images, respectively, and μ x ,μ y These are the mean pixel values for the two windows, σ and σ', respectively. x ,σ y σ represents the variance of the pixel values corresponding to the two windows, respectively. xy c1 and c2 are the covariance of the pixel values of the two windows, and are two constants, which are (0.01*255) by default. 2 and (0.03*255) 2 .
[0250] Perceptual similarity is calculated for two image patches based on the features of a deep network. The smaller the result, the more similar the two image patches are.
[0251] This application embodiment compares the image processing method provided in this embodiment with two other methods in related technologies for three metrics: peak signal-to-noise ratio (PSNR), structural similarity, and perceptual similarity, for nine target objects in a dataset, and calculates the average metric. Table 1 shows the identification codes of the nine target objects and the metric values of the three methods provided in this application embodiment (MeGA), PointAvatar, and GussianAvatars, respectively, for the three metrics. According to the average metric, the image processing method provided in this application embodiment outperforms the other two methods in all comparisons.
[0252]
[0253]
[0254] Table 1
[0255] See Figure 11 , Figure 11 This is a schematic diagram of the rendering effect provided in the embodiment of this application. Figure 11The rendering effects and real images of the image processing method (MeGA), PointAvatar, and GussianAvatars provided in the embodiments of this application are shown respectively. Figure 11 As shown, the rendering effect of the image processing method provided in this application embodiment is superior to the rendering effect of related technologies.
[0256] This application embodiment decodes the control signals and viewing angle vectors of the human head to obtain dynamic neural texture maps of the head region in the appearance dimension and viewing angle neural texture maps in the viewing angle dimension. The dynamic neural texture map, the viewing angle neural texture map, and the pre-trained diffuse neural texture map are then fused to obtain a neural texture map of the head region. Compared with the related technologies that directly obtain the overall texture map of the target object, this method can depict the texture of the target object in more detail from both the appearance and viewing angle dimensions, thereby improving the expressive power of the texture map used for image rendering of the target object.
[0257] The following description continues to illustrate the exemplary structure of the image processing apparatus 555 provided in the embodiments of this application as a software module. In some embodiments, such as... Figure 2 As shown, the software modules stored in the image processing device 555 in the memory 550 may include:
[0258] The acquisition module 5551 is used to acquire control signals and viewing angle vectors for the target object model.
[0259] The decoding module 5552 is used to decode based on control signals to obtain the dynamic neural texture map of the first region in the appearance dimension of the target object model. Decoding is also performed based on the viewing angle vector to obtain the viewpoint neural texture map of the first region in the viewpoint dimension.
[0260] The fusion module 5553 is used to fuse the dynamic neural texture map, the visual neural texture map, and the pre-trained diffuse neural texture map to obtain the texture map of the first region.
[0261] In some embodiments, the decoding module 5552 is further configured to extract features from the appearance parameters and pose parameters in the control signal to obtain feature vectors of the appearance parameters and pose parameters, respectively; and to decode the feature vectors of the appearance parameters and pose parameters to obtain a dynamic neural texture map that changes synchronously with the changes in the appearance dimension of the first region in the target object model.
[0262] In some embodiments, the decoding module 5552 is further configured to extract features from the observation view vector to obtain a feature vector of the observation view vector; and to decode the feature vector of the observation view vector to obtain a view neural texture map that changes synchronously with the change in the view dimension of the first region.
[0263] In some embodiments, the acquisition module 5551 is further configured to acquire an initialized diffuse neural texture map; fuse the dynamic neural texture map, the viewpoint neural texture map, and the initialized diffuse neural texture map to obtain a predicted texture map of a first region; construct a first mesh of the first region based on the shape parameters, appearance parameters, and pose parameters in the control signal, and acquire first offset information of the first region relative to the first mesh; perform an offset operation on each vertex in the first mesh based on the first offset information to obtain a second mesh of the first region; map the texture at each position in the predicted texture map to the corresponding vertex in the second mesh according to a preset mapping relationship to obtain a third mesh, wherein the mapping relationship is used to characterize the correspondence between the position of the vertex in the second mesh and the position of the texture in the predicted texture map; perform rasterization processing on the third mesh to obtain the coordinates of each pixel in the first region and the corresponding predicted latent vector; perform pixel-level decoding based on the coordinates of each pixel in the first region and the corresponding predicted latent vector to obtain a predicted rendered image of the first region; and update the initialized diffuse neural texture map based on the difference between the predicted rendered image and the preset reference image to obtain a pre-trained diffuse neural texture map of the first region.
[0264] In some embodiments, the acquisition module 5551 is further configured to perform feature extraction on the appearance parameters and pose parameters to obtain feature vectors of the appearance parameters and pose parameters; and decode the feature vectors of the appearance parameters and pose parameters to obtain first offset information of the first region relative to the first grid.
[0265] In some embodiments, the fusion module 5553 is further configured to generate a first image to be rendered in a first grid of a first region based on the first offset information and the texture map of the first region.
[0266] In some embodiments, the fusion module 5553 is further configured to map the texture of each position in the texture map to the vertex corresponding to the position in the second grid according to the mapping relationship, to obtain the grid to be rendered; to perform rasterization processing on the grid to be rendered to obtain the coordinates of each pixel in the first region and the corresponding latent vector; and to perform pixel-level decoding based on the coordinates of each pixel in the first region and the corresponding latent vector to obtain the first image to be rendered in the first region.
[0267] In some embodiments, the fusion module 5553 is further configured to extract features from the coordinates and corresponding latent vectors of each pixel to obtain a feature vector for each pixel; perform a linear transformation on the feature vector of each pixel to obtain a linear feature vector; perform a nonlinear mapping on the linear feature vector to obtain a color value for each pixel; and generate a predicted rendering image of the first region based on the color value of each pixel.
[0268] In some embodiments, the fusion module 5553 is further configured to obtain a static three-dimensional Gaussian representation of a second region of the target object model and obtain second offset information of the static three-dimensional Gaussian representation; based on the second offset information, transform the static three-dimensional Gaussian representation from a static space to a dynamic space to obtain a dynamic three-dimensional Gaussian representation; generate a second image to be rendered based on the dynamic three-dimensional Gaussian representation; and blend the first image to be rendered and the second image to be rendered to obtain a target rendering image of the target object model.
[0269] In some embodiments, the acquisition module 5551 is further configured to: determine the overlapping region between the first region and the second region; determine the rigid transformation information between the sampling point in the static space and the matching point corresponding to the sampling point in the dynamic space of the overlapping region; extract features from the appearance parameters in the control signal to obtain the feature vector of the appearance parameters; perform nonlinear mapping on the feature vector of the appearance parameters to obtain nonrigid offset information in static three-dimensional Gaussian representation; and add the rigid transformation information and the nonrigid offset information to obtain the second offset information in static three-dimensional Gaussian representation.
[0270] In some embodiments, the fusion module 5553 is further configured to determine a first rendering depth corresponding to each pixel in the first image to be rendered and a second rendering depth corresponding to each pixel in the second image to be rendered; in response to the second rendering depth of pixels at different positions being greater than or equal to the first rendering depth, the color value of the position corresponding to the first rendering depth is used as the color value of the pixel at the corresponding position in the target rendering image of the target object model; in response to the second rendering depth of pixels at the same position being less than the first rendering depth, the color value of the position corresponding to the second rendering depth is used as the color value of the pixel at the corresponding position in the target rendering image of the target object model; and generate the target rendering image of the target object model based on the color value of the pixel at each position in the target rendering image.
[0271] This application provides a computer program product comprising a computer program or computer-executable instructions stored in a computer-readable storage medium. The processor of an electronic device reads the computer-executable instructions from the computer-readable storage medium and executes the computer-executable instructions, causing the electronic device to perform the image processing method described in this application.
[0272] This application provides a computer-readable storage medium storing computer-executable instructions or a computer program. When the computer-executable instructions or the computer program are executed by a processor, the processor will execute the image processing method provided in this application. For example, ... Figure 3A The image processing method shown.
[0273] In some embodiments, the computer-readable storage medium may be a memory such as RAM, ROM, flash memory, magnetic surface memory, optical disk, or CD-ROM; or it may be a variety of devices including one or any combination of the above-mentioned memories.
[0274] In some embodiments, computer-executable instructions may take the form of programs, software, software modules, scripts, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as stand-alone programs or as modules, components, subroutines, or other units suitable for use in a computing environment.
[0275] As an example, computer-executable instructions may, but do not necessarily, correspond to files in a file system. They may be stored as part of a file that holds other programs or data, for example, in one or more scripts in a Hyper Text Markup Language (HTML) document, in a single file dedicated to the program in question, or in multiple co-located files (e.g., files that store one or more modules, subroutines, or code sections).
[0276] As an example, computer-executable instructions can be deployed to execute on a single electronic device, or on multiple electronic devices located at one location, or on multiple electronic devices distributed across multiple locations and interconnected via a communication network.
[0277] In summary, this embodiment decodes the control signals and viewing angle vectors of the target object model to obtain dynamic neural texture maps of the first region in the target object model in both appearance and viewing angle dimensions. The dynamic neural texture map, the viewing angle neural texture map, and a pre-trained diffuse neural texture map are then fused to obtain the texture map of the first region. Compared to directly obtaining the overall texture map of the target object in related technologies, this approach can more meticulously depict the texture of the target object from both appearance and viewing angle dimensions, improving the expressive power of the texture map used for image rendering of the target object. A first image to be rendered is generated based on the first offset information and the texture map of the first region. Based on the second offset information, the three-dimensional Gaussian representation of the second region is transformed from static space to dynamic space, thereby rendering the second image to be rendered. Finally, the first image to be rendered is blended with the first image to be rendered to obtain the target rendered image, which improves the accuracy of image rendering.
[0278] The above description is merely an embodiment of this application and is not intended to limit the scope of protection of this application. Any modifications, equivalent substitutions, and improvements made within the spirit and scope of this application are included within the scope of protection of this application.
Claims
1. An image processing method, characterized in that, The method includes: Obtain control signals and viewing angle vectors for the target object model; Based on the control signal, a dynamic neural texture map of the first region in the appearance dimension of the target object model is obtained by decoding. Decoding is performed based on the observation view vector to obtain the view neural texture map of the first region in the view dimension; The dynamic neural texture map, the visual neural texture map, and the pre-trained diffuse neural texture map are fused to obtain the texture map of the first region.
2. The method according to claim 1, characterized in that, The process of decoding based on the control signal to obtain the dynamic neural texture map of the first region in the appearance dimension of the target object model includes: Feature extraction is performed on the appearance parameters and pose parameters in the control signal to obtain the feature vectors of the appearance parameters and the pose parameters, respectively. The feature vectors of the appearance parameters and the feature vectors of the pose parameters are decoded to obtain a dynamic neural texture map that changes synchronously with the changes in the appearance dimension of the first region in the target object model.
3. The method according to claim 1, characterized in that, The process of decoding based on the observation view vector to obtain the view neural texture map of the first region in the target object model in the view dimension includes: Feature extraction is performed on the observation view vector to obtain the feature vector of the observation view vector; The feature vector of the observation view vector is decoded to obtain a view neural texture map that changes synchronously with the change in the view dimension of the first region.
4. The method according to any one of claims 1 to 3, characterized in that, Before fusing the dynamic neural texture map, the viewpoint neural texture map, and the pre-trained diffuse neural texture map to obtain the texture map of the first region, the method further includes: Obtain the initialized diffuse neural texture map; The dynamic neural texture map, the visual neural texture map, and the initialized diffuse neural texture map are fused to obtain the predicted texture map of the first region. Based on the shape parameters, appearance parameters, and pose parameters in the control signal, a first mesh is constructed for the first region, and the first offset information of the first region relative to the first mesh is obtained; Based on the first offset information, each vertex in the first grid is offset to obtain the second grid of the first region; According to a preset mapping relationship, the texture at each position in the predicted texture map is mapped to the vertex in the second grid corresponding to the position to obtain a third grid. The mapping relationship is used to characterize the correspondence between the position of the vertex in the second grid and the position of the texture in the predicted texture map. The third grid is rasterized to obtain the coordinates of each pixel in the first region and the corresponding predicted latent vector; Based on the coordinates of each pixel in the first region and the corresponding predicted latent vector, pixel-level decoding is performed to obtain the predicted rendered image of the first region. Based on the difference between the predicted rendered image and the preset reference image, the initialized diffuse neural texture map is updated to obtain the pre-trained diffuse neural texture map of the first region.
5. The method according to claim 4, characterized in that, The step of obtaining the first offset information of the first region relative to the first grid includes: Feature extraction is performed on the appearance parameters and the pose parameters to obtain the feature vectors of the appearance parameters and the pose parameters; The feature vectors of the appearance parameters and the pose parameters are decoded to obtain the first offset information of the first region relative to the first grid.
6. The method according to claim 4, characterized in that, After fusing the dynamic neural texture map, the viewpoint neural texture map, and the pre-trained diffuse neural texture map to obtain the texture map of the first region, the method further includes: Based on the first offset information and the texture map of the first region, a first image to be rendered for the first region is generated in the first grid of the first region.
7. The method according to claim 6, characterized in that, The step of generating a first image to be rendered in the mesh of the first region based on the first offset information and the texture map of the first region includes: According to the mapping relationship, the texture at each position in the texture map is mapped to the vertex in the second mesh corresponding to the position to obtain the mesh to be rendered; The mesh to be rendered is rasterized to obtain the coordinates of each pixel in the first region and the corresponding latent vector. The first image to be rendered in the first region is obtained by decoding the coordinates of each pixel in the first region and the corresponding latent vector at the pixel level.
8. The method according to claim 7, characterized in that, The step of decoding at the pixel level based on the coordinates of each pixel in the first region and its corresponding latent vector to obtain the predicted rendering image of the first region includes: Feature extraction is performed on the coordinates and corresponding latent vector of each pixel to obtain the feature vector of each pixel; A linear transformation is performed on the feature vector of each pixel to obtain a linear feature vector; The linear feature vector is non-linearly mapped to obtain the color value of each pixel. A predicted rendered image of the first region is generated based on the color value of each pixel.
9. The method according to any one of claims 6 to 8, characterized in that, After generating the first image to be rendered in the mesh of the first region based on the first offset information and the texture map of the first region, the method further includes: Obtain the static three-dimensional Gaussian representation of the second region of the target object model, and obtain the second offset information of the static three-dimensional Gaussian representation; Based on the second offset information, the static three-dimensional Gaussian representation is transformed from the static space to the dynamic space to obtain the dynamic three-dimensional Gaussian representation; A second image to be rendered is generated based on the dynamic three-dimensional Gaussian representation; The first image to be rendered and the second image to be rendered are blended to obtain the target rendered image of the target object model.
10. The method according to claim 9, characterized in that, The step of obtaining the second offset information of the static three-dimensional Gaussian representation includes: Determine the overlapping area between the first region and the second region; Determine the rigid transformation information between the sampling points of the overlapping region in the static space and the matching points corresponding to the sampling points in the dynamic space; The appearance parameters in the control signal are extracted to obtain the feature vector of the appearance parameters; By performing a nonlinear mapping on the feature vectors of the appearance parameters, the non-rigid offset information of the static three-dimensional Gaussian representation is obtained; The rigid transformation information and the non-rigid offset information are added together to obtain the second offset information of the static three-dimensional Gaussian representation.
11. The method according to claim 9, characterized in that, The step of blending the first image to be rendered and the second image to be rendered to obtain the target rendered image of the target object model includes: Determine the first rendering depth corresponding to each pixel in the first image to be rendered and the second rendering depth corresponding to each pixel in the second image to be rendered; In response to the second rendering depth of the pixel at the same position being greater than or equal to the first rendering depth, the color value at the position corresponding to the first rendering depth is used as the color value of the pixel at the corresponding position in the target rendering image of the target object model. In response to the second rendering depth of the pixel at the same position being less than the first rendering depth, the color value of the position corresponding to the second rendering depth is used as the color value of the pixel at the corresponding position in the target rendering image of the target object model. For each pixel at a different location, the color corresponding to that location is used as the color of the pixel at that location in the first target rendering image of the target object model; A target rendering image of the target object model is generated based on the color value of the pixel at each location of the target rendering image.
12. An image processing apparatus, characterized in that, The device includes: The acquisition module is used to acquire control signals and viewing angle vectors for the target object model. The decoding module is used to decode based on the control signal to obtain a dynamic neural texture map of the first region in the appearance dimension of the target object model; and to decode based on the viewing angle vector to obtain a viewing angle neural texture map of the first region in the viewing angle dimension. The fusion module is used to fuse the dynamic neural texture map, the visual neural texture map, and the pre-trained diffuse neural texture map to obtain the texture map of the first region.
13. An electronic device, characterized in that, The electronic device includes: Memory is used to store executable instructions for a computer; A processor, when executing computer-executable instructions stored in the memory, implements the image processing method according to any one of claims 1 to 11.
14. A computer-readable storage medium storing computer-executable instructions or a computer program, characterized in that, When the computer-executable instructions or computer program are executed by a processor, the image processing method according to any one of claims 1 to 11 is implemented.
15. A computer program product comprising computer-executable instructions or a computer program, characterized in that, When the computer-executable instructions or computer program are executed by a processor, the image processing method according to any one of claims 1 to 11 is implemented.