Image data processing method and device, equipment and storage medium

By combining ray tracing technology and facial behavior coefficients with prior features of facial features, this method solves the problem that the relationship between facial expression features and spatial regions is not fully considered in traditional head avatar reconstruction, and achieves efficient and high-quality facial rendering effects.

CN120833432APending Publication Date: 2025-10-24GUANGDONG LAB OF ARTIFICIAL INTELLIGENCE & DIGITAL ECONOMY (SZ)
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510717914.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-29
Publication Date
2025-10-24

AI Technical Summary

Technical Problem

Traditional head avatar reconstruction methods fail to fully consider the complex relationship between expression features and spatial regions, making it difficult to achieve high-fidelity reconstruction effects when dealing with complex scenes and diverse expressions.

Method used

Ray tracing technology is used to obtain sampling point data and view direction vector of the face in virtual 3D space. Combined with facial behavior coefficients and prior features of facial features, dynamic conditional feature vectors are obtained by splicing them together to predict the color feature vector of the face, and rendering is performed based on this.

Benefits of technology

It achieves highly realistic facial rendering effects, improving the efficiency and quality of avatar image creation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120833432A_ABST
    Figure CN120833432A_ABST
Patent Text Reader

Abstract

The invention discloses an image data processing method and device, equipment and a storage medium. The method comprises the following steps: acquiring sampling point data of a face in a face image in a virtual three-dimensional space and a visual angle direction vector of light passing through the face based on a ray tracing technology; performing facial behavior coefficient extraction on the face image to obtain a facial behavior coefficient; acquiring five-sense-organ prior features of the face image from the face five-sense-organ dictionary; splicing based on the facial behavior coefficient and the five-sense-organ prior features to obtain a dynamic condition feature vector; predicting to obtain a color feature vector of the face based on the multiple sampling groups, the dynamic condition feature vector and the visual angle direction vector; according to the method, the to-be-rendered face is rendered on the basis of the color feature vectors of the multiple sampling groups, the dynamic condition feature vectors and the visual angle direction vectors, the to-be-rendered face is rendered on the basis of the color feature vectors of the face, and the high-fidelity target rendered image is obtained. And the efficiency and the quality of avatar image creation are improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of computer, and in particular to an image data processing method and device, equipment and storage medium. BACKGROUND

[0002] In today's rapidly developing digital era, there is an increasing demand for high-quality head avatar reconstruction technology in many fields such as virtual reality, game development, and special effects. Head avatar reconstruction aims to create a realistic and expressive virtual head model to achieve a more immersive interactive experience and visual effect.

[0003] Traditional head avatar reconstruction methods mainly rely on multi-view stereo technology and photogrammetry, which do not fully consider the complex relationship between expression features and spatial regions. When dealing with complex scenes and diverse expressions, it is often difficult to achieve high-fidelity reconstruction results. SUMMARY

[0004] The present application provides an image data processing method, device, computer equipment and storage medium, which solves the problem that the complex relationship between expression features and spatial regions is not fully considered, and it is often difficult to achieve high-fidelity reconstruction results when dealing with complex scenes and diverse expressions.

[0005] In a first aspect, an image data processing method is provided, comprising:

[0006] Based on the ray tracing technology, the sampling point data of the face in the virtual three-dimensional space and the perspective direction vector of the light ray passing through the face in the face image are obtained, and the sampling point data includes a sampling group of each pixel point on the face, and the sampling group is obtained by collecting a preset number of sampling points of the light ray passing through the pixel point on the face;

[0007] The face behavior coefficient of the face image is extracted to obtain the face behavior coefficient;

[0008] The five features of the face image are obtained from the face feature dictionary;

[0009] Based on the face behavior coefficient and the five features, a dynamic conditional feature vector is obtained;

[0010] Based on a plurality of sampling groups, the dynamic conditional feature vector and the perspective direction vector, a color feature vector of the face is predicted, and the color feature vector includes a color feature representation and a body density of the pixel point on the face;

[0011] Based on the color feature vector, the face to be rendered is rendered to obtain a target rendering image.

[0012] In a second aspect, an image data processing device is provided, comprising:

[0013] a first obtaining module, configured to obtain, based on a ray tracing technology, sample point data of a face in a virtual three-dimensional space in a face image and a view direction vector of a light ray passing through the face, the sample point data comprising a sampling group of each pixel point on the face, the sampling group being obtained by collecting a preset number of sample points of the light ray passing through the pixel point on the face;

[0014] an extraction module, configured to extract a facial behavior coefficient from the face image to obtain the facial behavior coefficient;

[0015] a second obtaining module, configured to obtain, from a face feature dictionary, a feature of a facial feature of the face image;

[0016] a splicing module, configured to splice the facial behavior coefficient and the feature of the facial feature to obtain a dynamic conditional feature vector;

[0017] a prediction module, configured to predict a color feature vector of the face based on a plurality of sets of the sampling group, the dynamic conditional feature vector and the view direction vector, the color feature vector comprising a color feature representation and a body density of a pixel point on the face;

[0018] a rendering module, configured to render a face to be rendered based on the color feature vector to obtain a target rendering image.

[0019] Optionally, the extraction module comprises a splicing submodule, and the splicing submodule is configured to:

[0020] extract an expression coefficient and an eye movement coefficient from the face image to obtain the facial behavior coefficient;

[0021] splice the expression coefficient and the eye movement coefficient to obtain the facial behavior coefficient.

[0022] Optionally, the second obtaining module comprises an obtaining submodule, and the obtaining submodule is configured to:

[0023] extract a feature of a facial feature from the face image based on a region alignment operation;

[0024] obtain the feature of the facial feature from a face feature dictionary based on the feature of the facial feature.

[0025] Optionally, the prediction module comprises:

[0026] an encoding submodule, configured to encode the plurality of sets of the sampling group to obtain a spatial geometric feature vector;

[0027] The first prediction submodule is configured to input the spatial geometry vector into a lightweight multi-layer perception to perform weight prediction, so as to obtain a weight vector with the same dimension as the dynamic condition feature vector;

[0028] The operation submodule is configured to perform Hadamard product operation on the dynamic condition feature vector and the weight vector, so as to obtain a region perception feature vector.

[0029] The second prediction submodule is configured to predict a color feature vector of the face based on the spatial geometry vector, the region perception feature vector and the view direction vector.

[0030] Optionally, the encoding submodule comprises:

[0031] The projection unit is configured to project the sampling points in the plurality of sampling groups to a plurality of two-dimensional planes, each of the two-dimensional planes comprising a plurality of projection faces with different scales;

[0032] The calculation unit is configured to perform linear difference value calculation on each projection point on each projection face on each two-dimensional plane based on a linear difference value algorithm, so as to obtain a first feature representation of each projection face in each two-dimensional plane;

[0033] The first splicing unit is configured to splice the first feature representation of each projection face in each two-dimensional plane, so as to obtain a second feature representation of each two-dimensional plane;

[0034] The second splicing unit is configured to splice the second feature representation of each two-dimensional plane, so as to obtain the spatial geometry feature vector.

[0035] Optionally, the number of the sampling groups is less than the number of pixel points of the face image, and the device further comprises:

[0036] The up-sampling processing module is configured to perform up-sampling processing on the target rendering image.

[0037] Optionally, the device further comprises:

[0038] The pre-processing operation module is configured to perform image pre-processing operation on the face image, the image pre-processing operation comprising image cropping and image background setting.

[0039] In a third aspect, a computer device is provided, which comprises a memory, a processor and a computer program stored in the memory and executable on the processor, and the processor implements the steps of the image data processing method when executing the computer program.

[0040] In a fourth aspect, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above-mentioned image data processing method are implemented.

[0041] The present application provides an image data processing method, apparatus, computer device, and storage medium. The method uses ray tracing technology to obtain sampling point data of a face in a virtual three-dimensional space and the viewing direction vector of a light ray passing through the face. The sampling point data includes a sampling group for each pixel on the face, and the sampling group is obtained by collecting a preset number of sampling points of light rays passing through the pixel points on the face. The method also extracts facial behavior coefficients from the facial image to obtain facial behavior coefficients. Prior features of the facial features of the facial image are obtained from a dictionary of facial features. A dynamic conditional feature vector is obtained by concatenating the facial behavior coefficients and the prior features of the facial features. A color feature vector of the face is predicted based on multiple groups of the sampling groups, the dynamic conditional feature vectors, and the viewing direction vector. The color feature vector includes color feature representations and volume density of the pixels on the face. The method then renders the face to be rendered based on the color feature vector to obtain a target rendered image. The method uses ray tracing technology to accurately obtain sampling point data of a face in a virtual three-dimensional space and the viewing direction vector of a light ray passing through the face. Furthermore, by extracting facial behavior coefficients from facial images, key parameters reflecting individual expression changes are obtained. Combined with prior features of facial features obtained from a dictionary of facial features, dynamic conditional feature vectors can be spliced ​​together. These dynamic conditional feature vectors, combined with prior knowledge of expression changes and facial structure, can predict facial color feature vectors based on multiple sampling groups, the dynamic conditional feature vectors, and the viewing direction vector, providing a foundation for realistic facial rendering. Ultimately, based on these color feature vectors, the face to be rendered is rendered, resulting in a highly realistic target rendered image, improving the efficiency and quality of avatar image creation. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0043] Figure 1 A diagram illustrating an application environment of the image data processing method provided in an embodiment of the present application;

[0044] Figure 2 A flowchart of the image data processing method provided in an embodiment of the present application;

[0045] Figure 3 A light ray schematic diagram of a face image in a virtual three-dimensional space provided by an embodiment of the present application;

[0046] Figure 4 A structure block diagram of an image data processing apparatus provided by an embodiment of the present application;

[0047] Figure 5 A structure block diagram of a computer device provided by an embodiment of the present application. DETAILED DESCRIPTION

[0048] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are some but not all of the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the scope of the present application.

[0049] In addition, the described features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are provided to give a thorough understanding of embodiments of the present application. One of ordinary skill in the art, however, will recognize that the application can be practiced without one or more of the specific details, or with other methods, components, devices, steps, etc. In other instances, well-known methods, devices, implementations, or operations are not shown or described in detail to avoid obscuring aspects of the present application.

[0050] The block diagrams shown in the drawings are only functional entities, and do not necessarily correspond to physically independent entities. That is, the functional entities can be implemented in the form of software, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.

[0051] The flowcharts shown in the drawings are only exemplary illustrations, and do not necessarily include all contents and operations / steps, nor are they necessarily executed in the described order. For example, some operations / steps can be further decomposed, and some operations / steps can be combined or partially combined, so the actual execution order can be changed according to actual conditions.

[0052] The image data processing method provided by the embodiments of the present application can be applied to, for example, Figure 1application environment. Among them, the computer device 110 communicates with the server 120 through the network 130. The computer device 110 can obtain the sampling point data of the face in the virtual three-dimensional space in the face image and the perspective direction vector of the light ray passing through the face based on the ray tracing technology. The sampling point data includes a sampling group of each pixel point on the face, and the sampling group is obtained by collecting a preset number of sampling points of the light ray passing through the pixel point of the face. The face behavior coefficient of the face image is extracted to obtain the face behavior coefficient. The five feature prior features of the face image are obtained from the face five feature dictionary. The dynamic condition feature vector is spliced based on the face behavior coefficient and the five feature prior features. The color feature vector of the face is predicted based on multiple groups of the sampling group, the dynamic condition feature vector and the perspective direction vector. The color feature vector includes the color feature representation and the body density of the pixel point on the face. The target rendering image is obtained by rendering the face to be rendered based on the color feature vector, and the computer device 110 is used for display. In the present application, the sampling point data of the face in the virtual three-dimensional space in the face image and the perspective direction vector of the light ray passing through the face can be accurately obtained based on the method of ray tracing technology. Further, the key parameters reflecting individual expression changes are obtained by extracting the face behavior coefficient of the face image. Combined with the five feature prior features obtained from the face five feature dictionary, the dynamic condition feature vector can be spliced. The dynamic condition feature vector combines the prior knowledge of expression changes and five feature structures, and can predict the color feature vector of the face based on multiple groups of the sampling group, the dynamic condition feature vector and the perspective direction vector, which provides a basis for realistic rendering of the face. Finally, based on these color feature vectors, the face to be rendered is rendered to obtain a high-fidelity target rendering image, improving the efficiency and quality of avatar image creation. Among them, the computer device 110 can be but not limited to various smart phones 110-1, tablet computers 110-2 and notebook computers 110-3. The present application will be described in detail below through specific examples.

[0053] Please refer to Figure 2 as shown, Figure 2 A flowchart of an image data processing method provided by an embodiment of the present application is shown. The method can be applied to a terminal or a server. The embodiment is illustrated by taking the server as an example. The image data processing method comprises the following steps:

[0054] S101: Based on the ray tracing technology, the sampling point data of the face in the virtual three-dimensional space in the face image and the perspective direction vector of the light ray passing through the face are obtained. The sampling point data includes a sampling group of each pixel point on the face, and the sampling group is obtained by collecting a preset number of sampling points of the light ray passing through the pixel point of the face.

[0055] Ray Tracing Technology (RTT) is a method used in computer graphics to generate realistic images. It creates images by simulating the behavior of light rays in a virtual three-dimensional environment, including how light rays interact with object surfaces (such as reflection, refraction, scattering, etc.), and how they are affected by the environment (such as shadows, global illumination, etc.). Ray tracing technology can be implemented through various programming languages and graphics libraries, such as C++, Python, OpenGL, DirectX, Vulkan, etc.

[0056] In this application, the light rays of the face image in the virtual three-dimensional space can be simulated by the ray tracing technology, and the pixel points on the face of the face image at least pass through a light ray. The sampling range (t1, t2) of the light ray can be determined based on experience. For example, please refer to Figure 3 where t1 represents the sampling point on light ray 1 close to the light ray source device (usually a camera), and t2 represents the sampling point on light ray 1 away from the light ray source device (not shown in the figure). A predetermined number of sampling points (such as 64 sampling points) are uniformly sampled in this sampling range, obtaining the sampling group of light ray 1. Similarly, the sampling group of the light ray passing through each pixel point on the face can be obtained, so that the sampling point data can be obtained. Assuming that all light rays originate from the point P (0, 0, 0) of the light ray source device, the viewing direction of light ray 1 can be represented as a represents the direction of light ray 1 on the x-axis, b represents the direction of light ray 1 on the y-axis, and c represents the direction of light ray 1 on the z-axis. The viewing direction of light ray 1 is represented as

[0057]

[0058] where is represented as

[0059] Similarly, the viewing directions of the remaining light rays can be obtained, and thus the viewing direction vector [viewing direction of light ray 1, viewing direction of light ray 2,..., viewing direction of light ray 512] can be obtained according to all viewing directions.

[0060] In an embodiment, before the method of obtaining the sampling point data of the face in the virtual three-dimensional space in the face image and the viewing direction vector of the light ray passing through the face based on the ray tracing technology, the method further comprises:

[0061] performing image preprocessing operations on the face image, the image preprocessing operations including image cropping and image background setting.

[0062] ​The face image can be cropped to a preset target size, and the background color in the face image is set. The uniform size is beneficial to reduce the consumption of computing resources, make the face on the face image more clear, avoid the interference of background pixels, and improve the subsequent image data processing speed and analysis accuracy. For example, the face image is cropped according to the preset target size 512x 512 to obtain a face image with a size of 512x 512, and the background color of the cropped face image is set to white.

[0063] S102: face behavior coefficient extraction is performed on the face image to obtain a face behavior coefficient.

[0064] The face behavior coefficient is a quantitative index for measuring the degree of facial expression and can be used to identify and distinguish different facial expressions and evaluate the intensity of the expression.

[0065] For example, facial recognition technology (such as key point detection) is used to extract facial features such as the position and shape of the eyes, nose, and mouth; according to the Facial Action Coding System (FACS), Action Units (AUs) are detected, which are the basic muscle movements that make up facial expressions; the intensity of each action unit is evaluated, usually using a value between 0 and 1, where 0 indicates no action unit and 1 indicates the maximum intensity of the action unit. The intensity values of all action units are combined into an expression coefficient vector to describe the entire facial expression.

[0066] In an embodiment, the face behavior coefficient extraction on the face image includes:

[0067] The face behavior coefficient extraction on the face image includes:

[0068] The expression coefficient and the eye movement coefficient are spliced to obtain the face behavior coefficient.

[0069] The EMOCA tool is a tool that can generate a 3D reconstruction from a single face image. EMOCA can accurately capture and convey the emotional state of the input face image through single face image input. EMOCA tool introduces a new depth perception emotion consistency loss to ensure that the reconstructed 3D expression matches the expression in the input face image during training, thereby significantly improving the accuracy of monocular face reconstruction, especially in capturing subtle or extreme expressions.

[0070] OpenFace is an open-source toolkit that can perform facial landmark detection, head pose estimation, facial action unit recognition, and eye gaze estimation. OpenFace provides the recognition function of the Facial Action Coding System (FACS), which can recognize the intensity of 17 action units (from 0 to 5) and the existence of 18 action units (0 represents nonexistence, and 1 represents existence). In addition, OpenFace can also estimate the gaze direction of the eyes, which can be used to analyze eye movement. Specifically, OpenFace can directly extract facial expression action units from a face image and give the existence and intensity scores of each action unit. For the existence of action units, the column in the output file will encode 0 as nonexistence and 1 as existence. For the intensity of action units, the column in the output file ranges from 0 (nonexistence), 1 (existence with minimum intensity), to 5 (existence with maximum intensity).

[0071] In this embodiment, the expression coefficients and eye movement coefficients of all face images can be extracted using the emoca and openface tools respectively, and then the expression coefficients and eye movement coefficients are spliced together to obtain the facial behavior coefficients.

[0072] The face image is processed using the EMOCA tool. EMOCA outputs a series of expression-related coefficients that describe the activity level of different facial muscle groups. For example: expression coefficient 1 (smile): 0.8 (indicating a strong degree of smiling); expression coefficient 2 (frown): 0.2 (indicating a weak degree of frowning); expression coefficient 3 (eye opening degree): 0.9 (indicating that the eyes are almost completely open); expression coefficient 4 (mouth opening degree): 0.5 (indicating that the mouth is half open).

[0073] The face image is processed using the OpenFace tool to extract eye movement coefficients. OpenFace may output coefficients related to eye movement, such as: eye movement coefficient 1 (left eye gaze direction X): 0.1 (indicating that the left eye is gazing to the right); eye movement coefficient 2 (left eye gaze direction Y): 0.05 (indicating that the left eye's line of sight is slightly upward); eye movement coefficient 3 (right eye gaze direction X): 0.1 (indicating that the right eye is gazing to the right); eye movement coefficient 4 (right eye gaze direction Y): 0.05 (indicating that the right eye's line of sight is slightly upward).

[0074] The expression coefficients and eye movement coefficients obtained from EMOCA and OpenFace are spliced together to form a comprehensive facial behavior coefficient vector: facial behavior coefficients: [0.8, 0.2, 0.9, 0.5, 0.1, 0.05, 0.1, 0.05]

[0075] S103: Obtain the prior features of the facial features of the face image from the facial feature dictionary.

[0076] The facial feature dictionary is a dictionary storing feature information of eye, nose, mouth and other feature regions, and provides prior knowledge of facial feature details. The dictionary can provide reference for missing or blurred features, and enhance the accuracy of feature details. The facial feature dictionary contains depth features of different facial regions, and each facial region contains three different levels of features, so that it can reflect local details and maintain overall consistency. Taking the eye region features as an example, the three different levels of features include: the coarse level describes the basic shape and position of the eye, for example, including the approximate rectangular bounding box of the eye; the binary state of the eye opening or closing. The medium level describes the shape of the eyelid and the outline of the eye, for example, including the curve features of the eye corner point, upper eyelid and lower eyelid; the preliminary estimation of the eye gaze direction. The fine level describes the details of the iris, pupil and white of the eye. For example, including the texture features of the eye, such as wrinkles and fine wrinkles of the eyelid; the accurate estimation of the eye gaze direction.

[0077] The local feature of the facial feature can be extracted from the facial image, and then similarity calculation is performed based on the local feature of the facial feature and the feature of the facial feature region in the facial feature dictionary. If the calculated similarity meets the similarity threshold, it indicates that the local feature of the facial feature and the feature of the facial feature region are similar, and the feature of the facial feature region is taken as the prior feature of the facial feature.

[0078] In an embodiment, the obtaining of the prior feature of the facial feature from the facial feature dictionary comprises:

[0079] extracting the local feature of the facial feature from the facial image based on a region alignment operation;

[0080] obtaining the prior feature of the facial feature from the facial feature dictionary based on the local feature of the facial feature.

[0081] Through the region alignment (RoIAlign) operation, the local feature of the facial feature can be extracted from the facial image, matched with the feature of the facial feature region in the facial feature dictionary, and the matched feature of the facial feature region is taken as the prior feature of the facial feature to obtain more detailed facial feature details.

[0082] S104: Splicing a dynamic condition feature vector based on the facial behavior coefficient and the prior feature of the facial feature.

[0083] The facial behavior coefficient is combined with the prior feature of the facial feature read from the facial feature dictionary to form a final dynamic condition feature vector (efp) for driving facial expression changes.

[0084] The splicing process can be feature connection of the facial behavior coefficient and the prior feature of the facial feature, or a feature fusion method, such as using a neural network to integrate the facial behavior coefficient and the prior feature of the facial feature.

[0085] Exemplarily, the facial behavior coefficients include: expression coefficients: [0.8 (smile intensity), 0.2 (frown intensity)]; eye movement coefficients: [0.1 (left eye gaze direction X), 0.05 (left eye gaze direction Y)]; facial feature prior characteristics: [eye shape feature, nose shape feature, mouth shape feature]; and the spliced dynamic condition feature vector can be represented as [0.8, 0.2, 0.1, 0.05, eye shape feature, nose shape feature, mouth shape feature].

[0086] S105: predicting a color feature vector of the face based on the multiple groups of sampling groups, the dynamic condition feature vector, and the perspective direction vector, the color feature vector including color feature representation and volume density of a pixel point on the face.

[0087] The color feature vector includes color feature representation and volume density of each pixel point. The color feature representation describes the color attribute of the pixel point, for example, the position of the pixel point in a color space (such as RGB color space, HSV color space, LAB color space), the intensity of the color, etc. The volume density represents the transparency of the pixel point.

[0088] The multiple groups of sampling groups, the dynamic condition feature vector, and the perspective direction vector can be used as a pre-trained neural network to predict facial features to obtain a color feature vector of the face, the color feature vector including color feature representation and volume density of a pixel point on the face. In this process, the color feature representation of all sampling points in each group of sampling groups can be predicted, and the color feature representation of all sampling points in a group of sampling groups can be averaged (such as weighted average, the weight can be the volume density corresponding to each sampling point) to obtain the color feature representation and volume density of the pixel point corresponding to the sampling group. Similarly, the color feature representation and volume density of all sampling groups can be obtained, so as to construct the color feature vector of the face based on the color feature representation and volume density of each sampling group.

[0089] In an embodiment, the predicting the color feature vector of the face based on the multiple groups of sampling groups, the dynamic condition feature vector, and the perspective direction vector includes:

[0090] encoding the multiple groups of sampling groups to obtain a spatial geometry feature vector;

[0091] inputting the spatial geometry vector into a lightweight multilayer perceptron to predict a weight vector with the same dimension as the dynamic condition feature vector;

[0092] performing Hadamard product operation on the dynamic condition feature vector and the weight vector to obtain a region perception feature vector;

[0093] predict the color feature vector of the face based on the spatial geometry vector, the region-aware feature vector, and the view direction vector.

[0094] wherein, after encoding and calculating each sampling point in each group of sampling groups, a geometry feature of each sampling point is obtained, which reflects the position information of the sampling point in three-dimensional space, and finally a spatial geometry feature vector is formed based on the geometry features of the sampling points. The weight vector includes the weight of each sampling point, which is predicted based on the geometry feature of each sampling point. The lightweight multilayer perceptron is a simplified neural network structure, which usually contains fewer hidden layers and neurons to reduce computational complexity and memory usage. In this application, the mapping relationship between the input features (i.e. the spatial geometry vector) and the output weight vector is learned through the lightweight multilayer perceptron. Specifically, the spatial geometry vector is input into the lightweight multilayer perceptron for weight prediction, and a weight vector is output.

[0095] In order to more accurately capture the changes of facial expressions in different regions (such as eye, nose, mouth, etc.), an example of a three-plane hash encoder (H3) can be used, such as a multi-resolution hash encoder, to perform multi-scale encoding and calculation on the feature information (position coordinates and direction) of the sampling points in multiple groups of sampling groups on each plane to obtain a spatial geometry feature vector. The spatial geometry feature vector is input into a double-layer perceptron (MLP) to predict the weight corresponding to each sampling point, which can be represented as (We, x), where We represents the weight and x represents the geometry feature of the sampling point. In this way, a weight vector can be formed based on the weights of the sampling points, and the size of the weight vector is the same as the dimension of the dynamic condition feature vector (efp). The dynamic condition feature vector and the weight vector are subjected to Hadamard product operation, i.e. the elements at the same position in the dynamic condition feature vector and the weight vector are multiplied, thereby obtaining a region-aware feature vector. The color feature vector of the face is predicted based on the spatial geometry vector, the region-aware feature vector, and the view direction vector using a neural network.

[0096] In this embodiment, the degree of attention to different regions of the face can be adaptively adjusted according to the changes of the expression, so as to more accurately capture the subtle changes of the facial expression.

[0097] In an embodiment, the encoding and calculating the multiple groups of sampling groups to obtain a spatial geometry feature vector comprises:

[0098] projecting the sampling points in the multiple groups of sampling groups to multiple two-dimensional planes, each of the two-dimensional planes including multiple projection planes of different scales;

[0099] performing linear difference calculation on each projection point on each projection plane of each of the two-dimensional planes based on a linear difference algorithm to obtain a first feature representation of each projection plane in each of the two-dimensional planes;

[0100] stitching the first feature representation of each projection plane in each of the two-dimensional planes to obtain a second feature representation of each of the two-dimensional planes;

[0101] stitching the second feature representation of each of the two-dimensional planes to obtain the spatial geometric feature vector.

[0102] The linear difference algorithm can be a bilinear difference algorithm.

[0103] To improve the calculation speed, the sampling points in the three-dimensional space are encoded based on a "three-plane hash representation method" to obtain a spatial geometric feature vector. Specifically, the sampling points in the three-dimensional space are projected onto three two-dimensional planes, such as the X-Y plane, the Y-Z plane, and the X-Z plane, and then a three two-dimensional multi-resolution hash encoder tool (the hash encoder tool is implemented based on a linear difference algorithm) is used to perform linear difference calculation on the projection points on each two-dimensional plane. The multi-resolution hash encoder tool is used to capture the geometric features at different scales (such as 4*4 scale and 8*8 scale) of each two-dimensional plane as a first feature representation at each scale. After stitching the first feature representation of each scale of a two-dimensional plane, a second feature representation of the two-dimensional plane is obtained. Finally, the second feature representations of the three planes are stitched together to obtain a spatial geometric feature vector.

[0104] S106: rendering a face to be rendered based on the color feature vector to obtain a target rendering image.

[0105] The face to be rendered represents a three-dimensional face to be rendered. The face to be rendered is rendered by the color feature vector to render the expression in the face image onto the face to be rendered to obtain a target rendering image.

[0106] A neural network, such as a multilayer perceptron decoder (MLP), can be used to map the color feature representation and the body density of each pixel point in the color feature vector to the face to be rendered to obtain a target rendering image.

[0107] In an embodiment, the number of the sampling groups is less than the number of pixel points of the face image. After the face to be rendered is rendered based on the color feature vector to obtain a target rendering image, the method further includes:

[0108] performing up-sampling processing on the target rendering image.

[0109] The number of the sampling groups is less than the number of the pixel points of the face image, which means that the light of all the pixel points is not sampled during sampling, so as to reduce the data amount and the calculation amount, thereby improving the image data processing efficiency. However, in order to make the obtained target rendering image meet the source resolution, the target rendering image can be up-sampled. For example, assuming that the number of the pixel points of the face image is 512*512 and the number of the sampling groups is 128*128, the calculation amount can be reduced, and a target rendering image of a low resolution face image of 128*128 can be quickly generated. Then, the neural network is used to up-sample the resolution image, and a target rendering image of 512*512 which has the same resolution as the face image is obtained.

[0110] In the embodiment, when the number of the sampling groups is less than the number of the pixel points of the face image, the image processing calculation speed can be improved, and then the color feature vectors of the sampling points are restored to the face to be rendered by up-sampling the target rendering image, so as to enhance the face details and clarity, and make the target rendering image closer to the real face image.

[0111] Taking an example of cutting a face picture to a preset target size of 512*512, an image data processing method is provided, which is specifically as follows:

[0112] A face picture is obtained, the face picture is cut to 512*512, and a white background is set.

[0113] Expression coefficients and gaze coefficients are extracted by EMOCA and OpenFace tools, and are spliced to obtain face behavior coefficients;

[0114] Sampling points of the face image in a three-dimensional space are obtained based on a ray tracing method;

[0115] Points in the three-dimensional space are projected to three two-dimensional planes (XY, YZ, XZ), the projected points are encoded by using a multi-resolution hash encoder, the geometric features of the sampling points are captured, and a space geometric feature vector is spliced.

[0116] The facial features in a facial dictionary are read, and are matched with local facial features. The matched facial prior features are spliced with the face behavior coefficients to form a dynamic condition feature vector (efp, Expression Feature Prior), which is used to drive expression changes, so as to enhance the accuracy of facial details.

[0117] The spatial geometry feature vector, the view direction vector, and the dynamic condition feature vector are predicted by an MLP decoder to obtain a color feature vector of the face, which includes color feature representation and volume density of each sampling point.

[0118] The facial component dictionary stores feature information of eye, nose, mouth and other component regions. These features are divided into different scales (Scale-1, Scale-2, Scale-3) to ensure that both local details and overall consistency can be captured. Through RoIAlign, the corresponding component local features are extracted from the target face.

[0119] Optionally, the spatial geometry feature vector is predicted by a neural network such as MLP to generate a weight vector of the same dimension as the dynamic condition feature vector (efp, Expression Feature Prior).

[0120] The dynamic condition feature vector and the weight vector are subjected to Hadamard product operation to generate a region-aware feature vector (reweighted efp) to adaptively adjust the attention degree in key regions (such as eyes and mouth) and improve the expression detail capturing ability.

[0121] The region-aware dynamic feature vector, the spatial geometry feature vector, and the view direction vector are input into a neural network for prediction to obtain a color feature vector of the face, including a low-dimensional feature vector (Low-res Feature Map, 128x128x32), i.e., color feature representation and volume density (σ) of each pixel point, for volume rendering.

[0122] The low-dimensional feature vector (Low-res Feature Map, 128x128x32) and the volume density (σ) are rendered onto the face to be rendered by using volume rendering technology to generate a low-resolution face image.

[0123] Finally, the low-resolution face image is input into a super-resolution module (composed of a neural network) for up-sampling (Upsampling) to generate a high-resolution RGB portrait (512x512x3), enhancing the facial details and clarity, and making it closer to a real face image.

[0124] The above is the image data processing flow of the present application.

[0125] As above, the present application provides an image data processing method, device, computer equipment and storage medium. The sampling point data of a face in a virtual three-dimensional space and the perspective direction vector of a light ray passing through the face in a face image are obtained based on a ray tracing technology. The sampling point data includes a sampling group of each pixel point on the face, and the sampling group is obtained by collecting a preset number of sampling points of a light ray passing through the pixel point. A face behavior coefficient of the face image is extracted to obtain a face behavior coefficient. The prior feature of the five features of the face image is obtained from a face five feature dictionary. A dynamic conditional feature vector is spliced based on the face behavior coefficient and the prior feature of the five features. A color feature vector of the face is predicted based on a plurality of sampling groups, the dynamic conditional feature vector and the perspective direction vector. The color feature vector includes a color feature representation and a body density of the pixel point on the face. The face to be rendered is rendered based on the color feature vector to obtain a target rendering image. The method based on the ray tracing technology can accurately obtain the sampling point data of the face in the virtual three-dimensional space and the perspective direction vector of the light ray passing through the face in the face image. Further, the face behavior coefficient of the face image is extracted to obtain a key parameter reflecting individual expression changes. The dynamic conditional feature vector can be spliced by combining the prior feature of the five features obtained from the face five feature dictionary. The dynamic conditional feature vector combines the prior knowledge of expression changes and five feature structures, and can predict the color feature vector of the face based on a plurality of sampling groups, the dynamic conditional feature vector and the perspective direction vector, thereby providing a basis for realistic rendering of the face. Finally, the face to be rendered is rendered based on the color feature vector to obtain a high-fidelity target rendering image, thereby improving the efficiency and quality of avatar image creation.

[0126] It should be understood that the size of the serial number of each step in the above embodiment does not mean the order of execution. The execution order of each process should be determined by its function and inherent logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.

[0127] In an embodiment, an image data processing device is provided, which corresponds to the image data processing method in the above embodiment. Referring to FIG. 20, the image data processing device includes: Figure 4 As shown in the figure, the image data processing device includes:

[0128] The first acquisition module 201 is configured to obtain the sampling point data of a face in a virtual three-dimensional space and the perspective direction vector of a light ray passing through the face in a face image based on a ray tracing technology. The sampling point data includes a sampling group of each pixel point on the face, and the sampling group is obtained by collecting a preset number of sampling points of a light ray passing through the pixel point.

[0129] The extraction module 202 is configured to perform facial behavior coefficient extraction on the face image to obtain facial behavior coefficients.

[0130] The second acquisition module 203 is configured to acquire the facial feature prior from a face feature dictionary.

[0131] The splicing module 204 is configured to splice the facial behavior coefficients and the facial feature prior to obtain a dynamic condition feature vector.

[0132] The prediction module 205 is configured to predict a color feature vector of the face based on the multiple sampling groups, the dynamic condition feature vector and the view direction vector, wherein the color feature vector comprises color feature representation and volume density of a pixel point on the face.

[0133] The rendering module 206 is configured to render a face to be rendered based on the color feature vector to obtain a target rendering image.

[0134] In the embodiment, the sampling point data of the face in the face image in the virtual three-dimensional space and the view direction vector of the light ray passing through the face can be accurately obtained by the method based on the ray tracing technology. Further, the key parameters reflecting individual expression changes can be obtained by performing facial behavior coefficient extraction on the face image. In combination with the facial feature prior acquired from the face feature dictionary, the dynamic condition feature vector can be spliced. The dynamic condition feature vector combines the prior knowledge of expression changes and facial structure, and can predict the color feature vector of the face based on the multiple sampling groups, the dynamic condition feature vector and the view direction vector, thereby providing a basis for realistic rendering of the face. Finally, based on the color feature vectors, the face to be rendered is rendered to obtain a target rendering image with high fidelity, thereby improving the efficiency and quality of avatar image creation.

[0135] Optionally, the extraction module comprises a splicing sub-module, and the splicing sub-module is configured to:

[0136] perform facial behavior coefficient extraction on the face image to obtain expression coefficients and eye movement coefficients;

[0137] splice the expression coefficients and the eye movement coefficients to obtain the facial behavior coefficients.

[0138] Optionally, the second acquisition module comprises an acquisition sub-module, and the acquisition sub-module is configured to:

[0139] extract facial local features from the face image based on a region alignment operation;

[0140] acquire the facial feature prior from a face feature dictionary based on the facial local features.

[0141] Optionally, the prediction module comprises:

[0142] an encoding submodule, configured to encode the plurality of groups of sampling groups to obtain a spatial geometry feature vector;

[0143] a first prediction submodule, configured to input the spatial geometry vector into a lightweight multi-layer perception machine to perform weight prediction, to obtain a weight vector with the same dimension as the dynamic condition feature vector;

[0144] an operation submodule, configured to perform Hadamard product operation on the dynamic condition feature vector and the weight vector to obtain a region perception feature vector;

[0145] a second prediction submodule, configured to predict a color feature vector of the face based on the spatial geometry vector, the region perception feature vector and the view direction vector.

[0146] Optionally, the encoding submodule comprises:

[0147] a projection unit, configured to project sampling points in the plurality of groups of sampling groups to a plurality of two-dimensional planes, each of the two-dimensional planes comprising a plurality of projection faces with different scales;

[0148] a calculation unit, configured to perform linear difference value calculation on each projection point on each projection face on each of the two-dimensional planes based on a linear difference value algorithm, to obtain a first feature representation of each projection face in each of the two-dimensional planes;

[0149] a first splicing unit, configured to splice the first feature representation of each projection face in each of the two-dimensional planes to obtain a second feature representation of each of the two-dimensional planes;

[0150] a second splicing unit, configured to splice the second feature representation of each of the two-dimensional planes to obtain the spatial geometry feature vector.

[0151] Optionally, the number of the sampling groups is less than the number of pixel points of the face image, and the device further comprises:

[0152] an up-sampling processing module, configured to perform up-sampling processing on the target rendering image.

[0153] Optionally, the device further comprises:

[0154] a pre-processing operation module, configured to perform image pre-processing operation on the face image, the image pre-processing operation comprising image cropping and image background setting.

[0155] In one embodiment, a computer device is provided, and an internal structure diagram of the computer device can be as shown in Figure 5As shown in the figure. The computer device includes a processor, a memory, a network interface, a display screen and an input device connected by a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium, an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with the external server through the network connection. The computer program is executed by the processor to realize the function or step of the image data processing method.

[0156] In one embodiment, a computer device is proposed, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, the processor executing the computer program to implement the following steps:

[0157] Based on the light ray tracing technology, sample point data of the face in the virtual three-dimensional space and a perspective direction vector of a light ray passing through the face in the face image are obtained, the sample point data includes a sampling group of each pixel point on the face, and the sampling group is obtained by collecting a preset number of sample points of the light ray passing through the pixel point on the face; a face behavior coefficient of the face image is extracted to obtain the face behavior coefficient; a five-feature prior feature of the face image is obtained from a face five-feature dictionary; a dynamic conditional feature vector is spliced based on the face behavior coefficient and the five-feature prior feature; a color feature vector of the face is predicted based on the multiple sampling groups, the dynamic conditional feature vector and the perspective direction vector, and the color feature vector includes a color feature representation and a body density of the pixel point on the face; and the face to be rendered is rendered based on the color feature vector to obtain a target rendering image.

[0158] In the embodiment, by the method based on the light ray tracing technology, the sample point data of the face in the virtual three-dimensional space and the perspective direction vector of the light ray passing through the face in the face image can be accurately obtained. Further, by extracting the face behavior coefficient of the face image, a key parameter reflecting individual expression changes is obtained. Combined with the five-feature prior feature obtained from the face five-feature dictionary, a dynamic conditional feature vector can be spliced, the dynamic conditional feature vector combines the prior knowledge of expression changes and five-feature structure, and based on the multiple sampling groups, the dynamic conditional feature vector and the perspective direction vector, the color feature vector of the face can be predicted, which provides a basis for realistic rendering of the face. Finally, based on the color feature vector, the face to be rendered is rendered to obtain a high-fidelity target rendering image, improving the efficiency and quality of avatar image creation.

[0159] In one embodiment, a computer readable storage medium is proposed, the computer readable storage medium stores a computer program, and the computer program is executed by a processor to implement the following steps:

[0160] Based on the ray tracing technology, sample point data of a face in a virtual three-dimensional space in a face image and a perspective direction vector of a light ray passing through the face are obtained, the sample point data includes a sampling group of each pixel point on the face, and the sampling group is obtained by collecting a preset number of sample points of the light ray passing through the pixel point on the face; a face behavior coefficient of the face image is extracted to obtain a face behavior coefficient; a five-winknowledge dictionary of the face is used to obtain five-winknowledge prior features of the face image; a dynamic conditional feature vector is spliced based on the face behavior coefficient and the five-winknowledge prior features; a color feature vector of the face is predicted based on multiple sampling groups, the dynamic conditional feature vector, and the perspective direction vector, and the color feature vector includes color feature representation and body density of the pixel point on the face; and the face to be rendered is rendered based on the color feature vector to obtain a target rendering image.

[0161] In the embodiment, by the method based on the ray tracing technology, the sample point data of the face in the virtual three-dimensional space in the face image and the perspective direction vector of the light ray passing through the face can be accurately obtained. Further, by extracting the face behavior coefficient of the face image, a key parameter reflecting individual expression changes is obtained. In combination with the five-winknowledge prior features obtained from the five-winknowledge dictionary, a dynamic conditional feature vector can be spliced, the dynamic conditional feature vector combines prior knowledge of expression changes and five-winknowledge structures, and based on multiple sampling groups, the dynamic conditional feature vector, and the perspective direction vector, a color feature vector of the face can be predicted, which provides a basis for realistic rendering of the face. Finally, based on the color feature vector, the face to be rendered is rendered to obtain a target rendering image with high realism, improving the efficiency and quality of avatar image creation.

[0162] It should be noted that the functions or steps that the computer readable storage medium or the computer device can implement correspond to the related descriptions of the server side and the client side in the foregoing method embodiments, and to avoid repetition, they will not be described one by one here.

[0163] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by a computer program instructing relevant hardware, and the computer program can be stored in a non-volatile computer readable storage medium. When the computer program is executed, the processes of the above-mentioned embodiments of the methods can be included. Any reference to memory, storage, database or other medium used in each embodiment provided by the present application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration but not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0164] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the division of the above-mentioned functional units and modules is exemplified, and in actual application, the above-mentioned functions can be completed by different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the above-described functions.

[0165] The above embodiments are only used to illustrate the technical solutions of the present application, but not limit it; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that the technical solutions recorded in the foregoing embodiments can be modified, or some technical features can be replaced by equivalents; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should be included in the protection scope of the present application.

Claims

1. An image data processing method, characterized by, The method comprises: obtaining sample point data of a face in a virtual three-dimensional space and a perspective direction vector of a light ray passing through the face in a face image based on a ray tracing technology, the sample point data comprising a sampling group of each pixel point on the face, the sampling group being obtained by collecting a preset number of sample points of the light ray passing through the pixel point on the face; extracting a face behavior coefficient from the face image to obtain the face behavior coefficient; obtaining a prior feature of a facial feature of the face image from a facial feature dictionary; splicing a dynamic condition feature vector based on the face behavior coefficient and the prior feature of the facial feature to obtain the dynamic condition feature vector; predicting a color feature vector of the face based on a plurality of sampling groups, the dynamic condition feature vector and the perspective direction vector, the color feature vector comprising a color feature representation and a body density of a pixel point on the face; rendering a face to be rendered based on the color feature vector to obtain a target rendering image.

2. The image data processing method of claim 1, wherein, The method further comprises: extracting an expression coefficient and an eye movement coefficient from the face image to obtain the face behavior coefficient. The method further comprises:

3. The image data processing method of claim 1, wherein, extracting a local feature of a facial feature from the face image based on a region alignment operation; obtaining a prior feature of a facial feature from a facial feature dictionary based on the local feature of the facial feature. The method further comprises:

4. The image data processing method of claim 1, wherein, encoding the plurality of sampling groups to obtain a spatial geometric feature vector; inputting the spatial geometric vector into a lightweight multi-layer perception machine to predict a weight vector with the same dimension as the dynamic condition feature vector; performing a Hadamard product operation on the dynamic condition feature vector and the weight vector to obtain a region perception feature vector; predicting the color feature vector of the face based on the spatial geometric vector, the region perception feature vector and the perspective direction vector. The method further comprises:

5. The image data processing method of claim 4, wherein, projecting sample points in the plurality of sampling groups onto a plurality of two-dimensional planes, each two-dimensional plane comprising a plurality of projection faces with different scales; performing linear difference calculation on the projection points on each projection face on each two-dimensional plane based on a linear difference algorithm to obtain a first feature representation of each projection face in each two-dimensional plane; splicing the first feature representation of each projection face in each two-dimensional plane to obtain a second feature representation of each two-dimensional plane; splicing the second feature representation of each two-dimensional plane to obtain the spatial geometric feature vector. The number of the sampling groups is less than the number of pixel points of the face image, and the method further comprises:

6. The image data processing method of claim 1, wherein, performing up-sampling processing on the target rendering image. ​ 7. The image data processing method of claim 1, wherein, Before the method based on the ray tracing technology acquires the sampling point data of the face in the virtual three-dimensional space and the perspective direction vector of the light ray passing through the face in the face image, the method further comprises: performing image preprocessing operation on the face image, the image preprocessing operation comprising image cropping and image background setting.

8. An image data processing apparatus characterized by comprising: Comprise: The first acquisition module is used for acquiring the sampling point data of the face in the virtual three-dimensional space and the perspective direction vector of the light ray passing through the face in the face image based on the ray tracing technology, and the sampling point data comprises a sampling group of each pixel point on the face, and the sampling group is obtained by collecting a preset number of sampling points of the light ray passing through the pixel point of the face; The extraction module is used for extracting the face behavior coefficient of the face image to obtain the face behavior coefficient; The second acquisition module is used for acquiring the facial feature prior of the face image from the face feature dictionary; The splicing module is used for splicing to obtain a dynamic conditional feature vector based on the face behavior coefficient and the facial feature prior; The prediction module is used for predicting the color feature vector of the face based on a plurality of sampling groups, the dynamic conditional feature vector and the perspective direction vector, and the color feature vector comprises color feature representation and body density of the pixel point on the face; The rendering module is used for rendering the face to be rendered based on the color feature vector to obtain a target rendering image.

9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, The processor executes the computer program to realize the steps of the image data processing method according to any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, the computer program comprising instructions that, when executed by a computer, cause the computer to perform the method of any one of claims 1 to 9. The computer program is executed by the processor to realize the steps of the image data processing method according to any one of claims 1 to 7.