Three-dimensional model rendering method and device, equipment, storage medium and product

By using face key point offset calculation model and expression motion detection model in three-dimensional model rendering technology, calculating the expression coefficient of the face to be rendered and rendered, the problem of low rendering accuracy of the three-dimensional model in the existing technology is solved, and higher rendering accuracy and lower delay are achieved.

CN120070686APending Publication Date: 2025-05-30CHINA MERCHANTS BANK
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510151252.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-11
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

When the existing three-dimensional model rendering technology does not use high-precision face capture devices, high delay and inaccurate effects will occur through algorithms to render face expression information on the three-dimensional model, resulting in low accuracy of the three-dimensional model rendering technology.

Method used

By receiving face video frames, the face key point offset vector is calculated based on the preset face key point offset calculation model; the face key point offset vector is detected based on the preset face action detection model to obtain the expression action coefficient; then the face expression coefficient to be rendered is calculated based on the face key point offset vector and the expression action coefficient, and the preset three-dimensional model is rendered based on this.

Benefits of technology

It improves the accuracy of 3D model rendering technology, reduces delay, and improves the accuracy of rendering effects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070686A_ABST
    Figure CN120070686A_ABST
Patent Text Reader

Abstract

The invention discloses a three-dimensional model rendering method and device, equipment, a storage medium and a product, and relates to the technical field of computers, and the three-dimensional model rendering method comprises the steps: calculating a face key point offset vector of a face video frame based on a preset face key point offset calculation model when the face video frame is received; performing expression action detection on the face video frame according to a preset expression action detection model to obtain an expression action coefficient; according to the face key point offset vector and the expression action coefficient, calculating to obtain a to-be-rendered face expression coefficient; and rendering a preset three-dimensional model based on the to-be-rendered facial expression coefficient. The face key point offset calculation model and the expression action detection model are adopted, the face key point offset vector and the expression action coefficient are obtained, the high-accuracy to-be-rendered face expression coefficient is further obtained through calculation, and then the accuracy of the three-dimensional model rendering technology is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular to a three-dimensional model rendering method, apparatus, device, storage medium, and product. Background Art

[0002] In recent years, the three-dimensional model rendering technology based on facial capture has developed rapidly. It can present human facial expressions in digital form, bringing more realistic, natural, and interactive experiences to fields such as virtual reality, movie special effects, virtual character interaction, and social media.

[0003] Currently, the three-dimensional model rendering technology based on facial capture mainly relies on high-precision facial image capture, combined with deep learning algorithms, to record and simulate subtle human facial expressions, and finally render the three-dimensional model. However, the current three-dimensional model rendering technology is limited by hardware and algorithms. When not using high-precision facial capture devices, there will be high latency and inaccurate effects when rendering human face expression information onto the three-dimensional model through algorithms, resulting in low accuracy of the three-dimensional model rendering technology.

[0004] The above content is only used to assist in understanding the technical solution of this application, and does not represent an admission that the above content is prior art. Summary of the Invention

[0005] The main purpose of this application is to provide a three-dimensional model rendering method, apparatus, device, storage medium, and product, aiming to solve the technical problem of low accuracy of the three-dimensional model rendering technology.

[0006] To achieve the above object, this application proposes a three-dimensional model rendering method, and the method includes:

[0007] When receiving a human face video frame, based on a preset human face key point offset calculation model, calculate the human face key point offset vector of the human face video frame;

[0008] According to a preset expression action detection model, perform expression action detection on the human face video frame to obtain an expression action coefficient;

[0009] According to the human face key point offset vector and the expression action coefficient, calculate the human face expression coefficient to be rendered;

[0010] Based on the human face expression coefficient to be rendered, render a preset three-dimensional model.

[0011] In one embodiment, the step of when receiving a human face video frame, based on a preset human face key point offset calculation model, calculate the human face key point offset vector of the human face video frame includes:

[0012] Receive a face video frame, preprocess the face video frame to obtain a preprocessed face video frame;

[0013] Input the preprocessed face video frame into a preset face key point detection model to obtain face key points and corresponding face key point position vectors;

[0014] Calculate a point offset according to the face key points and the face key point position vectors;

[0015] Normalize the point offset to obtain a face key point offset vector.

[0016] In one embodiment, the step of preprocessing the face video frame to obtain a preprocessed face video frame includes:

[0017] Perform face recognition and positioning on the face video frame to obtain a first face position;

[0018] Correct the face image corresponding to the first face position according to a preset standard face position to obtain a preprocessed face video frame.

[0019] In one embodiment, the step of performing expression and action detection on the face video frame according to a preset expression and action detection model to obtain an expression and action coefficient includes:

[0020] Perform expression and action detection on the face video frame according to a preset expression and action detection model to obtain an expression and action intensity vector;

[0021] Map the expression and action intensity vector to an expression and action coefficient according to a preset mapping matrix.

[0022] In one embodiment, the step of calculating a face expression coefficient to be rendered according to the face key point offset vector and the expression and action coefficient includes:

[0023] Map the face key point offset vector to a preset standard face expression offset interval to obtain a face key point expression coefficient;

[0024] Perform weighted fusion on the expression and action coefficient and the face key point expression coefficient to calculate a face expression coefficient to be rendered.

[0025] In one embodiment, the step of rendering a preset 3D model based on the face expression coefficient to be rendered includes:

[0026] Based on the face expression coefficient to be rendered, adjust the head model parameters of the preset 3D model to obtain a 3D model with completed expression and action;

[0027] Render the 3D model through a preset rendering module, and collect the live streaming screen of the rendered 3D model;

[0028] Push the live streaming screen of the 3D model along a preset live streaming path.

[0029] In addition, to achieve the above object, the present application also proposes a 3D model rendering device, which includes:

[0030] An offset module, configured to calculate the face key point offset vector of the face video frame based on a preset face key point offset calculation model when receiving the face video frame;

[0031] A detection module, configured to perform expression and action detection on the face video frame according to a preset expression and action detection model to obtain an expression and action coefficient;

[0032] A calculation module, configured to calculate a face expression coefficient to be rendered according to the face key point offset vector and the expression and action coefficient;

[0033] A rendering module, configured to render a preset 3D model based on the face expression coefficient to be rendered.

[0034] In addition, to achieve the above object, the present application also proposes a 3D model rendering device, which includes: a memory, a processor, and a computer program stored on the memory and executable on the processor, and the computer program is configured to implement the steps of the 3D model rendering method as described above.

[0035] In addition, to achieve the above object, the present application also proposes a storage medium, which is a computer-readable storage medium, and a computer program is stored on the storage medium, and when the computer program is executed by a processor, it implements the steps of the 3D model rendering method as described above.

[0036] In addition, to achieve the above object, the present application also provides a computer program product, which includes a computer program, and when the computer program is executed by a processor, it implements the steps of the 3D model rendering method as described above.

[0037] One or more technical solutions proposed by the present application have at least the following technical effects:

[0038] In related technologies, current 3D model rendering technologies are restricted by hardware and algorithms. When not using high-precision facial capture devices, there will be high latency and inaccurate effects when rendering facial expression information onto a 3D model through algorithms, resulting in a low accuracy rate of the 3D model rendering technology. In contrast, in this application, when a face video frame is received, based on a preset facial key point offset calculation model, the facial key point offset vector of the face video frame is calculated; according to a preset expression action detection model, expression action detection is performed on the face video frame to obtain an expression action coefficient; according to the facial key point offset vector and the expression action coefficient, a to-be-rendered facial expression coefficient is calculated; based on the to-be-rendered facial expression coefficient, a preset 3D model is rendered. It can be understood that this application uses a facial key point offset calculation model to calculate the facial key point offset vector of a face video frame; detects the expression action coefficient of the face video frame according to the expression action detection model; calculates a to-be-rendered facial expression coefficient with high accuracy according to the facial key point offset vector and the expression action coefficient, thereby improving the accuracy rate of the 3D model rendering technology. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] The drawings herein are incorporated into and constitute a part of this specification, showing embodiments consistent with this application, and are used together with the specification to explain the principles of this application.

[0040] To more clearly illustrate the technical solutions in the embodiments of this application or in the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, other drawings can also be obtained based on these drawings without creative efforts.

[0041] Figure 1 It is a schematic flowchart provided for Embodiment 1 of the 3D model rendering method of this application;

[0042] Figure 2 It is a schematic flowchart provided for Embodiment 2 of the 3D model rendering method of this application;

[0043] Figure 3 It is a schematic flowchart provided for Embodiment 3 of the 3D model rendering method of this application;

[0044] Figure 4 It is a schematic module structure diagram of the 3D model rendering device for the embodiments of this application;

[0045] Figure 5 It is a schematic device structure diagram of the hardware operating environment involved in the 3D model rendering method for the embodiments of this application.

[0046] The realization of the purpose, functional features, and advantages of this application will be further described with reference to the embodiments and the drawings. Detailed implementation manners

[0047] It should be understood that the specific embodiments described herein are only used to explain the technical solutions of the present application and are not used to limit the present application.

[0048] To better understand the technical solutions of the present application, the following will be described in detail in conjunction with the accompanying drawings of the specification and specific implementation manners.

[0049] The main solution of the embodiments of the present application is as follows:

[0050] When a face video frame is received, based on a preset face key point offset calculation model, calculate the face key point offset vector of the face video frame;

[0051] According to a preset expression action detection model, perform expression action detection on the face video frame to obtain an expression action coefficient;

[0052] According to the face key point offset vector and the expression action coefficient, calculate the face expression coefficient to be rendered;

[0053] Based on the face expression coefficient to be rendered, render a preset 3D model.

[0054] In this embodiment, the present application takes a 3D model rendering device as the execution subject. For the convenience of description, it is hereinafter simply referred to as the "device" for specific description.

[0055] Due to the limitations of current 3D model rendering technology in the prior art by hardware and algorithms, when a high-precision facial capture device is not used, there will be high latency and inaccurate effects when rendering face expression information onto a 3D model through algorithms, resulting in low accuracy of 3D model rendering technology.

[0056] The present application provides a solution, enabling the present application to calculate the face key point offset vector of the face video frame when a face video frame is received, based on a preset face key point offset calculation model; perform expression action detection on the face video frame according to a preset expression action detection model to obtain an expression action coefficient; calculate the face expression coefficient to be rendered with high accuracy according to the face key point offset vector and the expression action coefficient, and achieve an improvement in the accuracy of 3D model rendering technology.

[0057] Based on this, the embodiments of the present application provide a three-dimensional model rendering method. Referring to Figure 1 , Figure 1 is a schematic flowchart of the first embodiment of the three-dimensional model rendering method of the present application.

[0058] In this embodiment, the three-dimensional model rendering method includes steps S10 to S40:

[0059] Step S10, when a face video frame is received, based on a preset face key point offset calculation model, calculate the face key point offset vector of the face video frame;

[0060] It should be noted that the face video frame is a single image intercepted from a video stream, which contains the image information of a face. In facial expression driving technology, these frames are the basic data for capturing and analyzing facial expressions. The face key point offset calculation model is an algorithm model used to identify and calculate the offset of the key points of a face (such as the positions of eyes, nose, mouth, etc.) relative to the standard position from the face video frame. These key points are the basis for facial expression analysis. The face key point offset vector represents the offset of the face key points relative to their positions in the standard or neutral expression state. This vector contains the displacement information of the key points in two-dimensional or three-dimensional space and is used to capture the changes in facial expressions.

[0061] It can be understood that, first of all, the device needs to receive a face video frame. This video frame can be captured in real time or extracted from a pre-recorded video. The video frame contains a series of pixel data, which represent the facial image of the person in the video. Inside the device, one or more face key point offset calculation models are preset. These models are trained based on a large amount of face data and can identify and understand the key feature points of a face, such as the corners of the eyes, the corners of the mouth, the tip of the nose, etc. These models can predict the positions of the key points according to the input face video frame.

[0062] After receiving the face video frame, the device will use the preset face key point offset calculation model to calculate the face key point offset vector of the face video frame. This step involves the following sub-steps:

[0063] Feature extraction: The device first performs feature extraction on the video frame to identify the key feature points of the face. This can be achieved through deep learning algorithms, such as using a convolutional neural network (CNN) to identify and locate the face feature points.

[0064] Key point positioning: Based on the feature extraction, the device further accurately locates these key points. This usually involves analyzing the pixels around the key points to determine the exact positions of the key points.

[0065] Offset Vector Calculation: Once the key points are located, the device will calculate the offset vectors of each key point relative to the reference positions in the preset model. These offset vectors describe the differences between the actual positions of the key points in the video frame and the expected positions in the model.

[0066] Data Correction: To improve accuracy, the device also needs to correct the calculated offset vectors. This can be achieved by comparing the changes in key points between different frames to eliminate errors caused by factors such as lighting changes and expression changes.

[0067] Exemplarily, for the generalization problem of different input video faces, based on the mediapipe AI model, it is possible to infer the position vectors of face key points through video frames, which can adapt to most faces. Therefore, it has generalization ability and can be generalized to different character images. According to the mediapipe face key point detection model, the 468-dimensional key point position vectors in the input video frame can be detected.

[0068] In a feasible implementation manner, step S10 may include:

[0069] Receive a face video frame, preprocess the face video frame to obtain a preprocessed face video frame;

[0070] Input the preprocessed face video frame into a preset face key point detection model to obtain face key points and corresponding face key point position vectors;

[0071] Calculate the point offset according to the face key points and the face key point position vectors;

[0072] Normalize the point offset to obtain a face key point offset vector.

[0073] It should be noted that the face video frame is a single frame image captured from a video, which contains the facial information of a person and is a static image in the video stream for subsequent facial expression analysis and processing. The preprocessing in facial expression analysis refers to a series of operations performed on the original video frame to improve the accuracy and efficiency of subsequent processing. This may include steps such as denoising, contrast enhancement, color correction, and face detection. The face key-point detection model is an algorithm model used to identify and locate the key feature points on the face, such as the corners of the eyes, the corners of the mouth, and the tip of the nose, from the preprocessed face video frame. The face key points are the points on the face with significant features, which are crucial for the recognition and simulation of facial expressions. The face key-point position vector describes the specific positions of the face key points in the video frame, usually represented in coordinate form. The point offset is the amount of change in the position of the face key point between different frames or compared with a preset model. The normalization process is a mathematical processing method used to scale the data to a specific range for easier comparison and processing. In facial expression analysis, the normalization process can ensure the consistency and comparability of different face key-point offset values.

[0074] It can be understood that the device starts to receive a face video frame, which is a static image containing a person's face extracted from the video. After receiving the face video frame, the device preprocesses it to improve the accuracy and efficiency of subsequent processing. The preprocessing may include operations such as denoising, contrast enhancement, color correction, and face detection to ensure that the face area is correctly identified and isolated. After the preprocessing is completed, the device inputs the preprocessed face video frame into a preset face key-point detection model. This model, based on deep learning technology, can accurately identify and locate the key feature points on the face, such as the corners of the eyes, the corners of the mouth, and the tip of the nose, and obtains the face key points and the corresponding face key-point position vectors. According to the detected face key points and their position vectors, the device calculates the offset of each key point. This offset describes the change of the key point between the current frame and the previous frame or the preset model. To ensure the consistency and comparability of different key-point offsets, the device normalizes the calculated point offset. This step scales the offset to a specific range, usually [-1, 1] or [0, 1], for subsequent facial expression simulation.

[0075] Exemplarily, first, the current video frame image is read from the camera or video through the opencv vision library, and the image is preprocessed to change the three channels from BGR to RGB.

[0076] After the preprocessing of the video frame is completed, it is input into the mediapipe face key-point detection model FaceMesh to obtain a [468, 3]-dimensional vector of face key points, for example:

[0077] Landmark Vector: L = (L1, L2, L3, ..., Lm).

[0078] In a feasible implementation, the step of preprocessing the face video frame to obtain a preprocessed face video frame includes:

[0079] Perform face recognition and localization on the face video frame to obtain the first face position;

[0080] According to the preset standard face position, correct the face image corresponding to the first face position to obtain a preprocessed face video frame.

[0081] It should be noted that the face video frame refers to a single-frame image containing a person's face intercepted from a video stream and is used for facial expression analysis and processing. The face recognition and localization is a process that detects and locates the face position in the image through a specific algorithm (such as a deep learning-based face recognition algorithm). The first face position is the initial position of the face recognized in the video frame, and this position may be different from the standard face position due to factors such as shooting angle and distance. The standard face position is a preset ideal face position, usually used in facial expression analysis and processing to ensure the consistency and comparability of different face images. The face image correction refers to the process of adjusting the recognized face to the standard face position, which involves operations such as rotation, scaling, and cropping. The preprocessed face video frame: refers to the face image that meets the requirements of the standard face position after face correction processing, and prepares for subsequent facial key point detection and expression analysis.

[0082] It can be understood that due to factors such as shooting angle, light conditions, and head pose, different face images may vary greatly. The device starts to receive a face video frame, which is a static image containing a person's face extracted from the video. The device performs face recognition and localization on the received face video frame, uses a deep learning algorithm to detect the face in the image, and determines its position in the image to obtain the first face position.

[0083] According to the preset standard face position, the device corrects the face image corresponding to the first face position. This may include the following operations:

[0084] Rotation: If the face is not facing forward, the system will rotate the image to align the facial features (such as eyes, nose, mouth) with the standard face position.

[0085] Scaling: If the size of the face is inconsistent with the standard face position, the system will adjust the image size to match the standard face position.

[0086] Cropping: If the background or extra parts around the face need to be removed, the device will crop the image to retain only the face part.

[0087] After the above correction operations, the device obtains a preprocessed face video frame. This video frame now meets the requirements of the standard face position and can be used for subsequent facial key point detection and expression analysis.

[0088] Exemplarily, through head correction, the face image can be standardized to be aligned in the horizontal and vertical directions, which can reduce the influence of these factors on subsequent image processing steps. To exclude interference when calculating the offset, the present application corrects the head, and the correction method is shown in the following formula:

[0089]

[0090] where x′, y′ represent the position of the face in the picture, x, y represent the position of the standard face, t x , t y represents the size of translation. At the same time, since the face box has determined the translation position, actually here, t x = t y = 0. Therefore, the rotation and translation matrix is represented by 4 vectors, so the positions of 5 points can find a unique solution to obtain the rotation matrix, and this rotation matrix is applied to subsequent key points.

[0091] Step S20: According to a preset expression action detection model, perform expression action detection on the face video frame to obtain expression action coefficients;

[0092] It should be noted that the expression action detection model is an algorithm model for identifying and quantifying the expression changes of faces in videos. It can detect specific expression actions based on the changes of facial key points and the movements of facial muscles. The expression action coefficients are a series of numerical values output by the expression action detection model, and these numerical values represent the intensity and characteristics of specific expression actions. These coefficients can be used to quantify the nuances of expressions, such as the degree of smiling or the depth of frowning.

[0093] It is understandable that the device starts to receive a face video frame, which is a static image containing a person's face extracted from the video. Before performing expression action detection, the device may need to preprocess the received face video frame to improve the accuracy and efficiency of subsequent processing. The preprocessing steps may include operations such as denoising, contrast enhancement, color correction, face detection, etc., as well as face alignment to ensure that the face area is correctly identified and isolated. The device performs expression action detection on the preprocessed face video frame according to a preset expression action detection model. This model may be trained based on machine learning or deep learning techniques and can recognize various facial expressions, such as happiness, sadness, surprise, etc., and quantify the characteristics of these expressions. Through expression action detection, the device obtains a series of expression action coefficients. These coefficients are parameters used to describe facial expression changes in the 3D model.

[0094] Exemplarily, the AU unit is a general standard for objectively describing facial expressions and is derived from the Facial Action Coding System (FACS). FACS is a system for classifying human facial movements according to facial expressions. FACS encodes the movements of each facial muscle according to the subtle changes in facial expressions. Using FACS, almost any anatomically possible facial expression can be encoded and decomposed into specific action units (AU units) that produce the expression.

[0095] The AU unit detection model used in this application is derived from the openface pre-trained model. At the same time, in order to achieve the effect of real-time inference in this application, openface is optimized and encapsulated to support python calls and image frame processing. The subset of AU units that openface can recognize is [1, 2, 4, 5, 7, 9, 10, 12, 14, 15, 17, 20, 23, 25, 26, 28, 45].

[0096] Among them, the description methods of AU units are:

[0097] Existence: Existence is 0, non-existence is 1;

[0098] Intensity: The intensity ranges from 1 to 5, and there is also an outlier value of 0, indicating non-existence.

[0099] In a feasible implementation manner, step S20 may include:

[0100] Performing expression action detection on the face video frame according to a preset expression action detection model to obtain an expression action intensity vector;

[0101] Mapping the expression action intensity vector to expression action coefficients according to a preset mapping matrix.

[0102] It should be noted that the facial expression action detection model is a preset algorithm model used to identify and quantify facial expression actions in a face video. This model can understand the characteristics of different facial expressions and convert them into quantifiable parameters. The facial expression intensity vector is a series of values obtained by analyzing the facial expression action detection model. These values represent the specific characteristics and intensities of the facial expressions of the people in the video and are used for subsequent facial expression simulation and 3D model driving. The mapping matrix is a preset mathematical model used to convert the facial expression intensity vector into facial expression action coefficients that can drive the 3D model. The mapping matrix contains the corresponding relationship from facial expression intensity to 3D model parameters.

[0103] It can be understood that the device performs facial expression action detection on the preprocessed face video frames according to the preset facial expression action detection model. This model is trained based on machine learning or deep learning techniques and can identify various facial expressions such as happiness, sadness, surprise, etc., and quantify the characteristics of these expressions. Through facial expression action detection, the device obtains a series of facial expression intensity vectors. The device maps the obtained facial expression intensity vectors into facial expression action coefficients according to the preset mapping matrix. The mapping matrix contains the corresponding relationship from facial expression intensity to 3D model parameters, ensuring the accurate conversion of facial expression actions.

[0104] Exemplarily, after the image is preprocessed, AU unit detection is performed. The AU unit detection model structure is a PCA dimensionality reduction model, mainly using an SVM classifier. Here, an openface pre-trained model is used, and the training data comes from databases such as BP4D, DISFA, FER2011, SEMMAINE, UNBC, etc. Using the detected AU intensity vector, it is mapped into the facial expression BS coefficient of the target 3D model. The BS calculation formula is as follows:

[0105] B au = wA + d

[0106] where A is the AU intensity vector. Blend shape vector Blend Vector: B au =(B1, B2, B3,..., Bm), Bi represents the i-th BS, and the intensity is in the closed interval [0, 1], where m is 35. The mapping matrix w is preset according to the correlation between the AU corresponding expression and the BS expression. d is the bias, which is also preset.

[0107] Step S30, calculate the face expression coefficient to be rendered according to the face key point offset vector and the facial expression action coefficient;

[0108] It should be noted that the face expression coefficients to be rendered are a set of parameters calculated based on the face key point offset vectors and expression action coefficients, and are used to guide the expression rendering of the 3D model. These coefficients ensure that the virtual face expression matches the actual expression in the video.

[0109] In a feasible implementation manner, step S30 may include:

[0110] Map the face key point offset vector to a preset standard face expression offset interval to obtain face key point expression coefficients;

[0111] Perform weighted fusion of the expression action coefficients and the face key point expression coefficients to calculate the face expression coefficients to be rendered.

[0112] Exemplarily, after obtaining a 468-dimensional face key point vector according to the mediapipe face key point detection model FaceMesh, the offset of the current face expression can be calculated, and the offset is mapped to the offset of the corresponding average face expression, so as to obtain the corresponding expression BS coefficient. During inference, the expression is processed separately by face part, divided into the eyes, eyebrows part and the mouth, cheeks part.

[0113] Key point motion mapping BS coefficient for eyes and eyebrows part:

[0114] After obtaining a 468-dimensional face key point vector, only take the key points related to the facial expression and following the facial movement. Let the key point vector of the eye part be L eye , where the key point vector numbers related to the left eye are [263, 362, 387, 386, 385, 373, 374, 380, 253, 450, 53, 52, 65, 9, 359, 342]; the key point vector numbers related to the right eye are [33, 133, 160, 159, 158, 144, 145, 153, 23, 230, 283, 282, 295, 6, 130, 113], and the function for processing the key points is expressed as follows:

[0115]

[0116] Among them, the blend shape represents the BS of the i-th expression related to the eyes and eyebrows part, and the intensity is in the closed interval [0, 1], and the interval [a i , b i is the offset interval of the average face key points in the corresponding expression. When f(xL eye ) is equal to b, it can be understood that the corresponding BS coefficient of the expression of this face is 1, that is, the expression of this face is exactly at the maximum intensity of the corresponding expression. The interval values [a i , b iAlthough it has been preset according to the average face, but a i ,b i Manual or automatic correction can be performed to improve the accuracy.

[0117] If the detected facial expression offset exceeds the ranges of ai and bi, the device automatically expands the ranges of bi and ai for automatic correction; during the initialization phase, the ranges of ai and bi are automatically corrected through videos or image frames containing extreme expressions, where the extreme expressions containing facial expressions include but are not limited to neutral, fully open chin, closed smile, downturned lips, pouting lips, raised eyebrows, downturned eyebrows, and closed eyes.

[0118] By manually or automatically correcting the ranges of ai and bi, the model can better fit the expression intensity.

[0119] Key point motion mapping BS coefficient for the mouth and cheek parts:

[0120] Let the key point vector of the eye part be L mou , and the key point vector numbers of the mouth part are [12, 13, 14, 291, 61, 1, 10, 152, 202, 422, 57, 287, 0, 17, 18, 164, 89, 319]. The function for processing the key points is expressed as follows:

[0121]

[0122]

[0123] Among them, the blend shape represents the BS of the i-th expression related to the mouth and cheek parts, and the intensity is also in the closed interval [0, 1]. The interval [a i ,b i is the motion interval of the key points of the average face for the corresponding expression. Similarly, the interval [a i ,b i can be adjusted.

[0124] Fuse the key point BS coefficient results B lneye and B bmouth with the AU detection BS coefficient results to obtain the fused BS coefficient result B merge :

[0125] B merge =αB lneye +αB bmouth +βB au

[0126] Among them,

[0127] α + β = 1

[0128] 0 < α, β < 1

[0129] Finally, send B merge to the rendering module via websocket for rendering.

[0130] Step S40: Render the preset 3D model based on the face expression coefficients to be rendered.

[0131] It should be noted that the 3D model is a digital 3D face model, which can be deformed and rendered according to the input expression coefficients to simulate the expressions of real human faces. This model usually includes vertices, textures, and possibly bone structures for accurately simulating facial expressions. The rendering refers to the process of converting the 3D model into a 2D image that matches the face expression in the video according to the face expression coefficients to be rendered.

[0132] Exemplarily, infer 56-dimensional BS coefficients in the 3D digital human general head model parameters based on the input video or camera. After sending the BS coefficients to the rendering module via websocket, the rendering module receives the parameters, and the rendering module can drive the head of the 3D digital human model by starting the MetaHuman blueprint of UE4. The specific operation process is as Figure 2 shown:

[0133] The WebRTC signaling service is responsible for managing the signaling process between the client and the server, including the establishment, maintenance, and termination of sessions. The signaling service receives BS parameters from the client, and these parameters are key data for facial expressions.

[0134] UE4 MetaHuman rendering is a game engine based on Unreal Engine 4 (UE4), which is specifically used for rendering high-quality 3D digital human models. The MetaHuman module allows the creation and rendering of highly realistic human characters.

[0135] Streaming is to push the video stream of the rendered 3D digital human model to the network for real-time playback in a browser or video player.

[0136] The browser or video player is where the end user receives and plays the real-time rendered 3D digital human video stream through these client devices.

[0137] In a feasible implementation, step S40 may include:

[0138] Based on the face expression coefficients to be rendered, adjust the head model parameters of the preset 3D model to obtain a 3D model with completed expression actions;

[0139] Render the 3D model through a preset rendering module, and collect the streaming screen of the rendered 3D model.

[0140] Push the streaming screen of the 3D model along a preset streaming path.

[0141] Exemplarily, first, start the WebRTC signaling service for handling the signaling between the client and the server.

[0142] The WebRTC signaling service receives the BS parameters from the client and transmits the received BS parameters to the UE4MetaHuman rendering module.

[0143] The UE4 MetaHuman rendering module adjusts the facial expressions of the 3D digital human model in real time according to the received BS parameters.

[0144] The rendering module processes the texture, lighting, and animation of the 3D model to ensure the authenticity of the expressions. The rendered video stream is sent to the network through the streaming service. The browser or video player acts as the client and receives the streamed video from the server. The client decodes and plays the video stream to display the real-time facial expressions of the 3D digital human.

[0145] The user can capture their own facial expressions through the camera, and these expressions are transmitted to the server in real time through the WebRTC signaling service. The server processes the expression data and updates the expressions of the 3D digital human model in real time through the UE4 MetaHuman rendering module.

[0146] When the user finishes the interaction or closes the application, the WebRTC signaling service is responsible for terminating the session and releasing resources.

[0147] This embodiment provides a 3D model rendering method, which uses a face key point offset calculation model to calculate the face key point offset vector of the face video frame; detects the expression action coefficient of the face video frame according to the expression action detection model; calculates the to-be-rendered face expression coefficient with high accuracy according to the face key point offset vector and the expression action coefficient, thereby improving the accuracy of the 3D model rendering technology.

[0148] Exemplarily, to help understand the implementation process of the 3D model rendering method obtained by combining this embodiment with the above Embodiment 1, please refer to Figure 3 , Figure 3 A brief flow schematic diagram of a 3D model rendering method is provided. Specifically:

[0149] The device receives video input, and the video input can be a real-time video stream or a pre-recorded video.

[0150] Preprocess the video, including denoising, color correction, contrast enhancement, etc., to improve the accuracy of subsequent processing.

[0151] Detect faces in the preprocessed video frames and identify the face regions.

[0152] Perform key point detection on the detected faces to identify key feature points such as the corners of the eyes, the corners of the mouth, the tip of the nose, etc.

[0153] According to the results of the key point detection, calculate the offset of the key points relative to the standard face model.

[0154] Map the offset to Blend Shape (BS) coefficients, which are used to control the facial expressions of the 3D model.

[0155] Perform weighted processing on the BS coefficients to simulate more natural or exaggerated facial expressions.

[0156] Use the weighted BS coefficients to drive the 3D model for video rendering.

[0157] Utilize the AU (Action Unit) detection algorithm to identify specific facial action units in the video, such as eyebrow raising, mouth corner raising, etc.

[0158] Map the detected AUs to the corresponding BS coefficients to more precisely control the expressions of the 3D model.

[0159] It should be noted that the above examples are only for understanding this application and do not constitute a limitation on the 3D model rendering method of this application. Based on this technical concept, more forms of simple transformations are within the protection scope of this application.

[0160] This application also provides a 3D model rendering device. Please refer to Figure 4 , and the 3D model rendering device includes:

[0161] Offset module 10, configured to calculate the face key point offset vector of the face video frame based on a preset face key point offset calculation model when receiving the face video frame;

[0162] Detection module 20, configured to perform expression action detection on the face video frame according to a preset expression action detection model to obtain expression action coefficients;

[0163] Calculation module 30, configured to calculate the face expression coefficients to be rendered according to the face key point offset vector and the expression action coefficients;

[0164] Rendering module 40, configured to render a preset 3D model based on the face expression coefficients to be rendered.

[0165] And / or, the offset module 10 includes:

[0166] A first preprocessing module, configured to receive a face video frame, preprocess the face video frame, and obtain a preprocessed face video frame;

[0167] A first detection module, configured to input the preprocessed face video frame into a preset face key point detection model, and obtain face key points and corresponding face key point position vectors;

[0168] A first calculation module, configured to calculate a point offset according to the face key points and the face key point position vectors;

[0169] A first normalization module, configured to normalize the point offset to obtain a face key point offset vector.

[0170] And / or, the first preprocessing module includes:

[0171] A first positioning module, configured to perform face recognition and positioning on the face video frame to obtain a first face position;

[0172] A first correction module, configured to correct the face image corresponding to the first face position according to a preset standard face position to obtain a preprocessed face video frame.

[0173] And / or, the detection module 20 includes:

[0174] A second detection module, configured to perform expression and action detection on the face video frame according to a preset expression and action detection model to obtain an expression and action intensity vector;

[0175] A first mapping module, configured to map the expression and action intensity vector to an expression and action coefficient according to a preset mapping matrix.

[0176] And / or, the calculation module 30 includes:

[0177] A second mapping module, configured to map the face key point offset vector to a preset standard face expression offset interval to obtain a face key point expression coefficient;

[0178] A second calculation module, configured to perform weighted fusion on the expression and action coefficient and the face key point expression coefficient to calculate a face expression coefficient to be rendered.

[0179] And / or, the rendering module 40 includes:

[0180] A first adjustment module, configured to adjust the head model parameters of a preset three-dimensional model based on the face expression coefficient to be rendered to obtain a three-dimensional model with completed expression and action;

[0181] A first rendering module, configured to render the 3D model through a preset rendering module and collect the streaming screen of the rendered 3D model.

[0182] A first streaming module, configured to stream the streaming screen of the 3D model according to a preset streaming path.

[0183] The 3D model rendering device provided by the present application adopts the 3D model rendering method in the above embodiment, and can solve the technical problem of low accuracy in the 3D model rendering technology. Compared with the prior art, the beneficial effects of the 3D model rendering device provided by the present application are the same as those of the 3D model rendering method provided by the above embodiment, and other technical features in the 3D model rendering device are the same as the features disclosed in the method of the above embodiment, which will not be elaborated herein.

[0184] The present application provides a 3D model rendering device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the 3D model rendering method in the first embodiment above.

[0185] Reference is made below to Figure 5 , which shows a schematic structural diagram of a 3D model rendering device suitable for implementing the embodiments of the present application. The 3D model rendering device in the embodiments of the present application may include, but is not limited to, mobile terminals such as mobile phones, tablet computers, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PMPs (Portable Media Players), in-vehicle terminals (such as in-vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 5 The 3D model rendering device shown is only an example and should not impose any limitations on the functions and usage scope of the embodiments of the present application.

[0186] As shown in Figure 5As shown in the figure, the 3D model rendering device may include a processing device 1001 (such as a central processing unit, a graphics processing unit, etc.), which may perform various appropriate actions and processes according to a program stored in a read-only memory (ROM: Read Only Memory) 1002 or a program loaded from a storage device 1003 into a random access memory (RAM: Random Access Memory) 1004. In the RAM 1004, various programs and data required for the operation of the 3D model rendering device are also stored. The processing device 1001, the ROM 1002, and the RAM 1004 are connected to each other through a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Generally, the following systems may be connected to the I / O interface 1006: an input device 1007 including, for example, a touch screen, a touchpad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; an output device 1008 including, for example, a liquid crystal display (LCD: Liquid Crystal Display), a speaker, a vibrator, etc.; a storage device 1003 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1009. The communication device 1009 may allow the 3D model rendering device to communicate with other devices wirelessly or wiredly to exchange data. Although the figure shows a 3D model rendering device having various systems, it should be understood that it is not required to implement or have all the shown systems. More or fewer systems may be alternatively implemented or had.

[0187] In particular, according to the embodiments disclosed in the present application, the processes described above with reference to the flowcharts may be implemented as computer software programs. For example, the embodiments disclosed in the present application include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program contains program codes for executing the methods shown in the flowcharts. In such an embodiment, the computer program may be downloaded and installed from a network through the communication device, or installed from the storage device 1003, or installed from the ROM 1002. When the computer program is executed by the processing device 1001, the above-mentioned functions defined in the methods of the embodiments disclosed in the present application are executed.

[0188] The 3D model rendering device provided by the present application adopts the 3D model rendering method in the above-mentioned embodiments, and can solve the technical problem of low accuracy in the 3D model rendering technology. Compared with the prior art, the beneficial effects of the 3D model rendering device provided by the present application are the same as those of the 3D model rendering method provided by the above-mentioned embodiments, and the other technical features in the 3D model rendering device are the same as the features disclosed in the method of the previous embodiment, and will not be elaborated here.

[0189] It should be understood that each part disclosed in this application can be implemented by hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in a suitable manner in any one or more embodiments or examples.

[0190] As described above, the above are only specific embodiments of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art within the technical scope disclosed in this application can easily think of changes or substitutions, which should all be covered within the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.

[0191] This application provides a computer-readable storage medium having computer-readable program instructions (i.e., computer programs) stored thereon, and the computer-readable program instructions are used to execute the three-dimensional model rendering method in the above embodiments.

[0192] The computer-readable storage medium provided by this application can be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or any combination of the above. More specific examples of computer-readable storage media can include, but are not limited to: electrical connections with one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM) or flash memory, optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In this embodiment, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or combined with an instruction execution system, device, or device. The program code contained on the computer-readable storage medium can be transmitted by any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.

[0193] The above computer-readable storage medium can be included in the three-dimensional model rendering device; it can also exist separately without being assembled into the three-dimensional model rendering device.

[0194] The above computer-readable storage medium carries one or more programs. When the one or more programs are executed by a 3D model rendering device, the 3D model rendering device is caused to: when receiving a face video frame, calculate a face key point offset vector of the face video frame based on a preset face key point offset calculation model; perform expression action detection on the face video frame according to a preset expression action detection model to obtain an expression action coefficient; calculate a to-be-rendered face expression coefficient according to the face key point offset vector and the expression action coefficient; and render a preset 3D model based on the to-be-rendered face expression coefficient.

[0195] Computer program code for performing the operations of the present application may be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., by connecting through the Internet using an Internet service provider).

[0196] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combinations of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system for performing the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.

[0197] The modules involved in the embodiments of the present application can be implemented in software or in hardware. In some cases, the name of the module does not constitute a limitation on the unit itself.

[0198] The readable storage medium provided by the present application is a computer-readable storage medium, and the computer-readable storage medium stores computer-readable program instructions (i.e., computer programs) for executing the above-mentioned three-dimensional model rendering method, which can solve the technical problem of low accuracy in three-dimensional model rendering technology. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided by the present application are the same as those of the three-dimensional model rendering method provided by the above embodiments, and will not be elaborated here.

[0199] The present application also provides a computer program product, including a computer program, and when the computer program is executed by a processor, the steps of the three-dimensional model rendering method as described above are implemented.

[0200] The computer program product provided by the present application can solve the technical problem of low accuracy in three-dimensional model rendering technology. Compared with the prior art, the beneficial effects of the computer program product provided by the present application are the same as those of the three-dimensional model rendering method provided by the above embodiments, and will not be elaborated here.

[0201] The above are only some embodiments of the present application, and thus do not limit the patent scope of the present application. Any equivalent structural transformation made under the technical concept of the present application by using the content of the specification and drawings of the present application, or direct / indirect application in other related technical fields, is included in the patent protection scope of the present application.

Claims

1. A three-dimensional model rendering method, characterized in that: The method includes: When a face video frame is received, a face key point offset vector of the face video frame is calculated based on a preset face key point offset calculation model; According to a preset expression and action detection model, expression and action detection is performed on the face video frame to obtain an expression and action coefficient; Calculating the facial expression coefficient to be rendered according to the facial key point offset vector and the facial expression action coefficient; Based on the facial expression coefficient to be rendered, a preset three-dimensional model is rendered.

2. The method according to claim 1, characterized in that When a face video frame is received, the step of calculating the face key point offset vector of the face video frame based on a preset face key point offset calculation model comprises: Receiving a face video frame, and preprocessing the face video frame to obtain a preprocessed face video frame; Input the preprocessed face video frame into a preset face key point detection model to obtain face key points and corresponding face key point position vectors; Calculating a point offset according to the standard face key point position vector of the face key point and the face key point position vector; The point offset is normalized to obtain a facial key point offset vector.

3. The method according to claim 2, characterized in that The step of preprocessing the face video frame to obtain the preprocessed face video frame comprises: Performing face recognition and positioning on the face video frame to obtain a first face position; According to a preset standard face position, the face image corresponding to the first face position is corrected to obtain a preprocessed face video frame.

4. The method according to claim 1, characterized in that The step of performing expression and action detection on the face video frame according to a preset expression and action detection model to obtain an expression and action coefficient comprises: According to a preset expression and action detection model, expression and action detection is performed on the face video frame to obtain an expression and action intensity vector; According to a preset mapping matrix, the expression action intensity vector is mapped into an expression action coefficient.

5. The method according to claim 1, characterized in that The step of calculating the facial expression coefficient to be rendered according to the facial key point offset vector and the facial expression action coefficient comprises: Mapping the facial key point offset vector to a preset standard facial expression offset interval to obtain a facial key point expression coefficient; The expression action coefficient and the facial key point expression coefficient are weightedly fused to calculate the facial expression coefficient to be rendered.

6. The method according to claim 1, characterized in that The step of rendering a preset three-dimensional model based on the facial expression coefficient to be rendered includes: Based on the facial expression coefficient to be rendered, adjusting the head model parameters of the preset three-dimensional model to obtain a three-dimensional model that completes the facial expression action; Rendering the three-dimensional model through a preset rendering module, and collecting a streaming image of the rendered three-dimensional model; According to the preset streaming path, the three-dimensional model streaming screen is streamed.

7. A three-dimensional model rendering device, characterized in that: The device comprises: An offset module is used to calculate the face key point offset vector of the face video frame based on a preset face key point offset calculation model when a face video frame is received; A detection module, used to perform expression and action detection on the face video frame according to a preset expression and action detection model to obtain an expression and action coefficient; A calculation module, used for calculating the facial expression coefficient to be rendered according to the facial key point offset vector and the facial expression action coefficient; The rendering module is used to render a preset three-dimensional model based on the facial expression coefficient to be rendered.

8. A three-dimensional model rendering device, characterized in that: The device comprises: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the three-dimensional model rendering method according to any one of claims 1 to 6.

9. A storage medium, characterized in that: The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, the steps of the three-dimensional model rendering method according to any one of claims 1 to 6 are implemented.

10. A computer program product, characterized in that The computer program product comprises a computer program, and when the computer program is executed by a processor, the steps of the three-dimensional model rendering method according to any one of claims 1 to 6 are implemented.