A three-dimensional image pronunciation head pose simulation method

By collecting and training two-dimensional head posture data, establishing a neural network model to form a three-dimensional pronunciation neural network, the problem of inconsistent head posture and pronunciation during the three-dimensional image pronunciation process is solved, and a natural three-dimensional pronunciation effect is achieved.

CN116129487BActive Publication Date: 2025-08-01JINDONG CULTURE TECHNOLOGY CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202211479715.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-24
Publication Date
2025-08-01
Estimated Expiration
2042-11-24

AI Technical Summary

Technical Problem

During the pronunciation of existing three-dimensional images, the head posture and pronunciation are lacking in linkage, resulting in unnatural and rigid pronunciation.

Method used

The two-dimensional head posture elements in the plane video library are collected, the neural network model is established, and the training data set is trained. The transfer learning algorithm is used to form a three-dimensional pronunciation neural network model, and the head posture of the three-dimensional virtual image is controlled by combining the three-dimensional vocal element set.

Benefits of technology

The linkage between head posture and pronunciation during the three-dimensional image pronunciation process is realized, and the naturalness of pronunciation is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116129487B_ABST
    Figure CN116129487B_ABST
Patent Text Reader

Abstract

The present invention provides a three-dimensional image pronunciation head pose simulation method, belonging to the technical field of three-dimensional image pronunciation. The method includes: collecting two-dimensional head pose elements of a person speaking in a plane video library as a training data set, and establishing a correspondence relationship between the two-dimensional head pose elements; establishing a neural network model and training it; performing three-dimensional recognition on the pronunciation processes of multiple testers to obtain three-dimensional head pose elements; using a transfer learning algorithm to update and optimize the neural network model with the three-dimensional head pose elements to form a three-dimensional pronunciation neural network model; taking a string that requires three-dimensional image pronunciation as text data, extracting phonemes from the text data to obtain a pronunciation phoneme set; using the three-dimensional pronunciation neural network model to calculate the obtained pronunciation phoneme set to obtain a three-dimensional vocalization element set; establishing a three-dimensional virtual image, and using the obtained three-dimensional vocalization element set to control the head pose of the three-dimensional virtual image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of three-dimensional image pronunciation, and specifically relates to a method for simulating the head posture of three-dimensional image pronunciation. Background Art

[0002] Virtual character modeling and rendering technologies have been widely used in industries such as animation, games, and movies. Enabling virtual characters to have natural and smooth lip-syncing movements when speaking is the key to enhancing the user experience. In a real-time system, it is necessary to synchronously play the audio obtained in real-time in the form of a stream and the virtual character image rendered synchronously. During this process, it is necessary to ensure the synchronization between the audio and the character's lip movements.

[0003] The Chinese invention patent with the publication number CN111081270B (application number: CN201911314031.3) discloses a method for synchronously controlling the lip movements of a virtual character driven by real-time audio, including the following steps: the step of identifying the viseme probability from the real-time speech stream; the step of filtering the viseme probability; the step of converting the sampling rate of the viseme probability to the same sampling rate as the virtual character rendering frame rate; the step of converting the viseme probability to a standard lip configuration and performing lip rendering. This method can avoid the requirement of synchronously transmitting phoneme sequences or lip sequence information when transmitting the audio stream, can significantly reduce the system complexity, coupling degree, and implementation difficulty, and is applicable to various application scenarios of rendering virtual characters on display devices.

[0004] In the above-mentioned invention and many current three-dimensional images during the pronunciation process, there are only simple lip movements, and there is no linkage between the head posture and the pronunciation, making the three-dimensional image pronunciation process rigid and unnatural. Summary of the Invention

[0005] In view of this, the present invention provides a method for simulating the head posture of three-dimensional image pronunciation, which can solve the technical problem that in the pronunciation process of three-dimensional images, there are only simple lip movements, and there is no linkage between the head posture and the pronunciation, making the three-dimensional image pronunciation process rigid and unnatural.

[0006] The present invention is implemented as follows:

[0007] The present invention provides a method for simulating the head posture of three-dimensional image pronunciation, which includes the following steps:

[0008] S100: Collect the two-dimensional head posture elements of a person speaking in a planar video library as a training data set, and establish the corresponding relationship between the two-dimensional head posture elements, where the two-dimensional head posture elements include phonemes and two-dimensional head postures;

[0009] S200: Establish a neural network model and train it using a training dataset, where the phonemes in the training dataset are used as inputs and the two-dimensional head pose is used as the output;

[0010] S300: Conduct three-dimensional recognition on the pronunciation processes of multiple testers to obtain three-dimensional head pose elements, where the three-dimensional head pose elements include phonemes and three-dimensional head poses;

[0011] S400: Adopt a transfer learning algorithm to update and optimize the neural network model using the three-dimensional head pose elements to form a three-dimensional pronunciation neural network model;

[0012] S500: Use the string that requires three-dimensional image pronunciation as text data, extract phonemes from the text data to obtain a pronunciation phoneme set;

[0013] S600: Use the three-dimensional pronunciation neural network model to calculate the obtained pronunciation phoneme set to obtain a three-dimensional vocalization element set;

[0014] S700: Establish a three-dimensional virtual image and use the obtained three-dimensional vocalization element set to control the head pose of the three-dimensional virtual image;

[0015] Among them, the head pose is described using the coordinates of the head pose key points.

[0016] Based on the above technical solutions, a method for simulating the head pose of three-dimensional image pronunciation of the present invention can be further improved as follows:

[0017] Among them, the specific steps of collecting the two-dimensional head pose elements of the person speaking in the planar video library in step S100 include:

[0018] The first step: Select the video clip of the person speaking in the planar video library, where the video clip includes video audio and multiple frames;

[0019] The second step: Establish a two-dimensional coordinate system for the image in each frame, and at the same time use a face recognition method to recognize the face;

[0020] The third step: Set two-dimensional key points on the recognized face and record the two-dimensional coordinates of the head pose key points in the plane.

[0021] The fourth step: Collect the phonemes in the corresponding video audio for each frame.

[0022] Further, the specific steps of establishing the correspondence relationship between the two-dimensional head pose elements in step S100 are: Establish the correspondence relationship between phonemes and two-dimensional head poses according to the time axis of the video.

[0023] Among them, the specific steps of three-dimensional recognition of the pronunciation processes of multiple testers in step S300 to obtain three-dimensional head pose elements include:

[0024] The first step: Set green three-dimensional recognition points on the faces of multiple testers;

[0025] The second step: Establish a three-dimensional coordinate system, and mark the three-dimensional coordinates of each three-dimensional recognition point to form the three-dimensional coordinates of the key points of the head pose;

[0026] The third step: On the five directions of front, back, left, right, and top of the testers' faces, set 5 cameras perpendicular to each other;

[0027] [[ID=!2]]The fourth step: The testers conduct conversation pronunciation or reading aloud, and use the cameras to collect videos of the conversation pronunciation or reading aloud processes of the testers;

[0028] The fifth step: Record the coordinates of each three-dimensional recognition point at different moments in the collected videos to form a three-dimensional recognition point coordinate sequence as the three-dimensional head pose element; recognize the pronunciation phonemes of the testers at different moments to form a pronunciation phoneme sequence; and establish the relationship between the three-dimensional head pose element and the pronunciation phoneme sequence according to the moments.

[0029] Further, the three-dimensional recognition points at least include the points where the tip of the nose, the midpoint of the chin, the left eye corner, the right eye corner, the left mouth corner, the right mouth corner, the center of the top of the head, the top of the left ear, the lower tip of the left ear, the top of the right ear, and the lower tip of the right ear of the testers' faces are located.

[0030] [[ID=!1]]Among them, the specific steps of using the obtained three-dimensional vocalization element set to control the head pose of the three-dimensional virtual image in step S700 include:

[0031] First, define a standard mouth shape configuration for each head pose, and the standard mouth shape configuration is a key frame or a parameter describing the mouth shape;

[0032] Second, convert the visual element probability into a mixing ratio of the standard mouth shape configuration through a mapping function; among them, in the key frame scenario, the mixing ratio is the interpolation ratio between different key frames;

[0033] In the scenario of key point parameters, bone parameters, or blenshape parameters, the mixing ratio is the mixing ratio of key point parameters, bone parameters, or blenshape parameters.

[0034] Further, the key points of the head pose at least include the points where the tip of the nose, the midpoint of the chin, the left eye corner, the right eye corner, the left mouth corner, the right mouth corner, the center of the top of the head, the top of the left ear, the lower tip of the left ear, the top of the right ear, and the lower tip of the right ear of the face are located.

[0035] Furthermore, the two-dimensional head posture element also includes two-dimensional visual elements, and the three-dimensional head posture element also includes three-dimensional visual elements.

[0036] Among them, the number of characters whose pronunciation lasts for more than 2 minutes in the planar video library is greater than 10,000.

[0037] Furthermore, the number of the multiple testers is greater than 10, and the duration of the testers' conversation pronunciation or reading is greater than 10 minutes.

[0038] Compared with the prior art, the beneficial effects of the three-dimensional image pronunciation head posture simulation method provided by the present invention are: using head posture key points to describe head movements, wherein the head posture key points include at least the points where the nose tip, the midpoint of the lower jaw, the left eye corner, the right eye corner, the left mouth corner, the right mouth corner, the center of the top of the head, the left ear top, the lower tip of the left ear, the right ear top, and the lower tip of the right ear are located; using a three-dimensional pronunciation neural network model to calculate the obtained pronunciation phoneme set, obtain a three-dimensional sound element set to establish a three-dimensional virtual image, and use the obtained three-dimensional sound element set to control the head posture of the three-dimensional virtual image; and realizing the linkage between pronunciation and head posture key point movements, solving the technical problem that the three-dimensional image pronunciation process is rigid and unnatural. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments of the present invention. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0040] Figure 1 A flowchart of a three-dimensional image pronunciation head posture simulation method provided by the present invention;

[0041] Figure 2 Parameter map for the main detection of 3D head posture;

[0042] Figure 3 To collect the positional relationship of each coordinate system in the camera when using the Dlib library;

[0043] In the accompanying drawings, the components represented by the reference numerals are as follows: DETAILED DESCRIPTION

[0044] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0045] Therefore, the detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely represents selected embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0046] It should be noted that like reference numerals and letters denote like items in the following drawings. Therefore, once an item is defined in one drawing, it does not require further definition and explanation in subsequent drawings.

[0047] In the description of the present invention, it should be understood that the terms "center", "longitudinal", "lateral", "length", "width", "thickness", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", "clockwise", "counterclockwise", etc. indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings. They are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus should not be construed as a limitation of the present invention.

[0048] In addition, the terms "first" and "second" are only used for descriptive purposes and should not be construed as indicating or implying relative importance or implicitly specifying the number of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the present invention, "a plurality" means two or more unless otherwise specifically defined.

[0049] As Figure 1 shown, it is a flowchart of a three-dimensional image pronunciation head pose simulation method provided by the present invention. The method includes the following steps:

[0050] S100: Collect two-dimensional head pose elements of a person speaking in a planar video library as a training data set, and establish the corresponding relationship between the two-dimensional head pose elements. The two-dimensional head pose elements include phonemes and two-dimensional head poses;

[0051] S200: Establish a neural network model and train it using a training dataset, where the phonemes in the training dataset are used as inputs and the two-dimensional head pose is used as the output;

[0052] S300: Conduct three-dimensional recognition on the pronunciation processes of multiple testers to obtain three-dimensional head pose elements, which include phonemes and three-dimensional head poses;

[0053] S400: Adopt a transfer learning algorithm to update and optimize the neural network model using the three-dimensional head pose elements to form a three-dimensional pronunciation neural network model;

[0054] S500: Use the string that requires three-dimensional image pronunciation as text data, extract phonemes from the text data to obtain a pronunciation phoneme set;

[0055] S600: Use the three-dimensional pronunciation neural network model to calculate the obtained pronunciation phoneme set to obtain a three-dimensional vocalization element set;

[0056] S700: Establish a three-dimensional virtual image and use the obtained three-dimensional vocalization element set to control the head pose of the three-dimensional virtual image;

[0057] Among them, the head pose is described using the coordinates of the head pose key points.

[0058] Among them, in the above technical solution, the specific steps of collecting two-dimensional head pose elements of a person speaking in the planar video library in step S100 include:

[0059] The first step: Select a video clip of a person speaking in the planar video library, where the video clip includes video audio and multiple frames;

[0060] The second step: Establish a two-dimensional coordinate system for the image in each frame, and at the same time use a face recognition method to recognize the face;

[0061] The third step: Set two-dimensional key points on the recognized face and record the two-dimensional coordinates of the head pose key points in the plane.

[0062] The fourth step: Collect the phonemes in the corresponding video audio for each frame.

[0063] Furthermore, in the above technical solution, the specific steps of establishing the correspondence relationship between two-dimensional head pose elements in step S100 are: Establish the correspondence relationship between phonemes and two-dimensional head poses according to the time axis of the video.

[0064] Among them, in the above technical solution, the specific steps of conducting three-dimensional recognition on the pronunciation processes of multiple testers to obtain three-dimensional head pose elements in step S300 include:

[0065] Step 1: Set green 3D recognition points on the faces of multiple testers;

[0066] Step 2: Establish a 3D coordinate system, and mark the 3D coordinates of each 3D recognition point to form the 3D coordinates of the key points of the head pose;

[0067] Step 3: On the five directions of front, back, left, right, and top of the tester's face, set 5 cameras perpendicular to each other;

[0068] Step 4: The tester conducts dialogue pronunciation or reading, and uses the cameras to collect videos of the tester's dialogue pronunciation or reading process;

[0069] Step 5: Record the coordinates of each 3D recognition point at different moments in the collected video to form a 3D recognition point coordinate sequence as 3D head pose elements; identify the pronunciation phonemes of the tester at different moments to form a pronunciation phoneme sequence; and establish the relationship between the 3D head pose elements and the pronunciation phoneme sequence according to the moments.

[0070] Further, in the above technical solution, the 3D recognition points at least include the points where the tip of the tester's nose, the midpoint of the chin, the left eye corner, the right eye corner, the left mouth corner, the right mouth corner, the center of the top of the head, the top of the left ear, the lower tip of the left ear, the top of the right ear, and the lower tip of the right ear are located on the tester's face.

[0071] Among them, in the above technical solution, the steps of controlling the head pose of the 3D virtual image by using the obtained 3D vocalization element set in step S700 specifically include:

[0072] First, define a standard mouth shape configuration for each head pose, and the standard mouth shape configuration is a key frame or a parameter describing the mouth shape;

[0073] Second, convert the visual element probability into a mixing ratio of the standard mouth shape configuration through a mapping function; among them, in the key frame scenario, the mixing ratio is the interpolation ratio between different key frames;

[0074] In the scenario of key point parameters, bone parameters, or blenshape parameters, the mixing ratio is the mixing ratio of key point parameters, bone parameters, or blenshape parameters. Among them, blenshape is a blend shape editor in Unity3D, which is a control that can be used for all blend shape deformers in the scene. Unity3D is a real-time 3D interactive content creation and operation platform. All creators, including game development, art, architecture, automotive design, and film and television, can turn their creativity into reality with the help of Unity. The Unity platform provides a complete set of software solutions for creating, operating, and monetizing any real-time interactive 2D and 3D content, and the supported platforms include mobile phones, tablets, PCs, game consoles, augmented reality, and virtual reality devices.

[0075] As shown Figure 2 in the figure, there are three main parameters for detecting the three-dimensional head pose, namely pitch (rotation around the X-axis), yaw (rotation around the Y-axis), and roll (rotation around the Z-axis). Their scientific names are pitch angle, yaw angle, and roll angle respectively, which are equivalent to raising the head, shaking the head, and turning the head.

[0076] In the above technical solution, the key points of the head pose at least include the points where the tip of the nose, the midpoint of the lower jaw, the left eye corner, the right eye corner, the left mouth corner, the right mouth corner, the center of the top of the head, the top of the left ear, the lower tip of the left ear, the top of the right ear, and the lower tip of the right ear of the human face are located.

[0077] Furthermore, it is also possible to first obtain 68 feature key points of the 2D human face through the dlib library, and then fit the 3D human face feature points through model matching algorithms such as the 3DMorphable Model. At this time, the positional relationships of the coordinate systems in the acquisition camera are as Figure 3 shown in the figure; in the figure, the world coordinate system is: O w -X w Y w Z w ; the camera coordinate system is: 0 c -X c Y c Z c ; the image coordinate system is: o-xy; the pixel coordinate system is: the upper left corner of the image - uv, with the unit of pixel. P is a point in the world coordinate system; p is the projection point of P in the image coordinate system; f is the focal length, that is, the distance from Oc to o.

[0078] Among them, there is dlib's face recognition. The dlib library is a complete machine learning library, and face recognition is just one of its subsets. The dlib library uses 68-point position marks to identify important parts of the face. For example, points 18 - 22 mark the right eyebrow, and points 51 - 68 mark the mouth. The implementation idea of Dlib is achieved through five parts: face detection, face alignment, face representation, and face matching. Face Detection. Detect the face area from the input image and return the coordinates of the face bounding box. Face Alignment. Detect the facial feature points from the face area and perform a normalization operation on the face based on the feature points, making the scale and angle of the face area consistent, which is convenient for feature extraction and face matching. The ultimate goal of face alignment is to locate the precise shape of the face within the known face bounding box, mainly divided into two categories: optimization-based methods and regression-based methods. Optimization-based methods mainly come from deep network models, such as convolutional neural networks (CNNs), deep autoencoders (DAEs), and restricted Boltzmann machines (RBMs), etc., to model the changes in face shape and appearance, and then obtain the non-linear mapping from face appearance to shape. Optimization-based methods can be regarded as learning a regression function, with the image as the input, outputting the positions of the feature points (face shape), and constructing a cascaded regression model. Face Representation. Extract features from the normalized face area to obtain feature vectors. For example, some deep neural network methods use 128 feature vectors to represent a face. Ideally, the feature vectors extracted from photos of different people are different, while similar feature vectors can be extracted from different photos of the same person. The main idea of deep learning for face recognition is that different faces are composed of different features. Simply put, features can include eyelids, noses, eyes, skin color, hair color, etc. In the dlib library, 68 relatively important feature points (landmarks) are adopted, as can be seen from the following picture, and they are vectorized. 3DMM is a relatively basic three-dimensional face statistical model, first proposed to solve the problem of restoring three-dimensional shapes from two-dimensional face images. In the two decades since the development of the 3DMM method, scholars have carried out data expansion and in-depth research on it. Also, due to the widespread use of neural networks, the optimization of 3DMM parameters has been simplified, and there have been an endless stream of articles on three-dimensional reconstruction based on the 3DMM method. However, such methods represent any face based on a statistical model of a set of face shapes and textures, and there are still problems such as poor discriminability of the reconstructed face and difficulty in solving parameters. Currently, it is also a key research direction in the academic community.

[0079] Furthermore, in the above technical solution, the two-dimensional head pose elements further include two-dimensional visual elements, and the three-dimensional head pose elements further include three-dimensional visual elements.

[0080] Among them, in the above technical solution, the number of people in the plane video library whose pronunciation exceeds 2 minutes is greater than 10,000.

[0081] Furthermore, in the above technical solution, the number of multiple testers is greater than 10, and the duration of the testers' dialogue pronunciation or reading aloud is greater than 10 minutes.

[0082] The above is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.

Claims

1. A three-dimensional image pronunciation head posture simulation method, characterized in that, It includes the following steps: S100: Collect two-dimensional head pose elements when a person in a planar video library is speaking as a training data set, and establish the corresponding relationship between the two-dimensional head pose elements. The two-dimensional head pose elements include phonemes and two-dimensional head poses; S200: Establish a neural network model and train it using the training data set, where the phonemes in the training data set are used as inputs and the two-dimensional head poses are used as outputs; S300: Perform three-dimensional recognition on the pronunciation processes of multiple testers to obtain three-dimensional head pose elements. The three-dimensional head pose elements include phonemes and three-dimensional head poses; S400: Use a transfer learning algorithm to update and optimize the neural network model using the three-dimensional head pose elements to form a three-dimensional pronunciation neural network model; S500: Use the string that needs to be pronounced in three-dimensional form as text data, extract phonemes from the text data to obtain a pronunciation phoneme set; S600: Use the three-dimensional pronunciation neural network model to calculate the obtained pronunciation phoneme set to obtain a three-dimensional vocalization element set; S700: Establish a three-dimensional virtual image and use the obtained three-dimensional vocalization element set to control the head pose of the three-dimensional virtual image; Among them, the head pose is described by the coordinates of the head pose key points.

2. The three-dimensional image pronunciation head pose simulation method according to claim 1, characterized in that The specific steps of collecting two-dimensional head pose elements when a person in a planar video library is speaking in step S100 include: The first step: Select a video clip of a person speaking in a planar video library, where the video clip includes video audio and multiple frames; The second step: Establish a two-dimensional coordinate system for each frame of the image, and at the same time use a face recognition method to recognize the face; The third step: Set two-dimensional key points on the recognized face and record the two-dimensional coordinates of the head pose key points in the plane; The fourth step: Collect the phonemes in the corresponding video audio for each frame.

3. A three-dimensional image pronunciation head pose simulation method according to claim 2, characterized in that The specific steps of establishing the corresponding relationship between the two-dimensional head pose elements in step S100 are: Establish the corresponding relationship between phonemes and two-dimensional head poses according to the time axis of the video.

4. A three-dimensional image pronunciation head posture simulation method according to claim 1, characterized in that The specific steps of performing three-dimensional recognition on the pronunciation processes of multiple testers to obtain three-dimensional head pose elements in step S300 include: The first step: Set green three-dimensional recognition points on the faces of multiple testers. The three-dimensional recognition points at least include the points where the tip of the nose, the midpoint of the chin, the left eye corner, the right eye corner, the left mouth corner, the right mouth corner, the center of the top of the head, the top of the left ear, the lower tip of the left ear, the top of the right ear, and the lower tip of the right ear of the tester's face are located; The second step: Establish a three-dimensional coordinate system and mark the three-dimensional coordinates of each three-dimensional key point to form the three-dimensional coordinates of the head pose key points; The third step: Set 5 cameras perpendicular to each other in the front, back, left, right, and up directions on the tester's face; The fourth step: The tester conducts a conversation pronunciation or reading, and use the camera to collect the conversation pronunciation or reading process of the tester; The fifth step: Record the coordinates of each three-dimensional key point at different times to form a three-dimensional key point coordinate sequence; recognize the pronunciation phonemes of the tester at different times to form a pronunciation phoneme sequence; and establish the relationship between the three-dimensional key point coordinate sequence and the pronunciation phoneme sequence according to the time.

5. A three-dimensional image pronunciation head pose simulation method according to claim 4, characterized in that, The steps of collecting the conversation pronunciation or reading process of the tester by using a camera specifically include: The two-dimensional facial key points of the facial image collected by the camera; Determine the three-dimensional target key points from the three-dimensional facial key points in the pre-made three-dimensional facial model according to the two-dimensional facial key points; Calculate the camera parameters according to the two-dimensional facial key points and the three-dimensional target key points, where the camera parameters include the following: rotation angle parameter, translation amount, and scaling value; Perform pronunciation transformation processing on the three-dimensional facial model according to the camera parameters and the three-dimensional target key points to obtain the three-dimensional pronunciation parameters corresponding to the two-dimensional facial key points; Perform sparsification processing on the three-dimensional pronunciation parameters to obtain sparsified pronunciation parameters; Migrate the sparsified pronunciation parameters and the rotation angle parameter to the animation model of the virtual character, so that the pronunciation of the virtual character is consistent with the pronunciation of the facial image; The three-dimensional facial model is reconstructed by parameterizing the facial model 3DMM through a pronunciation fusion model; the three-dimensional pronunciation parameters include pronunciation sub-parameters corresponding to different facial parts; The step of performing sparsification processing on the three-dimensional pronunciation parameters to obtain sparsified pronunciation parameters includes: Take the pronunciation sub-parameters corresponding to each facial part as target sub-parameters and perform the following operations: Query the target pronunciation sub-model corresponding to the target sub-parameters in the pre-stored pronunciation model mapping table; where the pronunciation model mapping table stores the sub-model identifiers corresponding to the pronunciation sub-models of different facial parts included in the pronunciation fusion model; Assign the target sub-parameters to the target pronunciation sub-model; Set the pronunciation sub-parameters corresponding to the other pronunciation sub-models in the pronunciation fusion model except the target pronunciation sub-model to preset parameters; Fuse the target pronunciation sub-model and the other pronunciation sub-models to obtain a target three-dimensional facial model; Calculate the vertex deformation amount corresponding to the target sub-parameters based on the target three-dimensional facial model and a preset three-dimensional facial model; where the preset three-dimensional facial model is obtained by fusing multiple pronunciation sub-models with the pronunciation sub-parameters being the preset parameters; Input the vertex deformation amounts corresponding to each target sub-parameter and the three-dimensional pronunciation parameters into an optimization model for iterative calculation until the loss value of the optimization model reaches a preset loss value, and output the optimization result; Take the optimization result as the sparsified pronunciation parameters.

6. A three-dimensional image pronunciation head posture simulation method according to claim 1, characterized in that, The steps of controlling the head posture of the three-dimensional virtual image by using the obtained three-dimensional voice generation element set in step S700 specifically include: First, define a standard mouth configuration for each head posture, and the standard mouth configuration is a key frame or a parameter describing the mouth shape; Secondly, convert the visual element probability into a mixing ratio of the standard mouth configuration through a mapping function; where in the key frame scenario, the mixing ratio is the interpolation ratio between different key frames; In the scenario of key point parameters, bone parameters, or blenshape parameters, the mixing ratio is the mixing ratio of key point parameters, bone parameters, or blenshape parameters.

7. A three-dimensional image pronunciation head posture simulation method according to any one of claims 1-6, characterized in that, The head pose key points at least include the points where the tip of the nose, the midpoint of the chin, the left eye corner, the right eye corner, the left mouth corner, the right mouth corner, the center of the top of the head, the top of the left ear, the lower tip of the left ear, the top of the right ear, and the lower tip of the right ear of the human face are located.

8. A three-dimensional image pronunciation head posture simulation method according to any one of claims 1-6, characterized in that, The two-dimensional head pose elements further include two-dimensional visual elements, and the three-dimensional head pose elements further include three-dimensional visual elements.

9. A three-dimensional image pronunciation head pose simulation method according to claim 1, characterized in that The number of people in the plane video library whose pronunciation exceeds 2 minutes is greater than 10,000.

10. A three-dimensional image pronunciation head posture simulation method according to claim 4, characterized in that The number of the multiple testers is greater than 10, and the duration of the testers' dialogue pronunciation or reading aloud is greater than 10 minutes.

Citation Information

Patent Citations

  • A Real-Time Audio-Driven Method for Lip-Sync Control of Virtual Characters

    CN111081270B

  • Voice synchronous-drive three-dimensional face mouth shape and face posture animation method

    CN103218842A

  • Driver posture recognition method based on depth images and virtual data

    CN108345869A