A three-dimensional virtual image lip matching method with high simulation degree, medium and system

CN116543078BActive Publication Date: 2026-09-15JINDONG CULTURE TECHNOLOGY CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202310433938.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-21
Publication Date
2026-09-15
Estimated Expiration
2043-04-21

AI Technical Summary

Technical Problem

[0005]上述专利中,采用测试人员朗读文本的方法,利用三台摄像机进行录像进行采集口型,这一过程采集工作量极大

Benefits of technology

[0059] 1. The three-dimensional mouth shapes of each phoneme are collected by the experimenters, and the obtained three-dimensional mouth shapes of the phonemes are used to process each mouth shape in the frontal pronunciation mouth shape set in three dimensions, which can convert the frontal pronunciation mouth shape from two dimensions to three dimensions; the side pronunciation mouth shapes are used to process the first pronunciation mouth shape set, making the mouth shape smoother and avoiding the unnatural phenomenon in the process of simply converting the planar mouth shape into a three-dimensional mouth shape in the existing technology.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116543078B_ABST
    Figure CN116543078B_ABST
Patent Text Reader

Abstract

The application provides a three-dimensional virtual image mouth shape matching method with high simulation degree, a medium and a system, and belongs to the technical field of three-dimensional virtual technology. The method comprises the following steps: obtaining a person's front speaking video and a person's side speaking video in a movie work and a life monitoring video, and processing the front speaking video and the side speaking video to obtain a front video and a side video; obtaining a mouth shape corresponding to each pronunciation, and recording the mouth shape as a front pronunciation mouth shape set; obtaining a mouth shape corresponding to each pronunciation, and recording the mouth shape as a side pronunciation mouth shape set; obtaining a three-dimensional mouth shape corresponding to each phoneme of an experimental staff, and recording the three-dimensional mouth shape as a phoneme three-dimensional mouth shape set; performing three-dimensional processing on each mouth shape in the front pronunciation mouth shape set according to the phoneme three-dimensional mouth shape set, to obtain a first pronunciation mouth shape set; performing fluency processing on each mouth shape in the first pronunciation mouth shape set according to the side pronunciation mouth shape set, to obtain a second pronunciation mouth shape set; performing mouth shape matching on each pronunciation in a speaking process of a three-dimensional image according to the second pronunciation mouth shape set, to generate a three-dimensional image speaking mouth shape sequence.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of three-dimensional virtual technology, and specifically relates to a method, medium, and system for lip-syncing of a highly realistic three-dimensional virtual image. Background Technology

[0002] Lip-syncing is a key perspective in character facial animation. The realism and naturalness of lip-syncing animation directly affect the overall realism of the character's facial animation. Therefore, lip-syncing animation plays an important role in human-computer interaction methods such as film, games, and virtual reality.

[0003] In existing technologies, most methods use phoneme lip-syncing to match the lip movements of 3D virtual avatars. However, since there are certain differences between the lip movements of each sound in human language and the lip movements of the sounds broken down into phonemes, the pronunciation of 3D virtual avatars is not smooth. In addition, there are unnatural phenomena in the process of simply converting planar lip movements into 3D lip movements.

[0004] The applicant's Chinese invention patent application, with publication number CN115690280B and application number CN202211687841.5, discloses a three-dimensional image pronunciation lip-shape simulation method. The specific steps include: attaching multiple small colored blocks to the mouth of a tester; the tester reading text aloud; and collecting a video recording of the tester's reading. The reading video is then split according to the phonemes in the audio to obtain a phoneme video set, and the movement trajectories of the small colored blocks corresponding to the phoneme changes in adjacent video recordings are recorded as a phoneme change small colored block trajectory set. A three-dimensional virtual human mouth model is established, and a lip-shape model corresponding to each phoneme is established based on the stable coordinate set of the single-phoneme small colored blocks. A lip-shape model sequence is established according to the text to be read, and the lip-shape change process is constructed for adjacent lip shapes in the lip-shape model sequence using the phoneme change small colored block trajectory set.

[0005] In the aforementioned patent, the method of having testers read the text aloud and using three cameras to record and capture lip movements is employed, a process that involves an enormous amount of work. Summary of the Invention

[0006] In view of this, the present invention provides a highly realistic 3D virtual image lip-syncing method, medium, and system, which can solve the technical problem that in the prior art, most 3D virtual image lip-syncing is performed using the phoneme lip-syncing method. However, since the lip shape of each sound in human language is different from the lip shape of the sound when it is broken down into phonemes, the 3D virtual image does not speak fluently.

[0007] This invention is implemented as follows:

[0008] The first aspect of the present invention provides a highly realistic lip-syncing method for a three-dimensional virtual image, comprising the following steps:

[0009] S10: Acquire video recordings of a person speaking from the front and from the side in film and television works and in daily life surveillance videos, wherein the video recording of a person speaking from the front is a video consisting of continuous frames in which the entire mouth area of ​​the person can be displayed; the video recording of a person speaking from the side is a video consisting of continuous frames in which the entire mouth area of ​​the person is partially displayed.

[0010] S20: Perform distortion correction on the video recording of the person speaking from the front and the video recording of the person speaking from the side to obtain the front video and the side video;

[0011] S30: Obtain the lip shape corresponding to each pronunciation based on the frontal video recording and record it as the frontal pronunciation lip shape set; obtain the lip shape corresponding to each pronunciation based on the side video recording and record it as the side pronunciation lip shape set.

[0012] S40: Obtain the three-dimensional lip movements corresponding to each phoneme from the experimenter, and denot it as the three-dimensional lip movement set of the phoneme;

[0013] S50: Perform three-dimensional processing on each mouth shape in the frontal pronunciation mouth shape set according to the three-dimensional mouth shape set of phonemes to obtain the first pronunciation mouth shape set;

[0014] S60: Based on the lateral articulation set, each articulation shape in the first articulation set is smoothed to obtain the second articulation set;

[0015] S70: Match the mouth shape of each sound in the three-dimensional image's speech process according to the second set of mouth shapes, and generate a three-dimensional image's speech mouth shape sequence.

[0016] The technical effects of the highly realistic 3D virtual image lip-syncing method provided by this invention are as follows: By collecting frontal and side-view recordings of people speaking from film and television works and daily surveillance videos, a large number of people speaking videos can be obtained at low cost, without the need for experimental personnel to collect lip-sync data for each individual lip shape; by collecting 3D lip-sync data for each phoneme and using the obtained 3D lip-sync data to process each lip shape in the frontal pronunciation lip shape set, the frontal pronunciation lip shape can be converted from 2D to 3D; by using side-view pronunciation lip shape to process the first pronunciation lip shape set, the lip-sync becomes smoother.

[0017] Based on the above technical solution, the highly realistic 3D virtual image lip-syncing method of the present invention can be further improved as follows:

[0018] The step of obtaining the lip shape corresponding to each pronunciation based on the frontal video recording and recording it as a frontal pronunciation lip shape set specifically includes:

[0019] Extract the audio from the frontal pronunciation lip shape set and denote it as frontal audio;

[0020] In the frontal audio, the frontal peak curve is formed by connecting the peak points that reflect the amplitude change trend of the speech signal in the frontal audio.

[0021] The first frontal curve is obtained by smoothing and normalizing the frontal peak curve.

[0022] The first positive curve is divided by amplitude to obtain a set of multiple curve segments, which is denoted as the first positive curve segment set.

[0023] The curve segments in the first positive curve segment set are used as the speech signals corresponding to each pronunciation, and the start and end times of the speech signals are the start and end times corresponding to the pronunciation.

[0024] Based on the start and end times corresponding to the pronunciation, the frontal video is divided into multiple frontal pronunciation lip shape sets.

[0025] Furthermore, the step of dividing the curve by amplitude to obtain a set of multiple curve segments specifically includes:

[0026] Obtain the troughs of the curve;

[0027] Using the point where the trough is located as the cutting point, the curve is divided by amplitude to obtain a set of multiple curve segments.

[0028] The step of obtaining the lip shape corresponding to each pronunciation based on the side video recording and recording it as a side pronunciation lip shape set specifically includes:

[0029] Extract the audio from the set of lateral pronunciation mouth shapes and denote it as lateral audio;

[0030] In the side audio, the side peak curve is formed by connecting the peak points that reflect the amplitude change trend of the speech signal in the side audio.

[0031] The first side curve is obtained by smoothing and normalizing the side peak curve;

[0032] The first side curve is divided by amplitude to obtain a set of multiple curve segments, which is denoted as the first side curve segment set;

[0033] The curve segments in the first side curve segment set are used as the speech signals corresponding to each pronunciation, and the start and end times of the speech signals are the start and end times corresponding to the pronunciation.

[0034] Based on the start and end times corresponding to the pronunciation, the side video is divided into multiple side pronunciation lip shape sets.

[0035] Furthermore, the step of obtaining the three-dimensional lip movements corresponding to each phoneme by the experimenter specifically includes:

[0036] Acquire laser data corresponding to the laser scanner during the process of the experimenter reading text aloud;

[0037] A point cloud model is established based on the laser data;

[0038] The points in the point cloud model include at least the key points corresponding to the mouth in the Dlib facial recognition algorithm.

[0039] The three-dimensional mouth shape corresponding to each phoneme is obtained based on the point cloud model;

[0040] The text read aloud by the experimenters was a collection of all phonemes, and the experimenters read each phoneme one by one at intervals.

[0041] Furthermore, the specific steps for obtaining the three-dimensional lip shape corresponding to each phoneme based on the point cloud model include:

[0042] The key points in the point cloud model are filtered out, and only the key points corresponding to the mouth in the Dlib face recognition algorithm are retained as the mouth shape point cloud model.

[0043] The point cloud model of the mouth shape corresponding to each phoneme read aloud by the experimenter is obtained as a three-dimensional mouth shape.

[0044] The step of performing three-dimensional processing on each mouth shape in the frontal articulation mouth shape set according to the three-dimensional mouth shape set of phonemes to obtain the first articulation mouth shape set specifically includes:

[0045] For each mouth shape in the aforementioned frontal pronunciation mouth shape set, obtain the corresponding phoneme;

[0046] The three-dimensional lip shapes corresponding to the three-dimensional lip shape set are obtained according to the corresponding phonemes;

[0047] The three-dimensional processing of each mouth shape in the frontal articulation mouth shape set is performed using the corresponding three-dimensional mouth shape, specifically including:

[0048] Obtain the key points corresponding to the mouth in the Dlib facial recognition algorithm for each mouth shape;

[0049] The 3D coordinates of each key point are adjusted based on the 3D coordinates of the corresponding key points in the 3D lip shape.

[0050] The step of performing smoothing processing on each lip shape in the first lip shape set based on the lateral lip shape set to obtain the second lip shape set specifically includes:

[0051] For each mouth shape in the first set of pronunciation mouth shapes, obtain the corresponding phoneme;

[0052] Obtain the side mouth shape corresponding to the side pronunciation mouth shape set according to the corresponding phoneme;

[0053] The first set of articulation mouth shapes is processed in three dimensions using the corresponding lateral mouth shapes, specifically including:

[0054] Obtain the key points of the mouth corresponding to each mouth shape in the Dlib facial recognition algorithm of the first set of pronunciation mouth shapes;

[0055] The 3D coordinates of each key point are adjusted based on the 3D coordinates of the corresponding key point in the side profile.

[0056] A second aspect of the present invention provides a computer-readable storage medium, wherein the computer-readable storage medium stores program instructions, which, when executed, are used to perform the above-described method for lip-syncing of a highly realistic three-dimensional virtual image.

[0057] A third aspect of the present invention provides a highly realistic three-dimensional virtual avatar lip-syncing system, wherein the highly realistic three-dimensional virtual avatar lip-syncing system includes the aforementioned computer-readable storage medium.

[0058] Compared with existing technologies, the beneficial effects of the highly realistic 3D virtual image lip-syncing method, medium, and system provided by this invention are:

[0059] 1. The three-dimensional mouth shapes of each phoneme are collected by the experimenters, and the obtained three-dimensional mouth shapes of the phonemes are used to process each mouth shape in the frontal pronunciation mouth shape set in three dimensions, which can convert the frontal pronunciation mouth shape from two dimensions to three dimensions; the side pronunciation mouth shapes are used to process the first pronunciation mouth shape set, making the mouth shape smoother and avoiding the unnatural phenomenon in the process of simply converting the planar mouth shape into a three-dimensional mouth shape in the existing technology.

[0060] 2. By collecting videos of people speaking from the front and side in film and television works and daily surveillance footage, a large number of videos of people speaking can be obtained at low cost, without the need for experimental personnel to collect every lip movement. Attached Figure Description

[0061] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments of the present invention will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0062] Figure 1 A flowchart of a highly realistic 3D virtual image lip-syncing method is provided as a first aspect of the present invention;

[0063] Figure 2 This is a schematic diagram of facial key points in the Dlib face recognition algorithm. Detailed Implementation

[0064] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0065] Therefore, the following detailed description of the embodiments of the invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the invention without inventive effort are within the scope of protection of the invention.

[0066] It should be noted that similar labels and letters in the following figures indicate similar items. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures.

[0067] In the description of this invention, it should be understood that the terms "center," "longitudinal," "lateral," "length," "width," "thickness," "upper," "lower," "front," "rear," "left," "right," "vertical," "horizontal," "top," "bottom," "inner," "outer," "clockwise," and "counterclockwise," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are only for the convenience of describing this invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on this invention.

[0068] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this invention, "a plurality of" means two or more, unless otherwise explicitly specified.

[0069] like Figure 1 The diagram shown is a flowchart of a highly realistic 3D virtual image lip-syncing method provided by the first aspect of this invention. This method includes the following steps:

[0070] S10: Acquire video recordings of people speaking from the front and from the side in film and television works and daily surveillance footage. The video recording of a person speaking from the front consists of continuous frames in which the entire mouth area of ​​the person is visible. The video recording of a person speaking from the side consists of continuous frames in which only part of the mouth area of ​​the person is visible.

[0071] S20: Perform distortion correction on the video recordings of a person speaking from the front and from the side to obtain the frontal and side video recordings;

[0072] S30: Obtain the lip shape corresponding to each sound from the front video recording and record it as the front sound lip shape set; obtain the lip shape corresponding to each sound from the side video recording and record it as the side sound lip shape set.

[0073] S40: Obtain the three-dimensional lip movements corresponding to each phoneme from the experimenter, and denot it as the three-dimensional lip movement set of the phoneme;

[0074] S50: Perform three-dimensional processing on each mouth shape in the frontal articulation mouth shape set according to the three-dimensional mouth shape set of phonemes to obtain the first articulation mouth shape set;

[0075] S60: Based on the lateral articulation set, each articulation shape in the first articulation set is smoothed to obtain the second articulation set;

[0076] S70: Match the mouth shape of each sound in the three-dimensional image's speech process according to the second set of mouth shapes, and generate a three-dimensional image's speech mouth shape sequence.

[0077] Distortion correction can be performed using the cv2.getOptimalNewCameraMatrix() function in OpenCV.

[0078] In the above technical solution, the step of obtaining the lip shape corresponding to each pronunciation from the frontal video recording and recording it as a frontal pronunciation lip shape set specifically includes:

[0079] Extract the audio from the frontal pronunciation lip shape set and denote it as frontal audio;

[0080] In the frontal audio, the frontal peak curve is formed by connecting the peak points that reflect the amplitude change trend of the speech signal in the frontal audio.

[0081] The first positive curve is obtained by smoothing and normalizing the positive peak curve;

[0082] The first positive curve is divided by amplitude to obtain a set of multiple curve segments, which is denoted as the first positive curve segment set;

[0083] The curve segments in the first positive curve segment set are taken as the speech signals corresponding to each pronunciation, and the start and end times of the speech signals are the start and end times corresponding to the pronunciation.

[0084] Based on the start and end times of the pronunciation, the frontal video is divided into multiple frontal pronunciation lip shape sets.

[0085] Furthermore, in the above technical solution, the step of dividing the curve by amplitude to obtain a set of multiple curve segments specifically includes:

[0086] Obtain the troughs of the curve;

[0087] By using the trough as the cutting point, the curve is divided by amplitude, resulting in a set of multiple curve segments.

[0088] In the above technical solution, the step of obtaining the lip shape corresponding to each pronunciation based on the side video recording and recording it as a side pronunciation lip shape set specifically includes:

[0089] Extract the audio from the set of lateral pronunciation mouth shapes and denote it as lateral audio;

[0090] In the side audio, the side peak curve is formed by connecting the peak points that reflect the amplitude change trend of the speech signal in the side audio.

[0091] The first side curve is obtained by smoothing and normalizing the side peak curve;

[0092] The first side curve is divided by amplitude to obtain a set of multiple curve segments, which is denoted as the first side curve segment set;

[0093] The curve segments in the first side curve segment set are taken as the speech signal corresponding to each pronunciation, and the start and end times of the speech signal are the start and end times corresponding to the pronunciation.

[0094] Based on the start and end times of the pronunciation, the side-view video is divided into multiple sets of side-view pronunciation mouth shapes.

[0095] Furthermore, in the above technical solution, the step of obtaining the three-dimensional lip movements corresponding to each phoneme by the experimenter specifically includes:

[0096] Acquire laser data corresponding to the laser scanner during the process of the experimenter reading text aloud;

[0097] A point cloud model was built based on laser data;

[0098] The points in the point cloud model should include at least the key points corresponding to the mouth in the Dlib face recognition algorithm;

[0099] The three-dimensional mouth shape corresponding to each phoneme is obtained based on the point cloud model;

[0100] The text read aloud by the experimenters was a collection of all phonemes, and the experimenters read each phoneme one by one at intervals.

[0101] like Figure 2As shown, Dlib is a machine recognition algorithm that can be used for facial recognition. In this algorithm, the key point numbers corresponding to the mouth are 49 to 68.

[0102] Furthermore, in the above technical solution, the specific steps for obtaining the three-dimensional lip shape corresponding to each phoneme based on the point cloud model include:

[0103] The key points in the point cloud model are filtered out, and only the key points corresponding to the mouth in the Dlib face recognition algorithm are retained as the mouth shape point cloud model.

[0104] The point cloud model of the mouth shape corresponding to each phoneme read aloud by the experimenter is obtained as a three-dimensional mouth shape.

[0105] In the above technical solution, the step of performing three-dimensional processing on each mouth shape in the frontal articulation mouth shape set to obtain the first articulation mouth shape set based on the three-dimensional phoneme mouth shape set specifically includes:

[0106] For each mouth shape in the frontal pronunciation mouth shape set, obtain the corresponding phoneme;

[0107] Obtain the corresponding three-dimensional lip shape from the three-dimensional lip shape set based on the corresponding phoneme;

[0108] The corresponding three-dimensional lip shapes are used to perform three-dimensional processing on each lip shape in the frontal pronunciation lip shape set, specifically including:

[0109] Obtain the key points corresponding to the mouth in the Dlib facial recognition algorithm for each mouth shape;

[0110] The 3D coordinates of each key point are adjusted based on the 3D coordinates of the corresponding key points in the 3D lip shape.

[0111] In the above technical solution, the step of smoothing each mouth shape in the first set of mouth shapes according to the lateral mouth shape set to obtain the second set of mouth shapes specifically includes:

[0112] For each mouth shape in the first set of pronunciation mouth shapes, obtain the corresponding phoneme;

[0113] Obtain the corresponding side mouth shape from the set of side mouth shapes based on the corresponding phonemes;

[0114] The first set of articulation mouth shapes is processed in three dimensions using the corresponding lateral mouth shapes, specifically including:

[0115] Obtain the key points of the mouth corresponding to each mouth shape in the Dlib facial recognition algorithm of the first set of pronunciation mouth shapes;

[0116] The 3D coordinates of each key point are adjusted based on the 3D coordinates of the corresponding key point in the side profile.

[0117] A second aspect of the present invention provides a computer-readable storage medium, wherein the computer-readable storage medium stores program instructions, which, when executed, are used to perform the above-described method for lip-syncing of a highly realistic three-dimensional virtual image.

[0118] A third aspect of the present invention provides a highly realistic three-dimensional virtual avatar lip-syncing system, wherein the highly realistic three-dimensional virtual avatar lip-syncing system includes the aforementioned computer-readable storage medium.

[0119] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A highly realistic 3D virtual image lip-syncing method, characterized in that, Includes the following steps: S10: Acquire video recordings of a person speaking from the front and from the side in film and television works and in daily life surveillance videos, wherein the video recording of a person speaking from the front is a video consisting of continuous frames in which the entire mouth area of ​​the person can be displayed; the video recording of a person speaking from the side is a video consisting of continuous frames in which the entire mouth area of ​​the person is partially displayed. S20: Perform distortion correction on the video recording of the person speaking from the front and the video recording of the person speaking from the side to obtain the front video and the side video; S30: Obtain the mouth shape corresponding to each pronunciation based on the frontal video recording and record it as the frontal pronunciation mouth shape set; obtain the mouth shape corresponding to each pronunciation based on the side video recording and record it as the side pronunciation mouth shape set. S40: Obtain the three-dimensional lip movements corresponding to each phoneme from the experimenter, and denot it as the three-dimensional lip movement set of the phoneme; S5 0: Perform three-dimensional processing on each mouth shape in the frontal pronunciation mouth shape set according to the three-dimensional mouth shape set of the phonemes to obtain the first pronunciation mouth shape set; S60: Based on the lateral articulation set, each articulation shape in the first articulation set is smoothed to obtain the second articulation set; S70: Match the mouth shape of each sound in the three-dimensional image's speech process according to the second set of mouth shapes, and generate a three-dimensional image's speech mouth shape sequence.

2. The highly realistic lip-syncing method for a three-dimensional virtual character according to claim 1, characterized in that, The step of obtaining the lip shape corresponding to each pronunciation from the frontal video recording and recording it as a frontal pronunciation lip shape set specifically includes: Extract the audio from the frontal pronunciation lip shape set and denote it as frontal audio; In the frontal audio, the frontal peak curve is formed by connecting the peak points that reflect the amplitude change trend of the speech signal in the frontal audio. The first frontal curve is obtained by smoothing and normalizing the frontal peak curve. The first positive curve is divided by amplitude to obtain a set of multiple curve segments, which is denoted as the first positive curve segment set. The curve segments in the first positive curve segment set are used as the speech signals corresponding to each pronunciation, and the start and end times of the speech signals are the start and end times corresponding to the pronunciation. Based on the start and end times corresponding to the pronunciation, the frontal video is divided into multiple frontal pronunciation lip shape sets.

3. The highly realistic lip-syncing method for a three-dimensional virtual image according to claim 2, characterized in that, The step of dividing the curve by amplitude to obtain a set of multiple curve segments specifically includes: Obtain the troughs of the curve; Using the point where the trough is located as the cutting point, the curve is divided by amplitude to obtain a set of multiple curve segments.

4. The highly realistic lip-syncing method for a three-dimensional virtual character according to claim 1, characterized in that, The step of obtaining the lip shape corresponding to each pronunciation based on the side video recording and recording it as a side pronunciation lip shape set specifically includes: Extract the audio from the set of lateral pronunciation mouth shapes and denote it as lateral audio; In the side audio, the side peak curve is formed by connecting the peak points that reflect the amplitude change trend of the speech signal in the side audio. The first side curve is obtained by smoothing and normalizing the side peak curve; The first side curve is divided by amplitude to obtain a set of multiple curve segments, which is denoted as the first side curve segment set; The curve segments in the first side curve segment set are used as the speech signals corresponding to each pronunciation, and the start and end times of the speech signals are the start and end times corresponding to the pronunciation. Based on the start and end times corresponding to the pronunciation, the side video is divided into multiple side pronunciation lip shape sets.

5. The highly realistic lip-syncing method for a three-dimensional virtual character according to claim 4, characterized in that, The steps for obtaining the three-dimensional lip movements corresponding to each phoneme from the experimenter specifically include: Acquire laser data corresponding to the laser scanner during the process of the experimenter reading text aloud; A point cloud model is established based on the laser data; The points in the point cloud model include at least the key points corresponding to the mouth in the Dlib facial recognition algorithm. The three-dimensional mouth shape corresponding to each phoneme is obtained based on the point cloud model; The text read aloud by the experimenters was a collection of all phonemes, and the experimenters read each phoneme one by one at intervals.

6. The highly realistic lip-syncing method for a three-dimensional virtual character according to claim 5, characterized in that, The specific steps for obtaining the three-dimensional lip shape corresponding to each phoneme based on the point cloud model include: The key points in the point cloud model are filtered out, and only the key points corresponding to the mouth in the Dlib face recognition algorithm are retained as the mouth shape point cloud model. The point cloud model of the mouth shape corresponding to each phoneme read aloud by the experimenter is obtained as a three-dimensional mouth shape.

7. The highly realistic lip-syncing method for a three-dimensional virtual image according to claim 1, characterized in that, The step of performing three-dimensional processing on each mouth shape in the frontal articulation mouth shape set according to the three-dimensional phoneme mouth shape set to obtain the first articulation mouth shape set specifically includes: For each mouth shape in the aforementioned frontal pronunciation mouth shape set, obtain the corresponding phoneme; The three-dimensional lip shapes corresponding to the three-dimensional lip shape set are obtained according to the corresponding phonemes; The three-dimensional processing of each mouth shape in the frontal articulation mouth shape set is performed using the corresponding three-dimensional mouth shape, specifically including: Obtain the key points corresponding to the mouth in the Dlib facial recognition algorithm for each mouth shape; The 3D coordinates of each key point are adjusted based on the 3D coordinates of the corresponding key points in the 3D lip shape.

8. The highly realistic lip-syncing method for a three-dimensional virtual character according to claim 1, characterized in that, The step of smoothing each lip shape in the first lip shape set according to the lateral lip shape set to obtain the second lip shape set specifically includes: For each mouth shape in the first set of pronunciation mouth shapes, obtain the corresponding phoneme; Obtain the side mouth shape corresponding to the side pronunciation mouth shape set according to the corresponding phoneme; The first set of articulation mouth shapes is processed in three dimensions using the corresponding lateral mouth shapes, specifically including: Obtain the key points of the mouth corresponding to each mouth shape in the Dlib facial recognition algorithm of the first set of pronunciation mouth shapes; The 3D coordinates of each key point are adjusted based on the 3D coordinates of the corresponding key point in the side profile.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores program instructions, which, when executed, are used to perform a highly realistic three-dimensional virtual image lip-syncing method as described in any one of claims 1-8.

10. A highly realistic 3D virtual image lip-syncing system, characterized in that, The highly realistic 3D virtual avatar lip-syncing system includes the computer-readable storage medium as described in claim 9.

Citation Information

Patent Citations

  • A three-dimensional image-based mouth shape simulation method

    CN115690280B

  • Three-dimensional image pronunciation mouth shape simulation method

    CN115690280A

  • Artificial intelligence (AI) character system capable of natural verbal and visual interactions with a human

    US20190095775A1