Implementation method, device and product for simulating speaking of others by digital person

By obtaining the phoneme information and facial feature information of the first and second users, calculating the mixed shape BS coefficients, generating synthetic audio and adjusting facial animation, the problem that digital humans cannot retain speaking characteristics when imitating others' speech is solved, achieving better imitation effects and reducing data collection costs.

CN120707706APending Publication Date: 2025-09-26MIGU CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510814733.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-18
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

In the existing technology, when digital humans imitate others' speech, they are unable to retain the imitator's speech characteristics, resulting in poor imitation effects and high costs.

Method used

By obtaining the phoneme information and facial feature information of the first and second users, calculating the mixed shape BS coefficients, generating synthetic audio and adjusting facial animation, the digital human can imitate the speaking movements of the second user.

Benefits of technology

The facial movement characteristics of the first user are retained, the imitation effect is better, and the data collection cost is reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120707706A_ABST
    Figure CN120707706A_ABST
Patent Text Reader

Abstract

The invention provides an implementation method, device and product for a digital person to imitate speaking of others, and the method comprises the steps: obtaining first phoneme information of a target text read by a first user and second phoneme information of the target text read by a second user; acquiring facial feature information when the first user reads each phoneme; acquiring a first mixed shape BS coefficient according to the first phoneme information, the second phoneme information and the facial feature information; obtaining a synthetic audio according to the first phoneme information and the second phoneme information; wherein the synthesized audio is the audio that the first user simulates the second user to read the target text; and according to the synthesized audio and the first BS coefficient, generating the facial animation that the digital person corresponding to the first user simulates the second user, so that the facial action characteristics of the first user during speaking can be reserved, the simulation effect is better, additional data collection is not needed, and the data collection cost can be reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of communication technology, and in particular to a method, device and product for realizing digital humans imitating other people's speech. Background Art

[0002] In existing technology, when a digital human of user A imitates user B's speech, the human's mouth movements differ from those of a normal human speaking. Current digital human speech imitation solutions typically collect data on different imitation styles and use this data as a condition to drive the human. However, because the audio properties of each person's speech vary, the human can only be roughly categorized based on the size of the differences. This results in a trained digital human that can only roughly imitate user B's behavior, resulting in poor imitation results, high costs, and a failure to retain the unique movements of user A (digital human A). Summary of the Invention

[0003] The purpose of the present invention is to provide a method, device and product for realizing digital human imitating the speech of others, so as to solve the technical problem in the prior art that when digital human imitates the speech of other users, it cannot retain the speech characteristics of the imitator himself.

[0004] To achieve the above-mentioned purpose, an embodiment of the present invention provides a method for a digital human to imitate another person's speech, which includes:

[0005] Acquire first phoneme information of a first user reading a target text and second phoneme information of a second user reading the target text;

[0006] Obtaining facial feature information of the first user when reading each phoneme;

[0007] acquiring a first mixed shape BS coefficient according to the first phoneme information, the second phoneme information, and the facial feature information;

[0008] Acquire synthesized audio according to the first phoneme information and the second phoneme information; wherein the synthesized audio is audio of the first user imitating the second user reading the target text;

[0009] A facial animation of a digital human corresponding to the first user imitating the second user is generated according to the synthesized audio and the first BS coefficient.

[0010] Optionally, the method, wherein obtaining first phoneme information of a first user reading a target text and second phoneme information of a second user reading the target text, includes:

[0011] Acquire first voice information of the first user, second voice information of the second user, and the target text;

[0012] Extracting speech features from the first speech information and the second speech information respectively to obtain a first speech latent vector and a second speech latent vector;

[0013] The first phoneme information is obtained based on the target text and the first speech latent vector, and the second phoneme information is obtained based on the target text and the second speech latent vector.

[0014] Optionally, the method, wherein obtaining facial feature information of the first user when reading each phoneme, includes:

[0015] Obtaining a facial video of the first user recorded by a camera corresponding to the pronunciation of multiple phonemes;

[0016] Segmenting the facial video into facial video segments corresponding to each of a plurality of phonemes, and obtaining a plurality of facial video segments;

[0017] Obtaining corresponding facial key points according to the facial video clip;

[0018] The facial blend shape BS coefficients are extracted from the facial key points to obtain the facial feature information; wherein the facial feature information includes the facial BS coefficients corresponding to each of a plurality of phonemes.

[0019] Optionally, the method, wherein obtaining a first mixed shape BS coefficient according to the first phoneme information of the first phoneme audio, the second phoneme information of the second phoneme audio, and the facial feature information, includes:

[0020] obtaining preliminary mixed shape BS coefficients according to the first phoneme information and the facial feature information;

[0021] acquiring mouth difference information according to the first phoneme information and the second phoneme information;

[0022] acquiring rhythm difference information according to a first pronunciation duration acquired through the first phoneme information and a second pronunciation duration acquired through the second phoneme information;

[0023] The first BS coefficient is acquired according to the mouth difference information, the rhythm difference information and the preliminary BS coefficient.

[0024] Optionally, the method, wherein obtaining mouth difference information according to the first phoneme information and the second phoneme information, includes:

[0025] Obtaining a jaw difference according to a first pitch and a second pitch; wherein the first phoneme information includes the first pitch, the second phoneme information includes the second pitch; and the mouth difference information includes the jaw difference;

[0026] Obtain lip difference according to the first pitch, the second pitch, the first loudness, and the second loudness; wherein the first phoneme information includes the first loudness, the second phoneme information includes the second loudness; and the mouth difference information includes the lip difference.

[0027] Optionally, the method, wherein obtaining synthesized audio according to the first phoneme information and the second phoneme information, includes:

[0028] Acquire a second pronunciation duration according to the second phoneme information;

[0029] The synthesized audio is obtained according to the second pronunciation duration and the first phoneme information.

[0030] In order to achieve the above-mentioned purpose, an embodiment of the present invention further provides a device for realizing a digital human imitating another person's speech, which includes:

[0031] A first acquisition module is configured to acquire first phoneme information of a first user reading a target text and second phoneme information of a second user reading the target text;

[0032] A second acquisition module is used to acquire facial feature information of the first user when reading each phoneme;

[0033] a third acquisition module, configured to acquire a first mixed shape BS coefficient according to the first phoneme information, the second phoneme information, and the facial feature information;

[0034] a fourth acquisition module, configured to acquire a synthesized audio according to the first phoneme information and the second phoneme information; wherein the synthesized audio is an audio of the first user imitating the second user reading the target text;

[0035] The first generating module is configured to generate a facial animation of a digital human corresponding to the first user imitating the second user based on the synthesized audio and the first BS coefficient.

[0036] In order to achieve the above-mentioned purpose, an embodiment of the present invention further provides an electronic device, comprising: a transceiver, a processor, a memory, and a program or instruction stored in the memory and executable on the processor; wherein, when the processor executes the program or instruction, the method for realizing a digital human imitating another person's speech as described above is implemented.

[0037] In order to achieve the above-mentioned purpose, an embodiment of the present invention further provides a readable storage medium on which a program or instruction is stored, wherein when the program or instruction is executed by a processor, the steps in the above-mentioned method for realizing a digital human imitating the speech of another person are implemented.

[0038] In order to achieve the above-mentioned purpose, an embodiment of the present invention further provides a computer program product, which includes computer instructions. When the computer instructions are executed by a processor, the steps of the above-mentioned method for implementing a digital human imitating the speech of another person are implemented.

[0039] The beneficial effects of the above technical solution of the present invention are as follows:

[0040] In an embodiment of the present invention, a first mixed shape BS coefficient is obtained based on the first phoneme information of the first user reading the target text, the second phoneme information of the second user reading the target text, and the facial feature information of the first user when reading each phoneme, and a synthesized audio is obtained based on the first phoneme information and the second phoneme information, thereby generating a facial animation of a digital human corresponding to the first user imitating the second user. This can retain the facial movement characteristics of the first user when speaking, achieve a better imitation effect, and does not require additional data collection, thereby reducing data collection costs. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] Figure 1 A schematic diagram of a method for implementing a digital human imitating another person's speech according to an embodiment of the present invention;

[0042] Figure 2 This is a schematic diagram of a device for implementing a digital human imitating another person's speech according to an embodiment of the present invention. DETAILED DESCRIPTION

[0043] In order to make the technical problems, technical solutions and advantages to be solved by the present invention clearer, a detailed description will be given below with reference to the accompanying drawings and specific embodiments.

[0044] It should be understood that references throughout this specification to "one embodiment" or "an embodiment" mean that a particular feature, structure, or characteristic associated with the embodiment is included in at least one embodiment of the present invention. Therefore, the appearances of "in one embodiment" or "in an embodiment" throughout this specification do not necessarily refer to the same embodiment. Furthermore, these particular features, structures, or characteristics may be combined in any suitable manner in one or more embodiments.

[0045] In various embodiments of the present invention, it should be understood that the size of the serial numbers of the following processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0046] Additionally, the terms "system" and "network" are often used interchangeably herein.

[0047] In the embodiments provided herein, it should be understood that "B corresponding to A" means that B is associated with A and B can be determined based on A. However, it should also be understood that determining B based on A does not mean determining B based solely on A; B can also be determined based on A and / or other information.

[0048] For ease of understanding, some contents involved in the embodiments of the present invention are described below:

[0049] like Figure 1 As shown, a method for implementing a digital human to imitate another person's speech according to an embodiment of the present invention includes:

[0050] S10, obtaining first phoneme information of a first user reading a target text and second phoneme information of a second user reading the target text;

[0051] It should be noted that the target text is converted into the corresponding first phoneme information and second phoneme information based on the voice information of the first user and the second user. The first phoneme information and the second phoneme information include the phonemes of the target text, as well as information such as the pronunciation duration, pitch, and loudness of each phoneme. A phoneme is the smallest phonetic unit divided according to the natural properties of speech, such as the initials and finals in Chinese.

[0052] S20, obtaining facial feature information of the first user when reading each phoneme;

[0053] It should be noted that the facial animation of the first user saying each phoneme is unique, and the BS (Blend Shape) coefficient corresponding to each phoneme is calculated through lip animation to obtain the facial feature information. The BS coefficient is a coefficient used to describe the state of different areas of the face (mouth, nose, left and right eyes, etc.). Any facial action can be decomposed into multiple facial states, that is, decomposed into multiple groups of BS coefficients (for example, a video can be decomposed into multiple pictures). Since each person's facial movements may be slightly different when pronouncing a word. In order to capture this difference, it is necessary to collect the facial feature information of the first user. The facial feature information reflects the unique phoneme-level BS coefficient when the first user pronounces each phoneme.

[0054] S30, acquiring a first mixed shape BS coefficient according to the first phoneme information, the second phoneme information and the facial feature information;

[0055] It should be noted that, by combining the first phoneme information with the facial feature information, the target text can be converted into the facial movements unique to the first user when reading the target text. Combined with the second phoneme information, a BS coefficient can be obtained based on the unique facial movements of the first user and the facial pronunciation differences between the first and second users, that is, the BS coefficient.

[0056] S40, obtaining a synthesized audio according to the first phoneme information and the second phoneme information; wherein the synthesized audio is an audio of the first user imitating the second user reading the target text;

[0057] It should be noted that the phoneme pronunciation duration of the second user in the second phoneme information is used to replace the phoneme pronunciation duration in the first phoneme information to synthesize the synthesized audio of the first user imitating the second user reading the target text.

[0058] S50, based on the synthesized audio and the first BS coefficient;

[0059] It should be noted that, by combining the synthesized audio and the first BS coefficient, a facial animation is generated in which the face of the digital human corresponding to the first user imitates the pronunciation of the second user.

[0060] In this embodiment, the first mixed shape BS coefficient is obtained based on the first phoneme information of the first user reading the target text, the second phoneme information of the second user reading the target text, and the facial feature information when the first user reads each phoneme, and the synthesized audio is obtained based on the first phoneme information and the second phoneme information, and then the facial animation of the digital human corresponding to the first user imitating the second user is generated. This can retain the facial movement characteristics of the first user when speaking, has a better imitation effect, does not require additional data collection, and can reduce data collection costs.

[0061] In the real world, it's common for one user to imitate another user's speech. For the same sentence, the first user's mouth movements differ from their normal speech. This is similar to how a girl imitating a boy might lower her voice, creating a different mouth position than normal. For two users with different voices, when applying this imitation behavior to a digital human, it's desirable to replicate the movements of the first user as closely as possible, making the digital human more realistic. This embodiment of the present invention leverages the differences in the two users' voices to automatically adjust the mouth amplitude and restore the imitator's mouth position.

[0062] Optionally, the method, wherein obtaining first phoneme information of a first user reading a target text and second phoneme information of a second user reading the target text, includes:

[0063] Acquire first voice information of the first user, second voice information of the second user, and the target text;

[0064] Extracting speech features from the first speech information and the second speech information respectively to obtain a first speech latent vector and a second speech latent vector;

[0065] The first phoneme information is obtained based on the target text and the first speech latent vector, and the second phoneme information is obtained based on the target text and the second speech latent vector.

[0066] In this embodiment, a pre-trained speech synthesis model is used to extract speech features from the first and second speech information using its speech feature extraction module, including speech rhythm, pitch, loudness, etc., and represent them in the form of latent vectors to obtain the first and second speech latent vectors. The target text is converted into the first and second phoneme information using the first and second latent vectors, respectively. The phoneme information includes the phonemes corresponding to the target text, as well as the pronunciation duration, pitch, and loudness information of each phoneme.

[0067] Optionally, the method, wherein obtaining facial feature information of the first user when reading each phoneme, includes:

[0068] Obtaining a facial video of the first user recorded by a camera corresponding to the pronunciation of multiple phonemes;

[0069] Segmenting the facial video into facial video segments corresponding to each of a plurality of phonemes, and obtaining a plurality of facial video segments;

[0070] Obtaining corresponding facial key points according to the facial video clip;

[0071] The facial blend shape BS coefficients are extracted from the facial key points to obtain the facial feature information; wherein the facial feature information includes the facial BS coefficients corresponding to each of a plurality of phonemes.

[0072] In this embodiment, a facial video of the first user pronouncing 27 phonemes is obtained by recording the facial video through a camera. The interval between each phoneme is about 1 second, the lips are kept closed during the interval, and the recording takes about 1 minute. The facial video is then uploaded to a server, and the server background segments the video according to the interval. Each phoneme corresponds to a facial video segment, and then the facial key points of each facial video segment are calculated. The facial BS coefficient is extracted from the facial key points, so that each phoneme corresponds to a BS coefficient containing the facial features of the first user. Any sentence of speech can be composed of these 27 phonemes. Therefore, these 27 phoneme-BS coefficient pairs constitute the facial feature information unique to the first user.

[0073] Optionally, the method, wherein obtaining a first mixed shape BS coefficient according to the first phoneme information of the first phoneme audio, the second phoneme information of the second phoneme audio, and the facial feature information, includes:

[0074] obtaining preliminary mixed shape BS coefficients according to the first phoneme information and the facial feature information;

[0075] acquiring mouth difference information according to the first phoneme information and the second phoneme information;

[0076] acquiring rhythm difference information according to a first pronunciation duration acquired through the first phoneme information and a second pronunciation duration acquired through the second phoneme information;

[0077] The first BS coefficient is acquired according to the mouth difference information, the rhythm difference information and the preliminary BS coefficient.

[0078] In this embodiment, the target text content can be converted into facial movements unique to the first user by combining the first phoneme information and the facial feature information. The facial movements are also represented by BS coefficients to obtain the preliminary BS coefficients. The preliminary BS coefficients describe the facial movements of the first user's digital human when narrating the target text according to the first user's voice characteristics. When the first user's digital human wants to imitate the second user (according to the second user's voice characteristics), the preliminary BS coefficients can be modified based on the audio difference information. Based on the first phoneme information and the second phoneme information, the mouth amplitude and rhythm of the first user's digital human are adjusted, and combined with the preliminary BS coefficients to obtain the first BS coefficients.

[0079] Because the duration of each phoneme pronunciation by the first and second users may differ, the rhythm of the first user's digital human needs to be adjusted based on the second user's rhythm. The rhythm difference information is the interpolation ratio calculated by dividing the second pronunciation duration by the first pronunciation duration. A phoneme interpolation ratio of 1 indicates no adjustment is required; a ratio greater than 1 indicates that frames need to be added to the facial video clip corresponding to the phoneme; a ratio less than 1 indicates that frames need to be deleted.

[0080] The coefficients corresponding to the chin and lips in the preliminary BS coefficients are adjusted according to the mouth difference information to obtain the temporary BS coefficients. The formula is as follows:

[0081] Temporary BS coefficient = min(|preliminary BS coefficient + chin difference + lip difference|, I),

[0082] Here, I represents an all-1 vector having the same shape as the preliminary BS coefficients.

[0083] The temporary BS coefficient is interpolated and adjusted according to the rhythm difference information, and the interpolated BS coefficient is recorded as the first BS coefficient.

[0084] Optionally, the method, wherein obtaining mouth difference information according to the first phoneme information and the second phoneme information, includes:

[0085] Obtaining a jaw difference according to a first pitch and a second pitch; wherein the first phoneme information includes the first pitch, the second phoneme information includes the second pitch; and the mouth difference information includes the jaw difference;

[0086] Obtain lip difference according to the first pitch, the second pitch, the first loudness, and the second loudness; wherein the first phoneme information includes the first loudness, the second phoneme information includes the second loudness; and the mouth difference information includes the lip difference.

[0087] In this embodiment, the mouth difference information includes the chin difference and the lip difference, wherein the lip difference is used to adjust the degree of lip opening and closing of the phoneme in the action; and the chin difference is used to adjust the degree of chin opening and closing of the phoneme in the action. The value range of the chin difference and the lip difference is between -1 and 1, where 0 indicates no adjustment, less than 0 indicates inhibition, and greater than 0 indicates stimulation. Since phonemes are divided into three categories according to fricatives, plosives, and vowels, only the lip amplitude of fricatives and plosives is affected by loudness and pitch; the lip amplitude and chin amplitude of vowels are both affected by loudness and pitch. Therefore, the calculation method of the chin difference and the lip difference is as follows:

[0088] Chin difference = L B -LA ;

[0089] Lip difference = L B / P B -L A / P A .

[0090] L represents the pitch of a phoneme, and P represents its loudness. Loudness is the square root of the mean of the squared audio sample values, followed by its logarithm. Pitch is the inverse of the fundamental frequency period. Real-time adjustments to the lip-sync amplitude allow for finer granularity.

[0091] Optionally, the method, wherein obtaining synthesized audio according to the first phoneme information and the second phoneme information, includes:

[0092] Acquire a second pronunciation duration according to the second phoneme information;

[0093] The synthesized audio is obtained according to the second pronunciation duration and the first phoneme information.

[0094] In this embodiment, the phoneme pronunciation duration of the second user in the second phoneme information is used to replace the phoneme pronunciation duration in the first phoneme information to synthesize the synthesized audio of the first user imitating the second user reading the target text.

[0095] like Figure 2 To achieve the above-mentioned purpose, an embodiment of the present invention further provides a device for enabling a digital human to imitate another person's speech, which includes:

[0096] A first acquisition module 201 is configured to acquire first phoneme information of a first user reading a target text and second phoneme information of a second user reading the target text;

[0097] A second acquisition module 202 is configured to acquire facial feature information of the first user when reading each phoneme;

[0098] A third acquisition module 203 is configured to acquire a first mixed shape BS coefficient according to the first phoneme information, the second phoneme information, and the facial feature information;

[0099] The fourth acquisition module 204 is configured to acquire a synthesized audio according to the first phoneme information and the second phoneme information; wherein the synthesized audio is an audio of the first user imitating the second user reading the target text;

[0100] The first generating module 205 is configured to generate a facial animation of a digital human corresponding to the first user imitating the second user based on the synthesized audio and the first BS coefficient.

[0101] Optionally, in the device, the first obtaining module 201 includes:

[0102] a first acquiring unit, configured to acquire first voice information of the first user, second voice information of the second user, and the target text;

[0103] a second acquiring unit, configured to extract speech features from the first speech information and the second speech information respectively, to acquire a first speech latent vector and a second speech latent vector;

[0104] A third acquisition unit is used to acquire the first phoneme information based on the target text and the first speech latent vector, and to acquire the second phoneme information based on the target text and the second speech latent vector.

[0105] Optionally, in the device, the second obtaining module 202 includes:

[0106] a fourth acquiring unit, configured to acquire a facial video corresponding to a plurality of phoneme pronunciations recorded by the first user through a camera;

[0107] a fifth acquiring unit, configured to segment the facial video into facial video segments corresponding to each of the plurality of phonemes, and acquire a plurality of facial video segments;

[0108] a sixth acquiring unit, configured to acquire corresponding facial key points according to the facial video clip;

[0109] A seventh acquisition unit is configured to extract facial blend shape BS coefficients from the facial key points to acquire the facial feature information; wherein the facial feature information includes facial BS coefficients corresponding to each of a plurality of phonemes.

[0110] Optionally, in the device, the third obtaining module 203 includes:

[0111] an eighth acquiring unit, configured to acquire preliminary mixed shape BS coefficients according to the first phoneme information and the facial feature information;

[0112] a ninth acquiring unit, configured to acquire mouth difference information according to the first phoneme information and the second phoneme information;

[0113] a tenth acquiring unit, configured to acquire rhythm difference information according to the first phoneme information and the second phoneme information;

[0114] An eleventh acquiring unit is configured to acquire the first BS coefficient according to the mouth difference information, the rhythm difference information, and a preliminary BS coefficient.

[0115] Optionally, in the device, the ninth obtaining unit includes:

[0116] A first acquisition component is configured to acquire a jaw difference according to a first pitch and a second pitch; wherein the first phoneme information includes the first pitch, the second phoneme information includes the second pitch, and the mouth difference information includes the jaw difference;

[0117] The second acquisition component is used to obtain lip difference based on the first pitch, the second pitch, the first loudness and the second loudness; wherein the first phoneme information includes the first loudness, the second phoneme information includes the second loudness; and the mouth difference information includes the lip difference.

[0118] Optionally, in the device, the tenth obtaining unit includes:

[0119] A third acquisition component is used to acquire a first pronunciation duration according to the first phoneme information;

[0120] a fourth acquisition component, configured to acquire a second pronunciation duration according to the second phoneme information;

[0121] A fifth acquisition component is configured to acquire the rhythm difference information according to the first pronunciation duration and the second pronunciation duration.

[0122] Optionally, in the device, the eleventh obtaining unit includes:

[0123] a sixth obtaining component, configured to obtain a temporary blend shape BS coefficient based on the mouth difference information and the preliminary BS coefficient;

[0124] A seventh acquisition component is configured to acquire the first BS coefficient according to the temporary BS coefficient and the rhythm difference information.

[0125] Optionally, in the device, the fourth obtaining module 204 includes:

[0126] a twelfth acquiring unit, configured to acquire a second pronunciation duration according to the second phoneme information;

[0127] A thirteenth acquisition unit is used to acquire the synthesized audio according to the second pronunciation duration and the first phoneme information.

[0128] It should be noted here that the above-mentioned device provided by the embodiment of the present invention can implement all the method steps implemented by the above-mentioned method embodiment and can achieve the same technical effect. The parts and beneficial effects of this embodiment that are the same as those of the method embodiment will not be described in detail here.

[0129] In order to achieve the above-mentioned purpose, an embodiment of the present invention further provides an electronic device, comprising: a transceiver, a processor, a memory, and a program or instruction stored in the memory and executable on the processor; wherein, when the processor executes the program or instruction, the method for realizing a digital human imitating another person's speech as described above is implemented.

[0130] In order to achieve the above-mentioned purpose, an embodiment of the present invention further provides a readable storage medium on which a program or instruction is stored, wherein when the program or instruction is executed by a processor, the steps in the above-mentioned method for realizing a digital human imitating the speech of another person are implemented.

[0131] In order to achieve the above-mentioned purpose, an embodiment of the present invention further provides a computer program product, which includes computer instructions. When the computer instructions are executed by a processor, the steps of the above-mentioned method for implementing a digital human imitating the speech of another person are implemented.

[0132] It should be further noted that the terminals described in this specification include but are not limited to smartphones, tablet computers, etc., and many functional components described are referred to as modules in order to more particularly emphasize the independence of their implementation methods.

[0133] In embodiments of the present invention, modules can be implemented in software so that they can be executed by various types of processors. For example, an identified executable code module can include one or more physical or logical blocks of computer instructions, for example, which can be constructed as objects, procedures, or functions. Nevertheless, the executable code of the identified module does not need to be physically located together, but can include different instructions stored in different locations, which, when logically combined together, constitute the module and achieve the specified purpose of the module.

[0134] In fact, executable code module can be a single instruction or many instructions, and can even be distributed on a plurality of different code segments, distributed in the middle of different programs, and distributed across a plurality of memory devices.Similarly, operating data can be identified in the module, and can be implemented and organized in the data structure of any appropriate type according to any appropriate form.Described operating data can be collected as a single data set, or can be distributed in different locations (including on different storage devices), and can only be present on a system or network as an electronic signal at least in part.

[0135] When a module can be implemented using software, given the current state of hardware technology, those skilled in the art can build corresponding hardware circuits to implement the corresponding functions of the module, regardless of cost. The hardware circuits may include conventional very large scale integration (VLSI) circuits or gate arrays, as well as existing semiconductors such as logic chips and transistors, or other discrete components. Modules may also be implemented using programmable hardware devices, such as field programmable gate arrays, programmable array logic, or programmable logic devices.

[0136] The above exemplary embodiments are described with reference to the accompanying drawings. Many different forms and embodiments are possible without departing from the spirit and teachings of the present invention. Therefore, the present invention should not be construed as limited to the exemplary embodiments set forth herein. Rather, these exemplary embodiments are provided so that this disclosure will be complete and perfect and will convey the scope of the invention to those skilled in the art. In the drawings, component sizes and relative sizes may be exaggerated for clarity. The terminology used herein is for purposes of describing specific exemplary embodiments only and is not intended to be limiting. As used herein, the singular forms "a," "an," and "the" are intended to include plural forms, unless the context clearly indicates otherwise. It will be further understood that the terms "comprising" and / or "including," when used in this specification, indicate the presence of stated features, integers, steps, operations, components, and / or elements, but do not preclude the presence or addition of one or more other features, integers, steps, operations, components, elements, and / or groups thereof. Unless otherwise indicated, when stated, a range of values ​​includes the upper and lower limits of that range and any subranges therebetween.

[0137] The above is a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications should also be regarded as within the scope of protection of the present invention.

Claims

1. A method for implementing a digital human to imitate another person's speech, characterized in that: include: Acquire first phoneme information of a first user reading a target text and second phoneme information of a second user reading the target text; Obtaining facial feature information of the first user when reading each phoneme; acquiring a first mixed shape BS coefficient according to the first phoneme information, the second phoneme information, and the facial feature information; Acquire synthesized audio according to the first phoneme information and the second phoneme information; wherein the synthesized audio is audio of the first user imitating the second user reading the target text; A facial animation of a digital human corresponding to the first user imitating the second user is generated according to the synthesized audio and the first BS coefficient.

2. The method according to claim 1, characterized in that Acquiring first phoneme information of a target text read by a first user and second phoneme information of the target text read by a second user includes: Acquire first voice information of the first user, second voice information of the second user, and the target text; Extracting speech features from the first speech information and the second speech information respectively to obtain a first speech latent vector and a second speech latent vector; The first phoneme information is obtained based on the target text and the first speech latent vector, and the second phoneme information is obtained based on the target text and the second speech latent vector.

3. The method according to claim 1, characterized in that Obtaining facial feature information of the first user when reading each phoneme, including: Obtaining a facial video of the first user recorded by a camera corresponding to the pronunciation of multiple phonemes; Segmenting the facial video into facial video segments corresponding to each of a plurality of phonemes, and obtaining a plurality of facial video segments; Obtaining corresponding facial key points according to the facial video clip; The facial blend shape BS coefficients are extracted from the facial key points to obtain the facial feature information; wherein the facial feature information includes the facial BS coefficients corresponding to each of a plurality of phonemes.

4. The method according to claim 1, wherein Acquiring a first mixed shape BS coefficient according to first phoneme information of the first phoneme audio, second phoneme information of the second phoneme audio, and the facial feature information, including: obtaining preliminary mixed shape BS coefficients according to the first phoneme information and the facial feature information; acquiring mouth difference information according to the first phoneme information and the second phoneme information; acquiring rhythm difference information according to a first pronunciation duration acquired through the first phoneme information and a second pronunciation duration acquired through the second phoneme information; The first BS coefficient is acquired according to the mouth difference information, the rhythm difference information and the preliminary BS coefficient.

5. The method according to claim 4, characterized in that Acquiring mouth difference information according to the first phoneme information and the second phoneme information includes: Obtaining a jaw difference according to a first pitch and a second pitch; wherein the first phoneme information includes the first pitch, the second phoneme information includes the second pitch; and the mouth difference information includes the jaw difference; Obtain lip difference according to the first pitch, the second pitch, the first loudness, and the second loudness; wherein the first phoneme information includes the first loudness, the second phoneme information includes the second loudness; and the mouth difference information includes the lip difference.

6. The method according to claim 1, characterized in that Acquiring synthesized audio according to the first phoneme information and the second phoneme information includes: Acquire a second pronunciation duration according to the second phoneme information; The synthesized audio is obtained according to the second pronunciation duration and the first phoneme information.

7. A device for realizing a digital human imitating another person's speech, characterized in that: include: A first acquisition module is configured to acquire first phoneme information of a first user reading a target text and second phoneme information of a second user reading the target text; A second acquisition module is used to acquire facial feature information of the first user when reading each phoneme; a third acquisition module, configured to acquire a first mixed shape BS coefficient according to the first phoneme information, the second phoneme information, and the facial feature information; a fourth acquisition module, configured to acquire a synthesized audio according to the first phoneme information and the second phoneme information; wherein the synthesized audio is an audio of the first user imitating the second user reading the target text; The first generating module is configured to generate a facial animation of a digital human corresponding to the first user imitating the second user based on the synthesized audio and the first BS coefficient.

8. An electronic device comprising: A transceiver, a processor, a memory, and a program or instruction stored in the memory and executable on the processor; wherein the processor implements the method for realizing a digital human imitating another person's speech as described in any one of claims 1 to 6 when executing the program or instruction.

9. A readable storage medium having a program or instruction stored thereon, characterized in that: When the program or instruction is executed by a processor, the steps of the method for realizing a digital human imitating another person's speech as described in any one of claims 1 to 6 are implemented.

10. A computer program product, characterized in that The method comprises computer instructions, which, when executed by a processor, implement the steps of the method for realizing a digital human imitating the speech of another person as described in any one of claims 1 to 6.