Method and device for synchronizing voice and mouth shape of digital human and electronic equipment

By obtaining natural language information and converting it into text information, determining video frames and audio frames, and adding timestamps to audio frames and video frames, the problem of voice and lip shapes is solved in digital human products, and the synchronization of voice and lip shapes is achieved.

CN120238712APending Publication Date: 2025-07-01WUHAN AOTUO INTELLIGENT TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202311845546.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-28
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

Voice and lip shapes are often processed asynchronously in digital human products, resulting in the problem of dissynchronization of voice and lip shapes.

Method used

By obtaining natural language information, converting it into text information, determining video frames and audio frames, and adding timestamps on audio frames and video frames, ensuring that the time of the audio frames and video frames is consistent, the digital person is controlled to speak.

Benefits of technology

The synchronization of the voice and lip shape of digital people is achieved to ensure that the voice and lip shape are consistent.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120238712A_ABST
    Figure CN120238712A_ABST
Patent Text Reader

Abstract

The invention relates to the field of digital people, and discloses a method and a device for synchronizing voice and mouth shape of a digital person, and electronic equipment. The method comprises the steps of obtaining natural language information, converting the natural language information into character information, determining a video frame and an audio frame based on the character information, and controlling the digital person to speak when the first frame time of the audio frame is the same as the first frame time of the video frame, so that the voice and mouth shape of the digital person can be kept consistent.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of digital humans, and particularly to a method, device, and electronic device for synchronizing digital human speech and lip movements. Background Art

[0002] With the development of cloud resource construction and streaming technology, cloud rendering technology has gradually emerged and extended from the design field to the immersive digital human application field.

[0003] Since digital humans simulate the behaviors and actions of real people, typical voices and lip movements are the basis for understanding natural language and imitating real people's behaviors. Due to the asynchronous processing of voice and lip movement simulation, digital human products often have problems with speech and lip movement asynchronization. Summary of the Invention

[0004] Based on this, in order to solve the above technical problems, it is necessary to provide a method, device, and electronic device for synchronizing digital human speech and lip movements, which can keep the speech and lip movements of digital humans consistent.

[0005] In a first aspect, an embodiment of this application provides a method for synchronizing digital human speech and lip movements, the method comprising:

[0006] Obtain natural language information;

[0007] Convert the natural language information into text information;

[0008] Based on the text information, determine video frames and audio frames;

[0009] When the first frame of the audio frame is at the same time as the first frame of the video frame, control the digital human to speak.

[0010] In some embodiments, the determining video frames and audio frames based on the text information includes:

[0011] Based on the text information, determine the text-to-audio delay time, text-to-lip movement delay time, speech rate, lip movement for speech, audio frames, and video frames.

[0012] In some embodiments, after determining the video frames and audio frames based on the text information, the method further comprises:

[0013] Add timestamps to the audio frames and the video frames respectively.

[0014] In some embodiments, the adding a timestamp to the audio frame includes:

[0015] Determine the time difference between the video and the video generation according to the text-to-audio delay time and the text-to-lip movement delay time;

[0016] Determine the audio frame step size according to the audio frame, the speech rate, and the lip movement rate.

[0017] Starting from 0, calculate the relative timestamp of each audio frame according to the audio frame step size.

[0018] Calculate the absolute timestamp of each audio frame according to the relative timestamp of each audio frame, the current time, the text-to-audio delay time, and the text-to-lip movement delay time.

[0019] In a second aspect, an embodiment of the present application further provides a digital human speech and lip movement synchronization device, and the device includes:

[0020] An acquisition module, configured to acquire natural language information;

[0021] A conversion module, configured to convert the natural language information into text information;

[0022] A processing module, configured to determine video frames and audio frames based on the text information;

[0023] A control module, configured to control the digital human to speak when the first frames of the audio frames and the video frames are at the same time.

[0024] In some embodiments, the processing module is specifically configured to:

[0025] Based on the text information, determine the text-to-audio delay time, the text-to-lip movement delay time, the speech rate, the lip movement rate, the audio frames, and the video frames.

[0026] In some embodiments, the device further includes:

[0027] An adding module, configured to add timestamps to the audio frames and the video frames respectively.

[0028] In some embodiments, the adding module is specifically configured to:

[0029] Determine the time difference between the video and the video generation according to the text-to-audio delay time and the text-to-lip movement delay time;

[0030] Determine the audio frame step size according to the audio frame, the speech rate, and the lip movement rate;

[0031] Starting from 0, calculate the relative timestamp of each audio frame according to the audio frame step size;

[0032] Calculate the absolute timestamp of each audio frame according to the relative timestamp of each audio frame, the current time, the text-to-audio delay time, and the text-to-lip movement delay time.

[0033] In a third aspect, an embodiment of the present application further provides an electronic device, including:

[0034] at least one processor; and,

[0035] a memory communicatively connected to the at least one processor; wherein,

[0036] the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the method described in the first aspect above.

[0037] In a fourth aspect, an embodiment of the present application further provides a non-volatile computer-readable storage medium, which stores computer-executable instructions, and when the computer-executable instructions are executed by a processor, the processor is enabled to execute the method described in the first aspect above.

[0038] Compared with the prior art, the beneficial effects of the present application are as follows: Different from the prior art, the method for synchronizing the voice and lip movement of a digital human provided by the embodiment of the present application obtains natural language information, then converts the natural language information into text information, and then determines video frames and audio frames based on the text information. When the first frames of the audio frames and the video frames are at the same time, the digital human is controlled to speak, so as to ensure that the voice and lip movement of the digital human are consistent. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] One or more embodiments are illustrated by way of example in the accompanying drawings, and these illustrative descriptions do not limit the embodiments. Elements with the same reference numerals in the drawings are represented as similar elements, unless otherwise stated, and the drawings in the figures do not constitute a scale limitation.

[0040] Figure 1 is a schematic flowchart of a method for synchronizing the voice and lip movement of a digital human provided by an embodiment of the present application;

[0041] Figure 2 is a schematic structural diagram of a device for synchronizing the voice and lip movement of a digital human provided by an embodiment of the present application;

[0042] Figure 3 is a schematic hardware structure diagram of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0043] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of this application. Apparently, the described embodiments are some, but not all, of the embodiments of this application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in this application without creative efforts shall fall within the protection scope of this application.

[0044] It should be noted that if there is no conflict, the various features in the embodiments of this application can be combined with each other and are all within the protection scope of this application. In addition, although functional module division is performed in the device schematic diagram and the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order from the module division in the device or the order in the flowchart. Furthermore, the terms "first", "second", "third", etc. adopted in this application do not limit the data and execution order, but are only used to distinguish identical items or similar items with basically the same functions and effects.

[0045] As Figure 1 shown, the embodiments of this application provide a method for synchronizing the speech and lip movements of a digital human. The method includes:

[0046] Step 102, obtain natural language information.

[0047] Step 104, convert the natural language information into text information.

[0048] In the embodiments of this application, natural language information is first obtained, and then the obtained natural language information is converted into text information by means of machine translation or in the form of intelligent response text.

[0049] Step 106, determine video frames and audio frames based on the text information.

[0050] Based on the text information, determine the text-to-audio delay time, text-to-lip movement delay time, speech speed, lip movement speed, audio frames, and video frames. Exemplarily, in the embodiments of this application, the text-to-audio delay time is represented by T1, the text-to-lip movement delay time is represented by T2, the speech speed is represented by S1, the lip movement speed is represented by S2, the audio frames are represented by F1, and the video frames are represented by F2. Among them, the audio frame F1 is the frame rate of the audio, and the video frame F2 is the frame rate of the video.

[0051] In some embodiments, after determining the video frames and audio frames based on the text information, the method further includes: adding timestamps to the audio frames and the video frames respectively.

[0052] Specifically, first, add a timestamp t1 to the video frame, and let t1 = tc, where tc represents the current time. Then add a timestamp t2 to the audio frame. Then calculate the time difference between the generation of the audio and the video, subtract the text-to-audio delay time from the text-to-lip delay time (T1 - T2), and the time difference between the generation of the audio and the video can be obtained. Next, determine the audio frame step size according to the audio frame F1, the speech rate S1, and the lip movement rate S2. The specific calculation formula is: (1 / F1) * (S1 / S2). After determining the audio frame step size, starting from 0, calculate the relative timestamp of each frame of audio according to the audio frame step size. The specific calculation formula is: N * (1 / F1) * (S1 / S2), where N is a natural number. Finally, calculate the absolute timestamp of each frame of audio according to the relative timestamp of each frame of audio, the current time, the text-to-audio delay time, and the text-to-lip delay time. The specific calculation formula is: t2 = tc - (T1 - T2) + N * (1 / F1) * (S1 / S2).

[0053] Step 108, when the first frame of the audio frame is at the same time as the first frame of the video frame, then control the digital human to speak.

[0054] After determining the absolute timestamp of each frame of audio, buffer 100S of data. When the first frame of the audio frame is at the same time as the first frame of the video frame, then control the digital human to speak. That is, when the first frame of the audio coincides with the first frame of the current video in time, start playing and play at the frame rate F1 until the end.

[0055] In the embodiment of the present application, by adding timestamps to the audio frame and the video frame, the time synchronization of the audio frame and the video frame is maintained during the playback on the terminal, so as to achieve the purpose of keeping the speech and lip movements of the digital human consistent.

[0056] Correspondingly, the embodiment of the present application also provides a device 200 for synchronizing the speech and lip movements of a digital human, as Figure 2 shown. The device 200 includes:

[0057] An acquisition module 202, configured to acquire natural language information;

[0058] A conversion module 204, configured to convert the natural language information into text information;

[0059] A processing module 206, configured to determine a video frame and an audio frame based on the text information;

[0060] A control module 208, configured to control the digital human to speak when the first frame of the audio frame and the first frame of the video frame are at the same time.

[0061] In the embodiment of the present application, the natural language information is obtained through an acquisition module, then the natural language information is converted into text information through a conversion module, and then based on the text information, a video frame and an audio frame are determined through a processing module. When the first frame of the audio frame and the first frame of the video frame are at the same time, the digital human is controlled to speak through a control module, so that the speech and lip movement of the digital human are kept consistent.

[0062] Optionally, in other embodiments of the device, the processing module 206 is specifically configured to:

[0063] Based on the text information, determine the text-to-audio delay time, the text-to-lip movement delay time, the speech speed, the lip movement speed, the audio frame, and the video frame.

[0064] Optionally, in other embodiments of the device, the device 200 further includes:

[0065] An adding module 210, configured to add timestamps to the audio frame and the video frame respectively.

[0066] Optionally, in other embodiments of the device, the adding module 210 is specifically configured to:

[0067] Determine the time difference between the video and the video generation according to the text-to-audio delay time and the text-to-lip movement delay time;

[0068] Determine the audio frame step size according to the audio frame, the speech speed, and the lip movement speed;

[0069] Starting from 0, calculate the relative timestamp of each frame of audio according to the audio frame step size;

[0070] Calculate the absolute timestamp of each frame of audio according to the relative timestamp of each frame of audio, the current time, the text-to-audio delay time, and the text-to-lip movement delay time.

[0071] It should be noted that the above device for synchronizing the speech and lip movement of the digital human can execute the method for synchronizing the speech and lip movement of the digital human provided by the embodiment of the present application, and has the corresponding functional modules and beneficial effects of the method. For technical details not described in detail in the embodiment of the device for synchronizing the speech and lip movement of the digital human, reference may be made to the method for synchronizing the speech and lip movement of the digital human provided by the embodiment of the present application.

[0072] Figure 3 is a schematic hardware structure diagram of an electronic device provided by an embodiment of the present application. As Figure 3 shown, the electronic device 300 includes:

[0073] One or more processors 301 and a memory 302, Figure 3 Taking one processor as an example.

[0074] The processor 301 and the memory 302 can be connected by a bus or other means. Figure 3 Taking the connection via the bus as an example.

[0075] As a non-volatile computer-readable storage medium, the memory 302 can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules, such as the program instructions / modules corresponding to the method of digital human voice and lip synchronization in the embodiments of the present application. By running the non-volatile software programs, instructions, and modules stored in the memory 302, the processor 301 executes various functional applications and data processing of the electronic device, that is, implements the method of digital human voice and lip synchronization in the above method embodiments.

[0076] The memory 302 can include a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store data created according to the use of the digital human voice and lip synchronization device, etc. In addition, the memory 302 can include high-speed random access memory, and can also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices. In some embodiments, the memory 302 may optionally include a memory remotely provided relative to the processor 301, and these memories can be connected to the digital human voice and lip synchronization device through a network. Examples of the above network include but are not limited to the Internet, enterprise intranet, local area network, mobile communication network, and their combinations.

[0077] The embodiments of the present application also provide a non-volatile computer-readable storage medium. The computer-readable storage medium stores computer-executable instructions. When the computer-executable instructions are executed by one or more processors, the one or more processors can execute the method for adjusting the single-module brightness color temperature coefficient in any of the above embodiments.

[0078] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution in this embodiment.

[0079] Through the description of the above embodiments, those of ordinary skill in the art can clearly understand that each embodiment can be implemented by means of software plus a general hardware platform, and of course, it can also be implemented by hardware. Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM), etc.

[0080] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; under the idea of the present application, the technical features in the above embodiments or different embodiments can also be combined, and the steps can be implemented in any order, and there are many other changes in different aspects of the present application as described above. For the sake of brevity, they are not provided in detail; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present application.

Claims

1. A method for synchronizing the speech and lip movements of a digital human, characterized in that, The method includes: Obtaining natural language information; Converting the natural language information into text information; Based on the text information, determining video frames and audio frames; When the first frame of the audio frame is at the same time as the first frame of the video frame, controlling the digital human to speak.

2. The method according to claim 1, characterized in that The determining video frames and audio frames based on the text information includes: Based on the text information, determining the text-to-audio delay time, the text-to-lip delay time, the speech rate, the lip rate, audio frames, and video frames.

3. The method according to claim 2, wherein After determining the video frames and audio frames based on the text information, the method further includes: Adding timestamps to the audio frames and the video frames respectively.

4. The method according to claim 3, wherein The adding a timestamp to the audio frame includes: Determining the time difference between the video and the video generation according to the text-to-audio delay time and the text-to-lip delay time; Determining the audio frame step size according to the audio frame, the speech rate, and the lip rate; Starting from 0, calculating the relative timestamp of each audio frame according to the audio frame step size; Calculating the absolute timestamp of each audio frame according to the relative timestamp of each audio frame, the current time, the text-to-audio delay time, and the text-to-lip delay time.

5. A device for synchronizing the speech and lip movements of a digital human, characterized in that, The apparatus includes: An obtaining module, configured to obtain natural language information; A conversion module, configured to convert the natural language information into text information; A processing module, configured to determine video frames and audio frames based on the text information; A control module, configured to control the digital human to speak when the first frame of the audio frame and the first frame of the video frame are at the same time.

6. The device according to claim 5, characterized in that, The processing module is specifically configured to: Based on the text information, determine the text-to-audio delay time, the text-to-lip delay time, the speech rate, the lip rate, audio frames, and video frames.

7. The device according to claim 6, characterized in that, The apparatus further includes: An adding module, configured to add timestamps to the audio frames and the video frames respectively.

8. The device according to claim 7, characterized in that, The adding module is specifically configured to: Determine the time difference between the video and the video generation according to the text-to-audio delay time and the text-to-lip delay time; Determine the audio frame step size according to the audio frame, the speech rate, and the lip rate; Starting from 0, calculate the relative timestamp of each audio frame according to the audio frame step size; Calculate the absolute timestamp of each audio frame according to the relative timestamp of each audio frame, the current time, the text-to-audio delay time, and the text-to-lip delay time.

9. An electronic device, characterized in that, Includes: At least one processor; And, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the method according to any one of claims 1-4.

10. A non-volatile computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, and when the computer-executable instructions are executed by a processor, the processor executes the method according to any one of claims 1-4.

Citation Information

Cited By

  • Virtual digital human multimedia teaching interaction method and system and storage medium

    CN121888060A

  • A virtual digital human multimedia teaching interaction method and system and a storage medium

    CN121888060B