Human-computer interaction method and device, computer equipment and storage medium
By obtaining the spectral characteristics of audio data and using the syllable mouth shape mapping table to determine the target mouth shape corresponding to the target syllable, the problem of the 3D facial mouth shape being out of sync with the audio is solved, the synchronization of the 3D facial mouth shape and audio is achieved, and the user experience is improved.
Patent Information
- Application Number
- CN202510913331.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-02
- Publication Date
- 2025-10-17
AI Technical Summary
In the prior art, the problem of lip shape and audio being out of sync with the 3D face results in a poor user experience.
By obtaining the spectral features of the audio data, the target lip shape corresponding to the target syllable is determined using the syllable lip shape mapping table, and the lip shape of the three-dimensional face is synchronized to achieve synchronization of audio and lip shape.
It achieves synchronization of 3D facial lip shape and audio, improving user experience.
Smart Images

Figure CN120803267A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of human-computer interaction, and in particular to a human-computer interaction method, device, computer equipment and storage medium. BACKGROUND
[0002] In the related art, a three-dimensional face of a digital person is displayed on the screen of a terminal, and in the interaction process between a user and the digital person, audio corresponding to the speaking content is played on a loudspeaker through text-to-speech (TTS), and the mouth shape of the digital person corresponding to the text is rendered. However, due to the huge search range of the mouth shape lookup table corresponding to Chinese characters, the mouth shape of the three-dimensional face is out of sync with the audio, and the user experience is poor. SUMMARY
[0003] Therefore, the present application provides a human-computer interaction method, device, computer equipment and storage medium to solve the problem of the out-of-sync mouth shape of the three-dimensional face and the audio in the related art.
[0004] In a first aspect, the present application provides a human-computer interaction method, which comprises:
[0005] obtaining audio data to be played, the audio data being generated based on text;
[0006] determining a target syllable corresponding to the audio data according to a spectral feature corresponding to the audio data;
[0007] determining a target mouth shape corresponding to the target syllable based on a syllable-mouth shape mapping table, the syllable-mouth shape mapping table being used to store a corresponding relationship between syllables and mouth shapes;
[0008] playing the audio data and synchronizing the mouth shape of the three-dimensional face displayed on the target screen to the target mouth shape.
[0009] In an optional implementation, determining a target syllable corresponding to the audio data according to a spectral feature corresponding to the audio data comprises:
[0010] converting the audio data into first frequency data in the frequency domain;
[0011] extracting a frequency domain feature of the first frequency data and an intensity feature of the audio data as the spectral feature, the frequency domain feature being a core frequency of the audio data, and the intensity feature being an amplitude of the sound corresponding to the core frequency, the core frequency being a frequency value corresponding to a peak point in a spectral graph of the audio data;
[0012] determining the target syllable from a syllable feature lookup table according to the frequency domain feature and the intensity feature, the syllable feature lookup table being used to store a corresponding relationship between syllables and the frequency domain feature and the intensity feature.
[0013] In an optional implementation, before the target syllable is looked up from the syllable feature lookup table according to the frequency domain feature and the intensity feature, the method further comprises:
[0014] According to the spectrogram of the second frequency data, a feature frequency is determined as the frequency domain feature, the feature frequency being a frequency mean of a section in which a frequency is greater than or equal to a frequency threshold in the spectrogram or a frequency peak in the spectrogram;
[0015] An intensity corresponding to the feature frequency is obtained as the intensity feature;
[0016] The syllable feature lookup table is constructed by using the frequency domain feature, the intensity feature, and the sample syllable.
[0017] In an optional implementation, before the target mouth shape corresponding to the target syllable is determined based on the syllable-mouth shape mapping table, the method further comprises:
[0018] An image dataset is obtained, the image dataset comprising images of mouth shapes or mouths corresponding to the sample syllables;
[0019] A preset number of feature parameters of the mouth corresponding to each sample syllable are extracted from the image dataset;
[0020] The syllable-mouth shape mapping table is constructed based on the feature parameters of the mouth and the sample syllables.
[0021] In an optional implementation, the target mouth shape corresponding to the target syllable is determined based on the syllable-mouth shape mapping table, comprising:
[0022] A target feature parameter of the mouth corresponding to the target syllable is determined from the syllable-mouth shape mapping table;
[0023] The target feature parameter is mapped into a to-be-generated mouth shape parameter in a coordinate system of the three-dimensional face, the to-be-generated mouth shape parameter corresponding to the target mouth shape;
[0024] The current mouth shape of the three-dimensional face is rendered into the target mouth shape by using the to-be-generated mouth shape parameter.
[0025] In an optional implementation, the mouth shape of the three-dimensional face displayed on the target screen is synchronized to the target mouth shape, comprising:
[0026] The mouth shape of the three-dimensional face displayed on the target screen is synchronously switched to the target mouth shape.
[0027] In an optional implementation, the audio data is played, comprising:
[0028] The audio data is played by using an audio link associated with the target screen, the audio link representing a connection path of a series of devices and components through which the audio data passes in a transmission process;
[0029] The audio data corresponding sound signal is sent back to the audio software module or the digital signal processing chip associated with the target screen through the loopback channel.
[0030] In a second aspect, the present application provides a human-computer interaction device, comprising:
[0031] An acquisition module is configured to acquire audio data to be played, the audio data being generated based on text.
[0032] A syllable determination module is configured to determine a target syllable corresponding to the audio data according to a spectral feature corresponding to the audio data.
[0033] A mouth shape determination module is configured to determine a target mouth shape corresponding to the target syllable based on a syllable-mouth shape mapping table, the syllable-mouth shape mapping table being configured to store a corresponding relationship between syllables and mouth shapes.
[0034] A synchronization module is configured to play the audio data and synchronize a mouth shape of a three-dimensional face displayed on a target screen to the target mouth shape.
[0035] In a third aspect, the present application provides a computer device, comprising a memory and a processor, the memory and the processor being communicatively connected with each other, the memory storing computer instructions, and the processor executing the computer instructions to perform the human-computer interaction method of the first aspect or any of the corresponding embodiments thereof.
[0036] In a fourth aspect, the present application provides a computer readable storage medium, the computer readable storage medium storing computer instructions, the computer instructions being configured to cause a computer to perform the human-computer interaction method of the first aspect or any of the corresponding embodiments thereof.
[0037] In a fifth aspect, the present application provides a computer program product, comprising computer instructions, the computer instructions being configured to cause a computer to perform the human-computer interaction method of the first aspect or any of the corresponding embodiments thereof.
[0038] The human-computer interaction method of the present application can achieve the following beneficial effects:
[0039] The audio data to be played is acquired, the audio data being generated based on text, thereby providing a data source for the generation of the mouth shape of the three-dimensional face and the playing of the sound. The target syllable corresponding to the audio data is determined according to the spectral feature corresponding to the audio data, thereby providing syllable data as a basis for the generation of the mouth shape. The target mouth shape corresponding to the target syllable is determined based on the syllable-mouth shape mapping table, the syllable-mouth shape mapping table being configured to store a corresponding relationship between syllables and mouth shapes, thereby directly utilizing the corresponding relationship between the syllables and the mouth shapes to quickly determine the target mouth shape corresponding to the syllable. The audio data is played, and the mouth shape of the three-dimensional face displayed on the target screen is synchronized to the target mouth shape, thereby realizing the synchronization of the mouth shape of the three-dimensional face and the audio and improving the user experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the specific embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0041] Figure 1 is a flowchart of a human-computer interaction method according to an embodiment of the present invention;
[0042] Figure 2 is a flow chart of another human-computer interaction method according to an embodiment of the present invention;
[0043] Figure 3 is a flowchart of another human-computer interaction method according to an embodiment of the present invention;
[0044] Figure 4 A schematic diagram of another human-computer interaction process according to an embodiment of the present invention;
[0045] Figure 5 A schematic diagram of the sound frequency of the syllable corresponding to "hello" according to an embodiment of the present invention;
[0046] Figure 6 A schematic diagram of the sound intensity of the syllable corresponding to "hello" according to an embodiment of the present invention;
[0047] Figure 7 is a structural block diagram of a human-computer interaction device according to an embodiment of the present invention;
[0048] Figure 8 Schematic diagram of the hardware structure of a computer device according to an embodiment of the present invention. DETAILED DESCRIPTION
[0049] To make the purpose, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without making creative efforts shall fall within the scope of protection of the present invention.
[0050] The technical solution of the embodiment of the present application has many application scenarios or fields, and can be applied in the field of virtual people or digital people, such as virtual anchors or hosts, virtual idols, virtual spokespersons or customer service personnel of enterprises, virtual avatars in the metaverse, and the like. It can also be used in the field of entertainment and games, such as game character dialogue, production of animated movies or series, and the like. It can also be used in the field of customer service and interaction, such as intelligent customer service, voice assistant visualization, virtual service ambassadors in bank or government halls. It can also be used in the field of social media and communication, such as personalized virtual image communication, special effects or virtual images generated according to real-time changes in lip shapes for voice, and the like. It can be used to improve the naturalness, immersion and efficiency of human-computer interaction or digital content presentation. The foregoing application scenarios are only examples, and the application scenarios of the present application are not limited thereto.
[0051] In a voice interaction human machine interface (HMI), a three-dimensional (3D) face is displayed on a screen, and interaction with a user can be performed. For example, a three-dimensional face is displayed on a screen of a car machine large screen, a mobile phone, a television, and the like, and according to audio data generated by TTS, changes in facial expressions and opening and closing of lips are matched to perform more realistic and immersive interaction with a user. In the related art, a Chinese character is used to bind a lip shape of a three-dimensional face, and a lip shape corresponding to the Chinese character is generated by using a generated Chinese character, and audio generated by TTS corresponding to the Chinese character is simultaneously displayed and played. Since the number of Chinese characters is large, the number of mappings generated by the lip shapes is also large, and the process of retrieving, mapping, and using the lip shape corresponding to the Chinese character is time-consuming. This leads to a problem of asynchronization between the lip shape of the three-dimensional face and the played audio (audio-visual asynchronization).
[0052] The embodiment of the present application provides a human-computer interaction method, which directly determines the lip shape of a three-dimensional face by using syllable features corresponding to audio data generated by TTS, so as to achieve synchronization between the lip shape of the three-dimensional face and the played audio data.
[0053] According to the embodiment of the present application, a human-computer interaction method is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a group of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0054] In the present embodiment, a human-computer interaction method is provided, which can be used in terminal devices such as mobile phones, tablet computers, personal computers, and the like, Figure 1 The flowchart of the human-computer interaction method according to the embodiment of the present application is shown in FIG. 1, which includes the following steps: Figure 1
[0055] Step S101: Acquire audio data to be played, where the audio data is generated based on text.
[0056] In this embodiment, the audio data to be played refers to the audio data that is about to be played by the terminal device that interacts with the user, and the terminal device also displays the three-dimensional face that interacts with the user. Preferably, the audio data is a sound generated based on text and by TTS. For example, when a user is talking to a three-dimensional face, the three-dimensional face responds after receiving the user's instructions. Optionally, the text content of the response can be a preset wording or default phrase, or it can be text generated by a large language model. The text can be Chinese text or text in other languages, such as English text. By generating audio data from text, an audio reference is provided for the mouth sound of the three-dimensional face.
[0057] Step S102: determining a target syllable corresponding to the audio data based on the frequency spectrum features corresponding to the audio data.
[0058] In this embodiment, spectral features refer to the mathematical representation of sound in the frequency domain and time domain, indicating the strength and amplitude of different frequency components. The target syllable refers to the smallest pronunciation unit identified from the audio data that currently needs to drive the lip animation. In linguistics, a syllable is a natural unit of speech, consisting of a vowel and a consonant, and can be pronounced independently. Optionally, in Chinese, there are 405 basic syllables that do not contain tones. For example, "hello" is two syllables, namely "ni" and "hao". Each basic syllable has a unique spectral feature (i.e., the strength distribution and amplitude in different frequency bands). By analyzing the spectral features of the audio data and matching them with the spectral features of the basic syllables, the target syllable corresponding to the audio data segment can be determined.
[0059] Step S103: determining a target mouth shape corresponding to the target syllable based on the syllable mouth shape mapping table, where the syllable mouth shape mapping table is used to store the correspondence between syllables and mouth shapes.
[0060] In this embodiment, the syllable-to-mouth shape mapping table refers to a database or table that maps mouth shape parameters to syllable data, and is used to store the correspondence between syllables and mouth shapes. The target mouth shape refers to the mouth shape corresponding to a human pronouncing the target syllable. Based on the target syllable in the syllable-to-mouth shape mapping table, the target mouth shape corresponding to the target syllable can be queried.
[0061] Step S104: play the audio data and synchronize the lip shape of the three-dimensional face displayed on the target screen to the target lip shape.
[0062] In this embodiment, the target screen refers to the display screen of the terminal device interacting with the user. The display screen can be a two-dimensional display screen or a three-dimensional holographic display system supporting the display of a three-dimensional image of a face. The three-dimensional face refers to a digital face model with a three-dimensional spatial structure and dynamic control, which is constructed by computer graphics technology and supports real-time deformation. The position of the vertex of the three-dimensional structure can be adjusted to change the mouth shape (such as opening the mouth or grinning). The terminal device plays audio data through the audio module, and synchronizes the mouth shape of the displayed three-dimensional face to the target mouth shape on the target screen.
[0063] The method for human-computer interaction provided in this embodiment acquires audio data to be played, which is generated based on text, thereby providing a data source for the generation of the mouth shape of the three-dimensional face and the voice playing. According to the spectral features corresponding to the audio data, the target syllable corresponding to the audio data is determined, thereby providing syllable data as a basis for the generation of the mouth shape. Based on the syllable-mouth shape mapping table, the target mouth shape corresponding to the target syllable is determined. The syllable-mouth shape mapping table is used to store the corresponding relationship between the syllable and the mouth shape, and the corresponding relationship between the syllable and the mouth shape is directly used, instead of using the mapping database of Chinese characters and mouth shapes in the related art, to quickly determine the target mouth shape corresponding to the syllable. The audio data is played, and the mouth shape of the three-dimensional face displayed on the target screen is synchronized to the target mouth shape, thereby realizing the synchronization of the mouth shape of the three-dimensional face and the audio and improving the user experience.
[0064] In this embodiment, a method for human-computer interaction is provided, which can be used in the terminal device as described above, such as a mobile phone, a tablet computer, a personal computer, and the like. Figure 2 FIG. 4 is a flowchart of another method for human-computer interaction according to an embodiment of the present application, as shown in FIG. 4, the flowchart includes the following steps: Figure 2
[0065] In step S201, audio data to be played is acquired, which is generated based on text.
[0066] For details, refer to step S101 of the embodiment shown in FIG. 1, which will not be described here again. Figure 1
[0067] In step S202, according to the spectral features corresponding to the audio data, the target syllable corresponding to the audio data is determined.
[0068] Specifically, step S202 includes the following steps.
[0069] In step S2021, the audio data is converted into first frequency data in the frequency domain.
[0070] In this embodiment, the first frequency data refers to the frequency data of the current audio data to be played in the frequency domain. The time-domain audio data is converted into the first frequency data in the frequency domain through Fast Fourier Transform (FFT).
[0071] In step S2022, the frequency domain feature of the first frequency data and the intensity feature of the audio data are extracted as the spectrum features. The frequency domain feature is the core frequency of the audio data, and the intensity feature is the amplitude of the sound corresponding to the core frequency. The core frequency is the frequency value corresponding to the peak point in the spectrum graph of the audio data.
[0072] In this embodiment, the core frequency is the frequency value corresponding to the peak point in the spectrum graph of the audio data. The frequency domain feature is the core frequency of the audio data. The intensity feature is the amplitude of the sound corresponding to the core frequency, which can be described by decibels (dB) as a unit. The spectrum features and the intensity features of the audio data are extracted from the first frequency data.
[0073] In step S2023, the target syllable is searched from the syllable feature lookup table according to the frequency domain feature and the intensity feature. The syllable feature lookup table is used to store the corresponding relationship between the syllable and the frequency domain feature and the intensity feature.
[0074] In this embodiment, the syllable feature lookup table refers to a data table or a database of the syllable, the frequency domain feature of the syllable in the frequency domain, and the intensity feature of the syllable, which is used to store the corresponding relationship between the syllable and the frequency domain feature and the intensity feature. According to the frequency domain feature and the intensity feature, the corresponding target syllable can be searched from the syllable feature lookup table.
[0075] In step S203, the target mouth shape corresponding to the target syllable is determined based on the syllable mouth shape mapping table. The syllable mouth shape mapping table is used to store the corresponding relationship between the syllable and the mouth shape.
[0076] For details, please refer to Figure 1 In step S103 of the embodiment shown in the figure, no further description is given here.
[0077] In step S204, the audio data is played, and the mouth shape of the three-dimensional face displayed on the target screen is synchronized to the target mouth shape.
[0078] For details, please refer to Figure 1 In step S104 of the embodiment shown in the figure, no further description is given here.
[0079] According to the spectrum features and the intensity features of the audio data in the frequency domain, the target syllable corresponding to the audio data is determined from the syllable feature lookup table, which provides a basis for generating the mouth shape of the three-dimensional face.
[0080] In some optional embodiments, before step S2023 described above, the following steps are included:
[0081] Step a1, according to the spectrogram of the second frequency data, determine the characteristic frequency as the frequency domain feature, the characteristic frequency is the frequency mean of the section whose frequency is greater than or equal to the frequency threshold in the spectrogram or the frequency peak in the spectrogram.
[0082] In this embodiment, the second frequency data refers to the frequency data corresponding to all syllables collected for creating the syllable feature lookup table. The spectrogram refers to the visual representation of the frequency of the sound signal changing with time, the X axis of the spectrogram is time, unit is second, the Y axis is frequency, unit is hertz (Hz). The intensity of the sound is represented by color or brightness. The characteristic frequency is the frequency mean of the section whose frequency is greater than the frequency threshold in the spectrogram or the frequency peak in the spectrogram, for example, the average value of all frequencies greater than 100 Hz in the spectrogram is taken as the characteristic frequency, or the frequency values of the multiple peaks of the frequency fluctuation curve in the spectrogram are taken as the characteristic frequency. The spectrogram of the second frequency data is generated, and the characteristic frequency in the spectrogram is taken as the frequency spectrum feature.
[0083] Step a2, obtain the sound intensity corresponding to the characteristic frequency as the intensity feature.
[0084] In this embodiment, according to the determined characteristic frequency, the sound intensity corresponding to the characteristic frequency in the spectrogram is determined as the intensity feature.
[0085] Step a3, use the frequency domain feature, the intensity feature and the sample syllable to construct the syllable feature lookup table.
[0086] In this embodiment, according to the obtained frequency domain feature, intensity feature and sample syllable corresponding to the sample syllable, a data table or database for storing syllables and corresponding frequency domain features and intensity features is created as the syllable feature lookup table. Wherein, the sample syllable refers to any basic syllable used to construct the lookup table.
[0087] Through this implementation, a syllable feature lookup table for looking up the corresponding sample syllable by sound features is constructed, which helps to quickly complete the recognition of the target syllable.
[0088] In some optional embodiments, before the above step S203, comprising:
[0089] Step b1, obtain an image data set, the image data set includes the image of the mouth shape or the mouth corresponding to the sample syllable.
[0090] In this embodiment, the image data set refers to the data set collected for extracting the image of the mouth shape or the mouth when the sample syllable is uttered. The data set containing the image of the mouth shape or the mouth when the sample syllable is uttered is obtained.
[0091] Step b2, extracting a preset number of feature parameters of the mouth corresponding to each sample syllable from the image data set.
[0092] In this embodiment, a preset number of feature parameters of the mouth corresponding to each sample syllable are extracted from each image of the image data set. Alternatively, a face is recognized from the image data set through face recognition, and then a mouth key point of the face is extracted through face key point detection to determine the feature parameters of the mouth. Alternatively, the mouth is directly recognized to extract the feature parameters of the mouth, which can be recorded in the form of a vector, a matrix or a tensor.
[0093] Step b3, constructing a syllable mouth shape mapping table based on the feature parameters of the mouth and the sample syllables.
[0094] In this embodiment, according to the feature parameters of the mouth and the sample syllables, the corresponding relationship between the two is stored in a data table and a database to construct a syllable mouth shape mapping table.
[0095] Through this embodiment, a mapping table of syllables and corresponding mouth shapes is constructed. Since the number of mouth shapes corresponding to syllables is much smaller than the number of mouth shapes mapped by Chinese characters, for example, the number of syllables for pronunciation of Chinese characters is 405, and the number of Chinese characters is more than 90,000. Through the syllable mouth shape mapping table, the mouth shape parameters or the feature parameters of the mouth corresponding to the target syllable can be quickly obtained.
[0096] In this embodiment, a human-computer interaction method is provided, which can be used in the terminal device described above, such as a mobile phone, a tablet computer and a personal computer, etc. Figure 3 is a flowchart of another human-computer interaction method according to an embodiment of the present application, as shown in Figure 3 , the flowchart includes the following steps:
[0097] Step S301, obtaining audio data to be played, the audio data being generated based on text.
[0098] For details, please refer to step S101 of the embodiment shown in Figure 1 , which will not be repeated here.
[0099] Step S302, determining a target syllable corresponding to the audio data according to a spectrum feature corresponding to the audio data.
[0100] For details, please refer to step S102 of the embodiment shown in Figure 1 , which will not be repeated here.
[0101] Step S303, determining a target mouth shape corresponding to the target syllable based on a syllable mouth shape mapping table, the syllable mouth shape mapping table being used to store a corresponding relationship between syllables and mouth shapes.
[0102] Specifically, the above step S303 includes:
[0103] Step S3031, determining a target feature parameter of a mouth corresponding to the target syllable from the syllable mouth shape mapping table.
[0104] In this embodiment, the target feature parameter refers to the feature parameter of the mouth corresponding to the target syllable in the syllable mouth shape mapping table. According to the target syllable, the target feature parameter of the mouth corresponding to the target syllable is found from the syllable mouth shape mapping table.
[0105] Step S3032, mapping the target feature parameter to a to-be-generated mouth shape parameter in a coordinate system of the three-dimensional face, the to-be-generated mouth shape parameter corresponding to the target mouth shape.
[0106] In this embodiment, the to-be-generated mouth shape parameter refers to a parameter for rendering the target mouth shape of the face in the coordinate system of the three-dimensional face. Mapping the target feature parameter to the to-be-generated mouth shape parameter in the coordinate system of the three-dimensional face provides a basis for three-dimensional display of the target mouth shape.
[0107] Step S3033, rendering the current mouth shape of the three-dimensional face to the target mouth shape by using the to-be-generated mouth shape parameter.
[0108] In this embodiment, the current mouth shape of the three-dimensional face is rendered to the target mouth shape in the screen of the terminal device by using the to-be-generated mouth shape parameter.
[0109] Step S304, playing the audio data and synchronizing the mouth shape of the three-dimensional face displayed on the target screen to the target mouth shape.
[0110] Through this implementation, the target mouth shape corresponding to the target syllable is determined based on the syllable mouth shape mapping table, and accurate mouth shape display is obtained.
[0111] In some optional embodiments, the step S304 of synchronizing the mouth shape of the three-dimensional face displayed on the target screen to the target mouth shape includes:
[0112] synchronously switching the mouth shape of the three-dimensional face displayed on the target screen to the target mouth shape.
[0113] In this embodiment, the mouth shape of the three-dimensional face displayed on the target screen is a default mouth shape or a mouth shape maintained after the pronunciation is just completed, and the mouth shape of the three-dimensional face needs to be switched. Alternatively, the mouth shape of the three-dimensional face displayed on the target screen is synchronously switched to the target mouth shape by adjusting the position of the vertex coordinates of the mouth in the three-dimensional model of the mouth through rotation or translation.
[0114] In some optional embodiments, the step S304 of playing the audio data includes:
[0115] The audio data is played by using an audio link associated with the target screen, and the audio link represents a connection path of a series of devices and components through which the audio data passes during transmission.
[0116] The audio data is played by using an audio link associated with the target screen, and the audio link represents a connection path of a series of devices and components through which the audio data passes during transmission.
[0117] In the embodiment, the loopback refers to an internal signal loopback mechanism in the software or hardware layer, which can acquire the audio data to be played and provide the audio data to the current function module.
[0118] Optionally, the loopback is directly connected to an audio data stream transmission layer by adding the loopback in an audio hardware abstraction layer (Hardware Abstraction Layer for Audio, audio HAL) layer or a digital signal processing chip of the system, so as to avoid system buffer delay and realize audio and picture frame level synchronization.
[0119] The embodiment of the application provides a human-computer interaction method. Figure 4 Another flowchart of the human-computer interaction of the embodiment of the application is shown in FIG. 4. Figure 4 As shown in FIG. 4, in the initial state, the method comprises the following steps.
[0120] In step S401, the three-dimensional face is directly faced to the user, the mouth shape is closed, and the expression is formal.
[0121] In the interaction state, the method comprises the following steps.
[0122] In step S402, if the three-dimensional face is woken up by a voice keyword, for example, "Hello, little P", the wake-up event and a wake-up sound area, or other multi-modal information, are transmitted to the three-dimensional face application. The three-dimensional face turns the face to the position of the wake-up sound area and enters the interaction state. Alternatively, when a key is woken up, for example, a key on a steering wheel is woken up, the wake-up event and the wake-up sound area are transmitted to the three-dimensional face application. The wake-up sound area can be set according to the position of the operator. For example, the key wake-up of the steering wheel of the car can be the main driver. When the three-dimensional face application turns the face to the position of the wake-up sound area, the interaction state is entered.
[0123] In step S403, the audio data generated by the TTS is played, the text data of the TTS is sent to the three-dimensional face application, and then the audio data generated by the TTS is started to be played.
[0124] Step S404, the three-dimensional face application maps the text to a specific expression according to the tone words or other text in the text data, and switches the three-dimensional face expression. Among them, for the three-dimensional face expression, the emotion data can be extracted from the text data through a large language model, so as to display the emotion data as different three-dimensional face expressions. For switching the three-dimensional face expression, dynamic drawing can be performed through a graphic rendering tool (such as Unity software). For example, when the text is "ha ha, you are right", the three-dimensional face expression is switched to a smile. When the text is "sorry, I don't know your question yet", the three-dimensional face expression is switched to regret.
[0125] Step S405, when the TTS generated voice audio is played, the pulse code modulation (PCM) signal of the voice data stream needs to be transmitted to the speaker for playing and to the three-dimensional face application at the same time. After receiving the voice data stream, the three-dimensional face application converts it into frequency domain data after FFT.
[0126] Step S406, the three-dimensional face application queries the corresponding syllable and the pronunciation lip shape corresponding to the syllable according to the current frequency domain data, draws the lip shape animation through Unity and other tools, and completes the animation display of the three-dimensional face lip shape.
[0127] Among them, the Chinese pinyin combination mainly contains 405 syllables. The camera collects the lip shape of a human saying the 405 syllables, records the parameters of the width, height, mouth corner up angle, and double-lip opening distance of each lip shape corresponding to the mouth, and saves them to the syllable lip shape mapping table. For example, the "ni" syllable corresponds to the lip shape information: width 5 cm (or pixels), height 2 cm (or pixels), mouth corner up angle 20 degrees, double-lip opening distance 0.5 cm (or pixels), etc. The 405 syllables generated by TTS are played, and the audio data of each syllable during playing is saved.
[0128] Step S407, after the TTS generated audio data stream stops playing, exit the interactive state and enter the initial state to wait for the next trigger and interaction.
[0129] In an example of the embodiment, Figure 5 The sound frequency diagram of the "hello" corresponding syllable of the embodiment of the application. The sound data of each syllable is converted into frequency data in the frequency domain through FFT, and the sound of "ni hao" in the figure corresponds to the frequency spectrum diagram. The horizontal axis is time and the vertical axis is sound frequency. Based on the frequency spectrum diagram, the characteristics of the sound in the frequency domain can be obtained. The frequency band of 100Hz to 400Hz in the "ni" corresponding frequency band is the characteristic frequency band, and the frequency band of 100Hz to 1000Hz in the "hao" corresponding frequency band is the characteristic frequency band.
[0130] In one example of the embodiment, Figure 6 The sound intensity diagram of the "ni" syllable is shown in the figure. The corresponding frequency domain data and intensity data are (114 Hz, -16.6 dB), (242 Hz, -12.3 dB), (371 Hz, -19.7 dB), and (1002 Hz, -39 dB). These spectral feature values and intensity feature values are summarized and saved to the syllable feature lookup table. Thus, two tables, the syllable mouth shape mapping table of 405 syllables and the spectral summary table of 405 syllables, are generated. The spectral summary table can be used as the syllable feature lookup table. When the audio data generated by the actual TTS is played, the data stream of the voice data being played by the TTS is subjected to FFT to convert it into frequency domain data. The voice data stream can be obtained by adding a loopback channel in the audio HAL layer or a digital signal processing (DSP) chip. In this way, the TTS voice data is closer to the actual speaker playing end, avoiding the audio-visual synchronization deviation caused by the data transmission delay in the system intermediate layer.
[0131] The spectral features and intensity features of the frequency data of the voice in the frequency domain are compared with the data of the syllable feature lookup table of 405 syllables. The extracted spectral features and intensity features are compared with each entry in the syllable feature lookup table to find the syllable entry with the highest matching degree. The syllable corresponding to the entry is taken as the TTS syllable being currently played. Then, the mouth shape information corresponding to the syllable is obtained from the syllable mouth shape mapping table. For example, the mouth shape parameters of the "ni" syllable are a width of 5 cm, a height of 2 cm, a mouth corner upwarping angle of 20 degrees, and a double-lip opening distance of 0.5 cm.
[0132] The three-dimensional face mouth shape parameters are adjusted and animations are generated and displayed on the three-dimensional face by calling a 3D drawing software such as Unity. Through the above processing, the data stream of the voice data being played by the TTS is obtained in real time, subjected to fast Fourier transform, and the corresponding syllable is obtained by table lookup. The mouth shape animation corresponding to the syllable is drawn and displayed on the three-dimensional face. Because the TTS audio data being played is obtained, the three-dimensional face mouth shape change is accurately synchronized with the actual speaker played voice.
[0133] In some examples of the embodiment, when the data stream of the TTS voice data stops playing, the three-dimensional face returns to the initial state. The face orientation can be restored to the front, the mouth shape can be restored to the closed state, and the expression can be restored to the formal state, waiting for the next voice interaction. In the initial state, the face is not completely still and can have some micro-expressions, which are more realistic. For example, the eyes occasionally blink and the nose occasionally breathes.
[0134] Through the above process, the three-dimensional face application can switch between the interactive state and the initial state, realizing continuous HMI interaction with the interlocutor. Among them, the three-dimensional face user can switch different types (such as handsome men, beautiful women, etc.), meeting the personalized needs.
[0135] For the voice wake-up sound area, the three-dimensional face orientation is realized. The three-dimensional face orientation transformation can also be completed according to other multi-modal recognition information. For example, through the camera of the passenger monitoring system, face recognition is performed to automatically detect the position of the passenger. When passengers get on the vehicle in different seats, the three-dimensional face can automatically face the passenger and greet.
[0136] Through the method of the embodiment, the audio data to be played is acquired, the audio data is generated based on text, and a data source is provided for three-dimensional face lip shape generation and sound playing; a target syllable corresponding to the audio data is determined according to a spectrum feature corresponding to the audio data, and syllable data is provided as a basis for lip shape generation; a target lip shape corresponding to the target syllable is determined based on a syllable lip shape mapping table, the syllable lip shape mapping table is used to store a corresponding relationship between syllables and lip shapes, the corresponding relationship between syllables and lip shapes is directly utilized to quickly determine the target lip shape corresponding to the syllable; the audio data is played, and the lip shape of the three-dimensional face displayed on the target screen is synchronized to the target lip shape, the synchronization of the three-dimensional face lip shape and the audio is realized, and the user experience is improved.
[0137] In the embodiment, a human-computer interaction device is also provided, which is used to implement the above-mentioned embodiments and preferred embodiments, and will not be described again. As used below, the term "module" can be a combination of software and / or hardware that implements a predetermined function. Although the device described in the following embodiments is preferably implemented in software, hardware, or a combination of software and hardware is also possible and contemplated.
[0138] The embodiment provides a human-computer interaction device, as shown in Figure 7 The device includes:
[0139] The acquisition module 701 is configured to acquire audio data to be played, the audio data being generated based on text.
[0140] The syllable determination module 702 is configured to determine a target syllable corresponding to the audio data according to a spectrum feature corresponding to the audio data.
[0141] The lip shape determination module 703 is configured to determine a target lip shape corresponding to the target syllable based on a syllable lip shape mapping table, the syllable lip shape mapping table being used to store a corresponding relationship between syllables and lip shapes.
[0142] The synchronization module 704 is configured to play the audio data and synchronize the lip shape of the three-dimensional face displayed on the target screen to the target lip shape.
[0143] In some optional embodiments, the syllable determination module 702 comprises:
[0144] a conversion unit configured to convert the audio data into first frequency data in a frequency domain;
[0145] an extraction unit configured to extract a frequency domain feature of the first frequency data and an intensity feature of the audio data as the spectral feature, the frequency domain feature being a core frequency of the audio data, and the intensity feature being an amplitude of a sound corresponding to the core frequency, the core frequency being a frequency value corresponding to a peak point in a frequency spectrum of the audio data;
[0146] a searching unit configured to search for a target syllable from a syllable feature lookup table according to the frequency domain feature and the intensity feature, the syllable feature lookup table being configured to store a corresponding relationship between a syllable and the frequency domain feature and the intensity feature.
[0147] In some optional embodiments, the searching unit comprises:
[0148] a spectral feature determination subunit configured to determine a feature frequency as the frequency domain feature according to a frequency spectrum of the second frequency data, the feature frequency being a frequency mean value of a frequency range greater than or equal to a frequency threshold in the frequency spectrum or a frequency peak value in the frequency spectrum;
[0149] an intensity feature acquisition subunit configured to acquire a sound intensity corresponding to the feature frequency as the intensity feature;
[0150] a lookup table construction subunit configured to construct the syllable feature lookup table by using the frequency domain feature, the intensity feature, and a sample syllable.
[0151] In some optional embodiments, the mouth shape determination module 703 comprises:
[0152] an image acquisition unit configured to acquire an image data set, the image data set comprising images of mouth shapes or mouth portions corresponding to sample syllables;
[0153] a feature extraction unit configured to extract, from the image data set, a preset number of feature parameters of the mouth portions corresponding to each sample syllable;
[0154] a mapping table construction unit configured to construct a syllable mouth shape mapping table based on the feature parameters of the mouth portions and the sample syllables.
[0155] In some optional embodiments, the mouth shape determination module 703 further comprises:
[0156] a mouth feature determination unit configured to determine, from the syllable mouth shape mapping table, a target feature parameter of a mouth portion corresponding to a target syllable;
[0157] The to-be-generated mouth shape parameter mapping unit is configured to map the target feature parameter into a to-be-generated mouth shape parameter in a coordinate system of the three-dimensional face, and the to-be-generated mouth shape parameter corresponds to the target mouth shape.
[0158] The target mouth shape rendering unit is configured to render the current mouth shape of the three-dimensional face into the target mouth shape by using the to-be-generated mouth shape parameter.
[0159] In some optional embodiments, the synchronization module 704 includes:
[0160] The mouth shape switching unit is configured to synchronously switch the mouth shape of the three-dimensional face displayed on the target screen into the target mouth shape.
[0161] In some optional embodiments, the synchronization module 704 includes:
[0162] The audio playing unit is configured to play the audio data by using an audio link associated with the target screen, and the audio link represents a connection path of a series of devices and components through which the audio data passes in a transmission process.
[0163] The loopback unit is configured to loop back a sound signal corresponding to the audio data into an audio software module or a digital signal processing chip associated with the audio link and the target screen by using a loopback path.
[0164] Further function descriptions of the above-mentioned various modules and units are the same as those of the corresponding embodiments, and will not be repeated here.
[0165] The man-machine interaction device in the embodiment is in the form of a functional unit, and the unit refers to an Application Specific Integrated Circuit (ASIC) circuit, a processor and a memory executing one or more software or fixed programs, and / or other devices that can provide the above-mentioned functions.
[0166] The embodiment of the application further provides a computer device having the above-mentioned Figure 7 man-machine interaction device.
[0167] Please refer to Figure 8 , Figure 8 is a structural schematic diagram of a computer device provided by an optional embodiment of the application, as Figure 8As shown, the computer device includes one or more processors 10, memory 20, and interfaces 30 for external devices such as a keyboard and a mouse and a disk drive. One or more of the interfaces 30 enable a user to interact with the computer device. In some embodiments, the interface 30 also includes an input device, such as a microphone, or output device, such as a speaker. Figure 8 The processor 10 is used in the description as an example.
[0168] The processor 10 can be a central processing unit, a network processor, or a combination thereof. The processor 10 can further include a hardware chip. The hardware chip can be an application specific integrated circuit, a programmable logic device, or a combination thereof. The programmable logic device can be a complex programmable logic device, a field programmable logic device, a general array logic, or any combination thereof.
[0169] The memory 20 stores instructions and receives data from the at least one processor 10. Some embodiments of the computer device include one or more memory devices, including a main memory, and a static random access memory (SRAM). In some embodiments, the computer device includes one or more disk memory devices (e.g., magnetic, optical or tape), or flash memory devices.
[0170] The memory 20 can include a program storage area and a data storage area. The program storage area can store the operating system 21 and application programs 22 needed by at least one function. The data storage area can store data that is used as needed by the computer device, and the like. In addition, the memory 20 can include a high-speed random access memory, and can further include a non-volatile memory, such as at least one disk memory device, a flash memory device, or other non-volatile solid state memory device. In some alternative embodiments, the memory 20 can optionally include a memory that is remotely located with respect to the processor 10, and these remote memories can be connected to the computer device through a network. Examples of the network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and a combination thereof.
[0171] The memory 20 can include a volatile memory, such as a random access memory, and can also include a non-volatile memory, such as a flash memory, a hard disk, or a solid state disk. The memory 20 can further include a combination of the above-mentioned types of memories.
[0172] The computer device also includes an input device 30 and an output device 40. The processor 10, the memory 20, the input device 30 and the output device 40 can be connected by a bus or other means, Figure 8 The bus connection is taken as an example.
[0173] The input device 30 can receive inputted digital or character information, and generate key signal input related to user settings and function control of the computer device, such as a touch screen, a keypad, a mouse, a trackpad, a touchpad, a pointing stick, one or more mouse buttons, a trackball, a joystick, etc. The output device 40 can include a display device, an auxiliary lighting device (e.g., an LED), a tactile feedback device (e.g., a vibration motor), etc. The display device includes but is not limited to a liquid crystal display, a light-emitting diode, a display, and a plasma display. In some optional embodiments, the display device can be a touch screen.
[0174] The embodiments of the present application also provide a computer readable storage medium, and the method according to the embodiments of the present application can be implemented in hardware, firmware, or recorded in a storage medium, or stored in a remote storage medium or a non-transitory machine readable storage medium downloaded through a network and stored in a local storage medium, so that the method described herein can be processed by such software on a storage medium using a general purpose computer, a special purpose processor, or programmable or special hardware. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only memory, a random access memory, a flash memory, a hard disk or a solid state disk, etc.; further, the storage medium can also include a combination of the above types of memories. It can be understood that the computer, processor, microprocessor controller or programmable hardware includes a storage component that can store or receive software or computer code, which, when accessed and executed by the computer, processor or hardware, implements the method shown in the above embodiments.
[0175] Part of the present application can be applied as a computer program product, for example, computer program instructions, when executed by a computer, the operation of the computer can invoke or provide the method and / or technical solutions according to the present application. Those skilled in the art should understand that the form of computer program instructions in computer readable medium includes but is not limited to source file, executable file, installation package file, etc., and accordingly, the way of computer program instructions executed by computer includes but is not limited to: the computer directly executes the instructions, or the computer executes the corresponding compiled program after compiling the instructions, or the computer reads and executes the instructions, or the computer reads and installs the instructions and then executes the corresponding installed program. Here, the computer readable medium can be any available computer readable storage medium or communication medium accessible to the computer.
[0176] While embodiments of the application have been described in connection with the preferred embodiments of the various figures, those of ordinary skill in the art will appreciate that various modifications and changes can be made without departing from the spirit and scope of the application, and that such modifications and changes fall within the scope of the appended claims.
Claims
1. A human-computer interaction method, characterized in that: The method comprises: Acquire audio data to be played, where the audio data is generated based on text; determining a target syllable corresponding to the audio data according to a frequency spectrum feature corresponding to the audio data; Determining a target mouth shape corresponding to the target syllable based on a syllable mouth shape mapping table, wherein the syllable mouth shape mapping table is used to store a correspondence between syllables and mouth shapes; The audio data is played, and the lip shape of the three-dimensional human face displayed on the target screen is synchronized to the target lip shape.
2. The method according to claim 1, characterized in that Determining the target syllable corresponding to the audio data according to the frequency spectrum features corresponding to the audio data includes: Converting the audio data into first frequency data in a frequency domain; Extracting a frequency domain feature of the first frequency data and an intensity feature of the audio data as the spectrum feature, wherein the frequency domain feature is a core frequency of the audio data, the intensity feature is an amplitude of the sound corresponding to the core frequency, and the core frequency is a frequency value corresponding to a peak point in a spectrum graph of the audio data; The target syllable is searched for in a syllable feature lookup table according to the frequency domain feature and the intensity feature, where the syllable feature lookup table is used to store the correspondence between the syllable and the frequency domain feature and the intensity feature.
3. The method according to claim 2, characterized in that Before searching the target syllable from a syllable feature lookup table according to the frequency domain feature and the intensity feature, the method further includes: Determining, according to the spectrum graph of the second frequency data, a characteristic frequency as the frequency domain feature, wherein the characteristic frequency is a frequency mean of a segment in the spectrum graph having a frequency greater than or equal to a frequency threshold or a frequency peak in the spectrum graph; Acquire the sound intensity corresponding to the characteristic frequency as the intensity feature; The syllable feature lookup table is constructed using the frequency domain features, the intensity features and sample syllables.
4. The method according to claim 3, characterized in that Before determining the target mouth shape corresponding to the target syllable based on the syllable mouth shape mapping table, the method further includes: Acquire an image dataset, wherein the image dataset includes an image of a mouth shape or a mouth corresponding to the sample syllable; Extracting a preset number of mouth feature parameters corresponding to each sample syllable from the image data set; The syllable mouth shape mapping table is constructed based on the characteristic parameters of the mouth and the sample syllable.
5. The method according to any one of claims 1 to 4, characterized in that The determining the target mouth shape corresponding to the target syllable based on the syllable mouth shape mapping table includes: Determining target feature parameters of the mouth corresponding to the target syllable from the syllable mouth shape mapping table; Mapping the target feature parameters into lip shape parameters to be generated in the coordinate system of the three-dimensional face, wherein the lip shape parameters to be generated correspond to the target lip shape; The current lip shape of the three-dimensional human face is rendered as the target lip shape using the lip shape parameters to be generated.
6. The method according to claim 1, characterized in that The step of synchronizing the lip shape of the three-dimensional human face displayed on the target screen to the target lip shape comprises: The lip shape of the three-dimensional human face displayed on the target screen is synchronously switched to the target lip shape.
7. The method according to claim 1, characterized in that The playing of the audio data includes: Playing the audio data using an audio link associated with the target screen, where the audio link represents a connection path of a series of devices and components through which the audio data passes during transmission; The sound signal corresponding to the audio data is sent back to the audio software module or digital signal processing chip associated with the audio link and the target screen by using a loopback path.
8. A human-computer interaction device, characterized in that: The device comprises: An acquisition module, configured to acquire audio data to be played, wherein the audio data is generated based on text; a syllable determination module, configured to determine a target syllable corresponding to the audio data based on a spectral feature corresponding to the audio data; a mouth shape determination module, configured to determine a target mouth shape corresponding to the target syllable based on a syllable mouth shape mapping table, wherein the syllable mouth shape mapping table is configured to store a correspondence between syllables and mouth shapes; The synchronization module is used to play the audio data and synchronize the lip shape of the three-dimensional face displayed on the target screen to the target lip shape.
9. A computer device, characterized in that: include: A memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the human-computer interaction method according to any one of claims 1 to 7 by executing the computer instructions.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a computer to execute the human-computer interaction method according to any one of claims 1 to 7.