A voice generation and expression driving method, client and server

CN115797515BActive Publication Date: 2026-09-25WELLINK TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202211577842.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-09
Publication Date
2026-09-25
Estimated Expiration
2042-12-09

AI Technical Summary

Technical Problem

[0002]随着科技的发展,计算机语音的合成技术越来越成熟,经过计算机技术合成的语音几乎拥有与真人发声一样的语速、音调和发音,通过语音播报几乎可以媲美真人发声,但是由于没有画面与合成语音相结合,只通过合成的语音进行信息传播,可能会导致用户的体验感较低

Benefits of technology

1.在客户端根据音频数据确定出表情数据后,将音频数据和表情数据发送至服务端,以使服务端将表情数据的各个点位与虚拟形象的面部骨骼中的各个点位进行绑定,再经服务端进行驱动后,使客户端可以同时播放音频数据以及虚拟形象所对应的面部表情,由于表情数据是通过音频数据确定的,因此播放的音频数据和虚拟形象的面部表情匹配度较高,也就是说,在播放音频数据时,虚拟形象可以同步显示对应的表情,从而可以便于丰富虚拟形象的面部表情,也便于提升虚拟形象的面部表情播放的流畅度,进而有助于提升用户的体验感。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115797515B_ABST
    Figure CN115797515B_ABST
Patent Text Reader

Abstract

The application relates to the field of data processing, in particular to a voice generation and expression driving method, a client and a server. The method comprises the following steps: acquiring audio data; determining facial expression data of a virtual image corresponding to the audio data, wherein the facial expression data of the virtual image comprises a plurality of point positions; sending the audio data and the facial expression data of the virtual image corresponding to the audio data to the server, so that the server binds each point position of the facial expression data to each point position in a facial skeleton of the virtual image; and outputting the facial expression of the virtual image while playing the audio data based on a driving instruction of the server. The application has the effect of improving user experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing, and in particular to a speech generation and expression-driven method, client, and server. Background Technology

[0002] With the development of technology, computer speech synthesis technology has become more and more mature. Speech synthesized by computer technology has almost the same speech speed, tone and pronunciation as real people. It can be almost comparable to real people when broadcasting. However, since there is no picture to combine with the synthesized speech, and information is transmitted only through synthesized speech, it may result in a lower user experience.

[0003] However, in related technologies, when combining synthesized speech with images, only the duration is considered to match, resulting in a low degree of matching between the person in the image and the synthesized speech. For example, a 10-second image may correspond to a 10-second synthesized speech, but the person in the image has a monotonous expression and only opens and closes their mouth, which may reduce the user's experience. Summary of the Invention

[0004] To address at least one of the above technical problems, embodiments of this application provide a speech generation and expression-driven method, a client, and a server.

[0005] Firstly, this application provides a speech generation and expression-driven method, employing the following technical solution: A speech generation and expression-driven method, executed by a client, includes: Acquire audio data; Determine the facial expression data of the virtual avatar corresponding to the audio data, wherein the facial expression data of the virtual avatar includes multiple points; The audio data and the facial expression data of the virtual image corresponding to the audio data are sent to the server so that the server can bind each point of the facial expression data to each point of the facial skeleton of the virtual image. Based on the server's driver instructions, the virtual avatar's facial expressions are output while the audio data is being played.

[0006] By adopting the above technical solution, after the client determines the facial expression data based on the audio data, the audio data and facial expression data are sent to the server. The server then binds each point of the facial expression data to each point in the facial skeleton of the virtual character. After the server drives the process, the client can simultaneously play the audio data and the corresponding facial expression of the virtual character. Since the facial expression data is determined by the audio data, the matching degree between the played audio data and the facial expression of the virtual character is high. In other words, when playing the audio data, the virtual character can synchronously display the corresponding expression, which can enrich the facial expressions of the virtual character and improve the smoothness of the playback of the virtual character's facial expressions, thereby helping to improve the user experience.

[0007] In one possible implementation, acquiring the audio data includes: Obtain text information and perform sentence segmentation on the text information to obtain a set of sentences; Each statement is parsed to obtain the first emotional information and / or the first voice attribute information corresponding to each statement; Based on the first emotional information and / or the first sound attribute information corresponding to each statement, audio data corresponding to the text information is generated to obtain the audio data.

[0008] By adopting the above technical solution, when converting text information into audio data, the emotional and vocal attributes of each sentence in the text information are analyzed. Furthermore, audio data is generated based on the emotional and vocal characteristics of each sentence, which makes the generated audio data more consistent with the real situation, helps to improve the authenticity of the audio data, so that the audio can rival the voice of a real person, and further improves the user experience.

[0009] In one possible implementation, the step of parsing each statement to obtain the first emotional information and / or first voice attribute information corresponding to each statement further includes: To obtain information about virtual avatars and / or secondary emotional information input by users; The step of parsing each statement to obtain the first emotional information and / or first voice attribute information corresponding to each statement includes: Based on the virtual avatar information and / or the second emotional information input by the user, each statement is parsed to obtain the first emotional information and / or the first voice attribute information corresponding to each statement.

[0010] By adopting the above technical solution, and by acquiring virtual avatar information and / or, user-inputted second emotional information, the virtual avatar information and / or, user-inputted second emotional information can be considered when parsing each sentence in the text information. This allows for the simultaneous acquisition of the first emotional information and / or first voice attribute information corresponding to each sentence, while also considering the context of each sentence and the virtual avatar and / or, user-inputted second emotional information. This improves the accuracy of sentence analysis in obtaining the first emotional information and / or first voice attribute information corresponding to each sentence, thereby further enhancing the user experience.

[0011] In one possible implementation, determining the facial expression data of the virtual avatar corresponding to the audio data includes: The audio data is segmented into sentences, and the third emotional information and / or second sound attribute information corresponding to each audio sentence are obtained; Based on the third emotional information and / or second voice attribute information corresponding to each audio statement, the facial expression data of the virtual image corresponding to each audio statement is determined.

[0012] By adopting the above technical solution, when determining the facial expression data corresponding to each audio sentence in the audio data, the facial expression corresponding to the pronunciation of each word can be enriched by the third emotional information and / or second sound attribute information of each audio sentence, rather than determining the facial expression of the virtual image solely by the pronunciation of the words. Instead, the third emotional information and / or second sound attribute information corresponding to each audio sentence are further considered to determine the facial expression data of the virtual image corresponding to each audio sentence. Thus, the facial expression data of the virtual image when reading each audio sentence can simultaneously conform to the third emotional information and / or second sound attribute information corresponding to each audio sentence, thereby making the facial expression data of the virtual image more accurate and further improving the user experience.

[0013] In one possible implementation, determining the facial expression data of the virtual avatar corresponding to the audio data includes: Determine the mouth movement information corresponding to the audio data; Based on the mouth movement information corresponding to the audio data, determine the movement information of other parts of the face; Based on the mouth movement information corresponding to the audio data and the movement information of other parts of the face, the various points of the facial expression data are determined to obtain the facial expression data of the virtual image corresponding to the audio data.

[0014] By adopting the above technical solution, since the positions of the mouth and other parts remain fixed, when the mouth of the virtual image moves, the other parts will also change with the movement of the mouth. After determining the movement information of the mouth, the movement information of other parts can also be determined. By determining the facial expression of the virtual image through the position between the mouth and other parts, it is easier to improve the smoothness of the displacement changes of various parts of the virtual image's face, and thus improve the accuracy of determining facial expression data.

[0015] In one possible implementation, the step of sending the audio data and the facial expression data of the virtual avatar corresponding to the audio data to the server further includes: The audio data and the facial expression data of the corresponding virtual avatar are used to generate a data stream through a specific network protocol; The step of sending the audio data and the facial expression data of the virtual avatar corresponding to the audio data to the server includes: The data stream is sent to the server.

[0016] By adopting the above technical solution, audio data and facial expression data are generated into a data stream by simulating a specific network protocol, and the data stream is transmitted to the server to realize free data transmission between the client and the server, so that the server can drive the facial expressions of the virtual character.

[0017] Secondly, this application also provides another method for speech generation and expression-driven processing, employing the following technical solution: A speech generation and expression-driven method, executed by the server, includes: The audio data and the facial expression data of the virtual avatar corresponding to the audio data are obtained, wherein the facial expression data of the virtual avatar includes multiple points; The facial expression data points are bound to the facial bones of the virtual image. Based on the binding relationship, the client is controlled to drive the facial expressions of the virtual avatar while playing audio data.

[0018] By adopting the above technical solution, facial expression data is bound to the facial bones of the virtual character. After being driven, the client can simultaneously play audio data and the facial expressions corresponding to the virtual character. Since the facial expression data is determined by the audio data, the matching degree between the played audio data and the facial expressions of the virtual character is high. Furthermore, by binding the facial expression data to various points in the facial bones of the virtual character, it is easier to enrich the facial expressions of the virtual character and improve the smoothness of the playback of the facial expressions of the virtual character, thereby helping to improve the user experience.

[0019] In one possible implementation, the binding of individual points of facial expression data to individual points in the facial skeleton of the virtual avatar further includes: Determine the fourth emotional information and / or third voice attribute information corresponding to each statement in the audio data; The process of binding each point of facial expression data to each point in the facial skeleton of the virtual avatar further includes: Based on the fourth emotional information and / or third voice attribute information corresponding to each statement and the binding relationship, the displacement information of each point in the facial skeleton corresponding to the virtual image is determined. The step of controlling the client to drive the facial expressions of the virtual avatar while playing audio data, based on the binding relationship, includes: Based on the binding relationship and the displacement information of each point in the facial skeleton, the client is controlled to drive the facial expressions of the virtual character while playing audio data.

[0020] By adopting the above technical solution, facial expression data corresponding to the audio data is determined by identifying the fourth emotional information and / or the third sound attribute information corresponding to the audio data. Then, by binding the facial expression data to the facial bones of the virtual image, it is easy to determine the facial expression corresponding to the audio data. Furthermore, the virtual image is driven by the audio data so that the corresponding facial expression data can be played simultaneously when the audio data is played. Since the facial expression data is determined by the audio data and the facial expression data is bound to the facial bones of the virtual image, it helps to improve the accuracy of audio-visual correspondence and achieve audio-visual synchronization.

[0021] In one possible implementation, the binding of individual points of facial expression data to individual points in the facial skeleton of the virtual avatar further includes: Obtain image information containing the virtual image; Based on the image information containing the virtual image, identify the various points of the facial bones of the virtual image; The step of binding each point of facial expression data with each point in the facial skeleton of the virtual image includes: The facial expression data points are bound to the facial bones of the identified virtual image.

[0022] By adopting the above technical solution, the points of the virtual image's facial bones are obtained by identifying image information containing the virtual image, and the points of the facial expression data are bound to the points of the identified facial bones. This makes the binding accuracy between the points of the facial expression data and the points of the virtual image's facial bones higher, which in turn makes the corresponding facial expressions of the virtual image more accurate when playing audio data, and further improves the user experience.

[0023] Thirdly, this application provides a speech generation and expression-driven device from the client's perspective, employing the following technical solution: A speech generation and expression-driven device, comprising: The audio data acquisition module is used to acquire audio data; The facial expression data determination module is used to determine the facial expression data of the virtual character corresponding to the audio data, wherein the facial expression data of the virtual character includes multiple points; The data sending module is used to send the audio data and the facial expression data of the virtual image corresponding to the audio data to the server, so that the server can bind each point of the facial expression data to each point of the facial skeleton of the virtual image. The execution driver module is used to output the facial expressions of the virtual avatar while playing the audio data, based on the driver instructions from the server.

[0024] In one possible implementation, the audio data acquisition module, when acquiring audio data, specifically performs the following functions: Obtain text information and perform sentence segmentation on the text information to obtain a set of sentences; Each statement is parsed to obtain the first emotional information and / or the first voice attribute information corresponding to each statement; Based on the first emotional information and / or first sound attribute information corresponding to each statement, audio data corresponding to the text information is generated to obtain the audio data.

[0025] In one possible implementation, the device further includes: The information acquisition module is used to acquire information about the virtual avatar and / or secondary emotional information input by the user; Specifically, the audio data acquisition module, when parsing each statement to obtain the first emotional information and / or first sound attribute information corresponding to each statement, is used for: Based on the virtual avatar information and / or the second emotional information input by the user, each statement is parsed to obtain the first emotional information and / or the first voice attribute information corresponding to each statement.

[0026] In one possible implementation, when determining the facial expression data of the virtual avatar corresponding to the audio data, the expression data determination module is specifically used for: The audio data is segmented into sentences, and the third emotional information and / or second sound attribute information corresponding to each audio sentence are obtained; Based on the third emotional information and / or second voice attribute information corresponding to each audio statement, the facial expression data of the virtual image corresponding to each audio statement is determined.

[0027] In one possible implementation, when determining the facial expression data of the virtual avatar corresponding to the audio data, the expression data determination module is specifically used for: Determine the mouth movement information corresponding to the audio data; Based on the mouth movement information corresponding to the audio data, determine the movement information of other parts of the face; Based on the mouth movement information corresponding to the audio data and the movement information of other parts of the face, the various points of the facial expression data are determined to obtain the facial expression data of the virtual image corresponding to the audio data.

[0028] In one possible implementation, the device further includes: The data stream generation module is used to generate a data stream from the audio data and the facial expression data of the virtual image corresponding to the audio data through a specific network protocol; Specifically, when the data sending module sends the audio data and the facial expression data of the virtual avatar corresponding to the audio data to the server, it is used for: The data stream is sent to the server.

[0029] Fourthly, this application provides a speech generation and expression-driven device from a server-side perspective, employing the following technical solution: A speech generation and expression-driven device, comprising: The data information module is used to acquire the audio data and the facial expression data of the virtual image corresponding to the audio data, wherein the facial expression data of the virtual image includes multiple points; The point binding module is used to bind each point of the facial expression data to each point in the facial skeleton of the virtual image; The control driver module is used to control the client to drive the facial expressions of the virtual avatar while playing audio data, based on the binding relationship.

[0030] In one possible implementation, the device further includes: The information determination module is used to determine the fourth emotional information and / or the third voice attribute information corresponding to each statement in the audio data. The device also includes: The displacement information determination module is used to determine the displacement information of each point in the facial skeleton corresponding to the virtual image based on the fourth emotional information and / or third voice attribute information corresponding to each statement and the binding relationship. The control driver module, based on the binding relationship, controls the client to drive the facial expressions of the virtual avatar while playing audio data. Specifically, it is used for: Based on the binding relationship and the displacement information of each point in the facial skeleton, the client is controlled to drive the facial expressions of the virtual character while playing audio data.

[0031] In one possible implementation, the device further includes: An image information acquisition module is used to acquire image information containing the virtual image; The identification point module is used to identify various points of the facial bones of the virtual image based on the image information containing the virtual image; Specifically, the point binding module, when binding each point of facial expression data to each point in the facial skeleton of the virtual image, is used for: The facial expression data points are bound to the facial bones of the identified virtual image.

[0032] Fifthly, this application provides a client application that adopts the following technical solution: A client comprising: At least one processor; Memory; At least one application, wherein the at least one application is stored in memory and configured to be executed by at least one processor, the at least one application being configured to: perform the speech generation and expression-driven methods described in the first aspect above.

[0033] Sixthly, this application provides a server-side solution, which adopts the following technical solution: A server-side component, comprising: At least one processor; Memory; At least one application, wherein the at least one application is stored in memory and configured to be executed by at least one processor, the at least one application being configured to: perform the speech generation and expression-driven methods described in the second aspect above.

[0034] Seventhly, this application provides a computer-readable storage medium, which adopts the following technical solution: A computer-readable storage medium includes: a computer program that can be loaded by a processor and execute the speech generation and expression-driven methods described in the first and / or second aspects above.

[0035] In summary, this application includes at least one of the following beneficial technical effects: 1. After the client determines the facial expression data based on the audio data, the audio data and facial expression data are sent to the server. The server then binds each point of the facial expression data to each point in the facial skeleton of the virtual character. After the server drives the process, the client can simultaneously play the audio data and the corresponding facial expression of the virtual character. Since the facial expression data is determined by the audio data, the matching degree between the played audio data and the virtual character's facial expression is high. In other words, when playing the audio data, the virtual character can synchronously display the corresponding expression, which can enrich the virtual character's facial expressions and improve the smoothness of the playback of the virtual character's facial expressions, thereby helping to improve the user experience.

[0036] 2. By binding facial expression data to the facial skeleton of the virtual avatar and then driving it, the client can simultaneously play audio data and the corresponding facial expressions of the virtual avatar. Since the facial expression data is determined by the audio data, the matching degree between the played audio data and the facial expressions of the virtual avatar is high. Furthermore, by binding the facial expression data to various points in the facial skeleton of the virtual avatar, it is easier to enrich the facial expressions of the virtual avatar and improve the smoothness of the playback of the facial expressions of the virtual avatar, thereby helping to improve the user experience. Attached Figure Description

[0037] Figure 1 This is a flowchart illustrating a speech generation and expression-driven method according to an embodiment of this application; Figure 2 This is a flowchart illustrating another speech generation and expression-driven method in an embodiment of this application; Figure 3 This is a scenario diagram of client-server interaction in an embodiment of this application; Figure 4 This is a schematic diagram of the device results for voice generation and expression driving from a client perspective in an embodiment of this application; Figure 5 This is a schematic diagram of the device results for voice generation and expression driving from a server-side perspective in an embodiment of this application; Figure 6 This is a schematic diagram of the structure of a client in an embodiment of this application. Detailed Implementation

[0038] The following is in conjunction with the appendix Figure 1-6 This application will be described in further detail.

[0039] After reading this specification, those skilled in the art may make modifications to this embodiment without contributing any inventive step, but such modifications are protected by patent law as long as they fall within the scope of the claims of this application.

[0040] Since the birth of the first computer-based speech synthesis system, to the dominance of statistical machine learning speech synthesis represented by Hidden Markov Models, and then to the rapid development of neural network speech synthesis, computer speech synthesis technology can now rival human voices and is moving towards large-scale commercialization. AI speech can automatically generate natural speech that matches the tone and emotion of a human voice based on custom text, supporting detailed customization and easily adjusting speech rate, pitch, pronunciation, and pauses to optimize speech output. However, currently, the audio and facial expressions output in this field lack corresponding hyper-realistic virtual avatars, limiting the application scenarios of this technology.

[0041] To address the aforementioned technical issues, this application provides a voice generation and expression-driven method that integrates AI voice generation technology and binds it to hyper-realistic digital avatars, satisfying scenarios such as film and television, advertising, live streaming, and virtual shooting. To make the objectives, technical solutions, and advantages of this application's embodiments clearer, the technical solutions in this application's embodiments will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0042] Furthermore, the term "and / or" in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. Additionally, the character " / " in this article, unless otherwise specified, generally indicates that the preceding and following related objects have an "or" relationship.

[0043] This application provides a voice generation and expression-driven method, which is executed by a client. The client can be a terminal device, such as a smartphone, tablet, laptop, or desktop computer, but is not limited to these. This application does not impose any restrictions on this method.

[0044] refer to Figure 1 , Figure 1This is a flowchart illustrating a speech generation and expression-driven method according to an embodiment of this application, executed by a client. The method may include steps S110, S120, S130, and S140, wherein: Step S110: Obtain audio data.

[0045] Specifically, the audio data is digitized sound data, containing speech information such as tone and emotion of the human voice. The audio data can be input by the user, automatically crawled from relevant web pages by the client using web crawler technology, obtained from local storage by the client, or obtained by converting text information into audio. In this embodiment of the application, no limitation is made.

[0046] Step S120: Determine the facial expression data of the virtual character corresponding to the audio data. The facial expression data of the virtual character includes multiple points.

[0047] Specifically, the virtual avatar can be anime characters, cartoon characters, or virtual figures, etc. The specific virtual avatar can be selected and determined by the user, or it can be pre-specified by the system. Facial expression data can be facial expression data containing multiple points; for example, facial expression data can contain 55 points, as shown in Table 1. Table 1 1 Left eye blink EyeBlinkLeft 2 Left eye looking down EyeLookDownLeft 3 Left eye focused on the tip of the nose EyeLookInLeft 4 Look to the left with your left eye EyeLookOutLeft 5 Left eye looking upwards EyeLookUpLeft 6 Squinting with left eye EyeSquintLeft 7 Left eye wide open EyeWideLeft 8 Right eye blink EyeBlinkRight 9 Looking down with right eye EyeLookDownRight 10 Right eye looking at the tip of the nose EyeLookInRight 11 Looking to the left with the right eye EyeLookOutRight 12 Looking upwards with the right eye EyeLookUpRight 13 Squinting with right eye EyeSquintRight 14 Right eye wide open EyeWideRight 15 When pursing your lips, your chin moves forward. JawForward 16 When pouting, the chin moves to the left. JawLeft 17 When pouting, the chin moves to the right. JawRight 18 When you open your mouth, your chin should point downwards. JawOpen 19 Shut up MouthClose 20 Slightly open your mouth and part your lips. MouthFunnel 21 pursed lips MouthPucker 22 Pouting to the left MouthLeft 23 Pouting to the right MouthRight 24 Left-handed smile MouthSmileLeft 25 Right-handed smile MouthSmileRight 26 Press down on the left lip MouthFrownLeft 27 Press down on the right lip MouthFrownRight 28 Left lip backward MouthDimpleLeft 29 Right lip backward MouthDimpleRight 30 Left corner of the mouth to the left MouthStretchLeft 31 Right corner of the mouth to the right MouthStretchRight 32 The lower lip curls inward MouthRollLower 33 The lower lip curls upwards MouthRollUpper 34 Lower lip down MouthShrugLower 35 upper lip up MouthShrugUpper 36 Press the lower lip to the left MouthPressLeft 37 Press the lower lip to the right MouthPressRight 38 Press the lower lip down to the left MouthLowerDownLeft 39 Press the lower lip down to the right MouthLowerDownRight 40 Press the upper lip towards the upper left MouthUpperUpLeft 41 Press the upper lip towards the upper right MouthUpperUpRight 42 Left eyebrow outward BrowDownLeft 43 Right eyebrow outward BrowDownRight 44 Frowning BrowInnerUp 45 Left eyebrow to the upper left BrowOuterUpLeft 46 Right eyebrow to the upper right BrowOuterUpRight 47 cheeks outward CheekPuff 48 Left cheek up and rotate CheekSquintLeft 49 Right cheek up and rotate CheekSquintRight 50 Left wrinkled nose NoseSneerLeft 51 Right wrinkle nose NoseSneerRight 52 sticking out tongue TongueOut 53 Turning head HeadRoll 54 Turn your left eye LeftEyeRoll 55 Turn right eye RightEyeRoll Each data point is represented by a value between 0 and 1. The magnitude of the value represents the range of motion at each point. For example, the data point could represent the range of motion of the left eye blinking, the range of motion of the left eye looking down, or the range of motion of the left eye looking at the tip of the nose. When the data point is 0, it indicates that no change in expression has occurred.

[0048] When a virtual avatar makes a voice broadcast, different words correspond to different sounds, and the mouth will change accordingly. Furthermore, since the facial bones are interconnected, the movement of the mouth bones will also cause changes in the bones of other parts. Therefore, when determining facial expression data, the movement state of the virtual avatar's mouth can be determined first based on the audio data, and then the movement state of other parts can be determined based on the movement state of the mouth to form facial expression data.

[0049] Step S130: Send audio data and the facial expression data of the virtual character corresponding to the audio data to the server so that the server can bind each point of the facial expression data to each point in the facial skeleton of the virtual character.

[0050] Specifically, the client and server can exchange information through data transmission. When the client sends audio data and facial expression data to the server, it can first package the audio data and facial expression data according to the transmission protocol, and then transmit the packaged data to the server so that the server can receive the complete audio data and complete facial expression data. After receiving the facial expression data, the server binds the facial expression data to the facial bone points of the virtual image.

[0051] Step S140: Based on the server-side driver instructions, output the facial expressions of the virtual avatar while playing audio data.

[0052] Specifically, the driver commands are used to control the simultaneous playback of audio data and the virtual avatar's facial expressions on the client side, so that the audio and facial expressions correspond when the virtual avatar makes a voice broadcast.

[0053] In this embodiment of the application, after the client determines the facial expression data based on the audio data, the audio data and facial expression data are sent to the server. The server then binds each point of the facial expression data to each point in the facial skeleton of the virtual character. After the server drives the process, the client can simultaneously play the audio data and the facial expression corresponding to the virtual character. Since the facial expression data is determined by the audio data, the matching degree between the played audio data and the facial expression of the virtual character is high. In other words, when playing the audio data, the virtual character can synchronously display the corresponding expression, which can enrich the facial expressions of the virtual character and improve the smoothness of the playback of the virtual character's facial expressions, thereby helping to improve the user experience.

[0054] Specifically, in this embodiment of the application, obtaining audio data in step S110 may include steps S1101 (not shown in the figures), S1102 (not shown in the figures), and S1103 (not shown in the figures), wherein: Step S1101: Obtain text information and perform sentence segmentation on the text information to obtain a set of sentences.

[0055] Specifically, the text information can be content to be broadcast, policy announcements, or product introductions; the specific content of the text information is not specifically limited in this embodiment. The text information can be obtained by user input, by retrieving pre-stored content, or by scraping from a webpage.

[0056] After acquiring the text information, it is processed by sentence segmentation, that is, the acquired text information is segmented into sentences. Sentence segmentation can be performed by recognizing punctuation marks, by semantic recognition, and then using the semantic recognition results for sentence segmentation. Regular expressions can also be used for sentence segmentation. The specific sentence segmentation method is not limited in this embodiment. Furthermore, after sentence segmentation, multiple segmented sentences are obtained to form a sentence set.

[0057] Step S1102: Parse each statement to obtain the first emotional information and / or the first voice attribute information corresponding to each statement.

[0058] Specifically, the first emotional information can be happiness or anger, which can express real human emotions. The first sound attribute information includes tone, intonation, and amplitude when speaking. Among them, tone is the sound form of a specific sentence under the control of certain thoughts and feelings; intonation is the tone of speech, that is, the configuration and changes of speed and emphasis in a sentence; amplitude is the range of sound.

[0059] For example, when the statement is "Thank you", parsing the statement will likely yield the emotional information of happiness, and the corresponding first sound attribute information will show a relatively gentle tone and an amplitude of 10dB; when the statement is "Why do you always make mistakes", the first emotional information will likely show the emotional information of anger, and the corresponding sound attribute information will show a relatively urgent tone and an amplitude of 100dB.

[0060] To further refine the converted audio data to better suit user needs, specifically to better align with the virtual image and / or emotional information required by the user, each sentence is parsed to obtain the first emotional information and / or first sound attribute information corresponding to each sentence. This may further include: Obtain information about the virtual avatar and / or secondary emotional information input by the user.

[0061] In this embodiment of the application, the virtual avatar information can be input by the user, or input by the user through a selection operation, or it can be preset. For example, the virtual avatar information can be female avatar information, or male avatar information, or some animal avatar information.

[0062] Furthermore, the virtual avatar information can also include the avatar's vocal information, including its tone, frequency, and vocal habits. This information helps determine the avatar's personality, and the corresponding audio data for the text is determined based on the avatar's personality traits, making the audio data more closely match the virtual avatar. The vocal requirements for each sentence can also be input by the user. When inputting vocal requirements, the user can specify criteria such as gender, age, emotion, state, and personality to ensure the generated audio data better meets the user's needs. Gender can be male or female; age can be old, middle-aged, young, or infant; emotion can be joy, anger, sorrow, or happiness; state can include talking while walking or talking while running; and personality can be soft-spoken or gruff.

[0063] Furthermore, after obtaining the virtual avatar information and / or the second emotional information input by the user, each statement is parsed to obtain the emotional information and / or voice attribute information corresponding to each statement, which may specifically include: Based on virtual avatar information and / or second emotional information input by the user, each statement is parsed to obtain the first emotional information and / or first voice attribute information corresponding to each statement.

[0064] Specifically, each statement is parsed based on the virtual avatar and / or the second emotional information input by the user. That is, the statement parsing results are optimized according to the virtual avatar and / or the second emotional information input by the user. Since the initial emotional information and / or initial voice attribute information corresponding to each statement can be obtained by parsing each statement, the final emotional information and / or voice attribute information is obtained after optimizing the initial emotional information and / or initial voice attribute information based on the vocal requirements input by the virtual avatar and / or the user.

[0065] After parsing the sentence that has undergone sentence segmentation, the first emotional information obtained may be the same as or different from the second emotional information input by the user. For example, when the sentence is "Why do you always make mistakes?", parsing the sentence may determine that the first emotional information corresponding to the sentence is anger, but the second emotional information input by the user is happiness. The first emotional information is different from the second emotional information. In this case, in order to improve the authenticity of the sentence after it is converted into audio, the context can be parsed to determine the current context of the sentence, and the final emotional information can be determined based on the context. The final emotional information is then determined as the first emotional information.

[0066] Step S1103: Based on the first emotional information and / or the first sound attribute information corresponding to each statement, generate the audio data corresponding to the text information to obtain the audio data.

[0067] Specifically, when generating audio data from sentences, the sentences can be converted into audio using text-to-speech (TTS) tools or other text-to-audio methods. During the conversion process, the first emotional information and / or first sound attribute information corresponding to each sentence are considered so that each converted audio data has the corresponding emotional information and / or sound attribute information.

[0068] Further, after obtaining the audio data, the facial expression data of the virtual avatar corresponding to the audio data is determined, which may specifically include steps S1201 (not shown in the figures) and S1202 (not shown in the figures), wherein: Step S1201: Segment the audio data and obtain the third emotional information and / or second sound attribute information corresponding to each audio sentence.

[0069] Specifically, when determining each audio sentence from the audio data, speech recognition is performed to obtain pause positions in the audio data. The audio data is then segmented into sentences based on these pause positions. Semantic recognition is then performed on each sentence to determine the corresponding third emotional information and / or second voice attribute information. The first emotional information and first voice attribute information are obtained by segmenting the text into sentences and parsing each segmented sentence. The second emotional information is input by the user and may be the same as or different from the first emotional information. Since the audio data is generated from the determined first emotional information and first voice attribute information, if the audio data is converted from the aforementioned text information, then the third emotional information is the same as the first emotional information, and the second voice attribute information is the same as the first voice attribute information. If the audio data is not converted from the aforementioned text information, then the third emotional information may be different from the first emotional information, and the second voice attribute information may also be different from the first voice attribute information.

[0070] In addition, if the audio data is obtained by converting text information to audio, the text information is segmented into sentences, and after determining the first emotional information and / or first sound attribute information corresponding to the multiple sentences obtained after the segmentation, the first emotional information and / or first sound attribute information corresponding to each sentence is bound to that sentence. When the sentence is converted into audio data, the first emotional information and / or first sound attribute information corresponding to each sentence can also be bound to the audio sentence corresponding to each sentence. When obtaining the emotional information and / or sound attribute information corresponding to each audio sentence, the third emotional information and / or second sound attribute information corresponding to each audio sentence can be obtained.

[0071] Step S1202: Based on the third emotional information and / or second voice attribute information corresponding to each audio statement, determine the facial expression data of the virtual image corresponding to each audio statement.

[0072] Specifically, facial expression data refers to the facial expressions corresponding to the virtual avatar's recitation of audio data. Since each word corresponds to a different pronunciation, the contraction and relaxation of facial muscles differ depending on the word. For example, when the word "ah" is recited, the mouth opens and facial muscles relax. However, different emotional information and / or vocal attributes result in different facial expressions. For instance, when the third emotional information is "happy," the facial expression data includes not only the mouth opening and facial muscle relaxation corresponding to the word's pronunciation, but also facial muscle relaxation and an upward-turned corner of the mouth. When the emotional information is "angry," the facial expression data includes not only the mouth opening and facial muscle relaxation corresponding to the word's pronunciation, but also facial muscle contraction and a downward-turned corner of the mouth. By using the facial expression data corresponding to each word, the facial expression data corresponding to each audio statement is determined.

[0073] In this embodiment of the application, the facial expressions corresponding to the pronunciation of each character are enriched by the third emotional information and / or the second sound attribute information of each audio statement, rather than determining the facial expressions of the virtual character solely by the pronunciation of the characters. The use of emotional information and / or sound attribute information facilitates the increase in the number of facial expressions, thereby making the facial expressions more diverse.

[0074] Furthermore, since a sentence may correspond to multiple expressions, and the mouth state is different when each word is spoken, the mouth state affects the state of other parts of the face. Specifically, determining the facial expression data of the virtual avatar corresponding to the audio data may include steps S1 (not shown in the attached figure), S2 (not shown in the attached figure), and S3 (not shown in the attached figure), wherein: Step S1: Determine the mouth movement information corresponding to the audio data.

[0075] Specifically, mouth movement information can include the displacement of the mouth. When determining the mouth movement information corresponding to audio data, the mouth movement process corresponding to each character can be determined based on the pronunciation rules after identifying the characters in the audio data. The pronunciation rules refer to the mouth shape corresponding to each character when it is pronounced correctly. The mouth movement images before and after each character is pronounced are recorded. The boundary points of the mouth features in the pre-pronunciation and post-pronunciation mouth movement images are determined and labeled, forming pre-pronunciation and post-pronunciation boundary point images. These images are then imported into a pre-established coordinate system to determine the coordinates of each boundary point in both the pre-pronunciation and post-pronunciation boundary point images, and the displacement of each boundary point is also determined to obtain the mouth movement information corresponding to each character.

[0076] Furthermore, the mouth movement information corresponding to the audio data can be determined through a network model. Other methods for determining mouth movement information corresponding to audio data may also be included in this embodiment, and are not limited thereto.

[0077] Step S2: Based on the mouth movement information corresponding to the audio data, determine the movement information of other parts of the face.

[0078] Specifically, facial expressions refer to the expression of various emotional states through changes in the eye muscles, facial muscles, and mouth muscles. The direction of the muscles is also the movement of the bones. The movement of the eye muscles corresponds to the movement of the eye bones, and the movement of the mouth muscles corresponds to the movement of the mouth bones. Since the muscle groups near the eyes and mouth are the most expressive parts of the face, and the mouth shape corresponding to each word is different during broadcasting, the movement information of other parts can be determined by first determining the mouth movement information of the virtual image, and then determining the movement information of other parts based on the determined mouth movement information.

[0079] When determining the motion information of other parts, the fixed distance between the mouth and other parts can be determined by the mouth features. When the mouth changes, the displacement of other parts can be determined by the displacement of the mouth and the fixed distance between the mouth and other parts.

[0080] Step S3: Based on the mouth movement information and other facial movement information corresponding to the audio data, determine the various points of the facial expression data to obtain the facial expression data of the virtual image corresponding to the audio data.

[0081] Specifically, because there are connections between facial bones, when the bones of the mouth move due to vocalization, the bones of other parts of the face will also change accordingly. To determine the changes in other parts based on the changes in the bones of the mouth, the amplitude of the changes in the bones of the mouth and the direction of the movement can be recorded. By measuring the distance between the bones, the amplitude and direction of the changes in other bones can be determined.

[0082] Each point in the facial expression data is used to represent the facial expressions of the virtual character. For example, when a word in the audio data is read, the corresponding facial expression may be blinking the left eye, blinking the right eye, pressing down the left lip, pressing down the right lip, etc. Each word in the audio data corresponds to at least one facial expression. Based on the facial expression data corresponding to multiple words, it is easy to determine the facial expression data of the virtual character corresponding to the audio data.

[0083] In this embodiment of the application, since the positions of the mouth and other parts are fixed, when the mouth of the virtual image moves, the other parts will also change with the movement of the mouth. The facial expression of the virtual image is determined by the position between the mouth and other parts, which makes it easier to improve the smoothness when the various parts of the virtual image's face change displacement, and also makes it easier to improve the accuracy of determining facial expression data.

[0084] Further, after obtaining the audio data and the corresponding facial expression data, step Sa (not shown in the figure) is executed, or the audio data and the facial expression data of the virtual avatar corresponding to the audio data are sent to the server, before executing step Sa (not shown in the attached figure), wherein: Step Sa: Generate a data stream by combining the audio data and the facial expression data of the corresponding virtual avatar through a specific network protocol.

[0085] Specifically, a network protocol is a set of rules, standards, or conventions established for data exchange in a computer network. For example, a client and a server in a network need to communicate, but because the character sets used by the client and the server may be different, the server may not be able to recognize the audio data and facial expression data after the client transmits them to the server. In order to enable communication between the client and the server, it can be stipulated that both the client and the server first transform the characters in their respective character sets into characters in a standard character set before communicating. For example, before the client needs to transmit audio data and facial expression data to the server, it can simulate this specific network protocol to generate the corresponding data stream.

[0086] For example, the data format of the client before sending may be inconsistent with the data format that the server can receive. The client has 55 data points before sending, while the server has 52 data points. Therefore, it is necessary to convert the data format before sending according to a specific network protocol to form a data stream before forwarding.

[0087] Furthermore, through the above embodiments, a data stream is generated from the audio data and the facial expression data of the virtual avatar corresponding to the audio data. Therefore, sending the audio data and the facial expression data of the virtual avatar corresponding to the audio data to the server can specifically include: sending the data stream to the server.

[0088] In this embodiment of the application, a data stream is generated by simulating a specific network protocol to generate audio data and facial expression data, and the data stream is transmitted to the server to realize free data transmission between the client and the server, so that the server can drive the facial expressions of the virtual character.

[0089] refer to Figure 2 , Figure 2This is a flowchart illustrating another speech generation and expression-driven method in this application embodiment, executed by the server. The method may include steps S210, S220, and S230, wherein: Step S210: Obtain audio data and the facial expression data of the virtual character corresponding to the audio data. The facial expression data of the virtual character includes multiple points.

[0090] In this embodiment of the application, the server receives the data stream sent by the client, and obtains the audio data and the facial expression data of the virtual image corresponding to the audio data by parsing the data stream.

[0091] Step S220: Bind each point of the facial expression data to each point in the facial skeleton of the virtual image.

[0092] Specifically, the process involves acquiring a facial image of a virtual avatar, performing skeletal point recognition on the image, and determining the various points of the virtual avatar's facial bones. These virtual avatar facial bone points are pre-defined points on the virtual character's facial bone model that correspond to the positions of the user's facial points. The real-time coordinates of the user's facial feature points can be converted into the real-time coordinates of the virtual avatar's facial bone feature points by establishing corresponding coordinate systems for both the user's face and the virtual avatar's facial bone model, thus obtaining the positional data of the virtual character's facial bone feature points. This application does not impose any limitations on this process.

[0093] Since facial expression data contains multiple points, binding each point in the facial expression data to the corresponding points in the facial skeleton of the virtual character is simple. For example, in the XXX engine, the Live Link plugin can provide a universal interface to add facial expression data to the XXX engine, so that each point in the facial expression data is bound to the corresponding point in the facial skeleton of the virtual character. The XXX engine can consist of multiple animation tools and editors, including an animation blueprint editor, an animation editor, a skeleton mesh editor, a skeleton editor, and a pre-animation tool, used to play and blend pre-prepared facial expression data, directly controlling the deformation of the virtual character's skeleton to bind each point in the facial expression data to the corresponding point in the facial skeleton of the virtual character, making the facial expressions of the virtual character more realistic after driving the virtual character.

[0094] Step S230: Based on the binding relationship, control the client to drive the facial expressions of the virtual character while playing audio data.

[0095] Specifically, the binding relationship refers to the correspondence between points in the facial expression data and points in the facial bones of the virtual avatar. Since the facial expression data is determined based on the audio data, the virtual avatar's facial expressions are driven while the audio data is played, facilitating audio-visual synchronization.

[0096] In this embodiment of the application, the facial expression data is bound to the facial skeleton of the virtual character, and then driven so that the client can play the audio data and the facial expression corresponding to the virtual character at the same time. By binding the facial expression data to each point in the facial skeleton of the virtual character, it is easier to enrich the facial expressions of the virtual character and improve the smoothness of the playback of the facial expressions of the virtual character, thereby helping to improve the user experience.

[0097] Furthermore, the various points of the facial expression data are bound to the various points in the facial skeleton of the virtual avatar. This process includes step S2a (not shown in the attached diagram), in which: Step S2a: Determine the fourth emotional information and / or third voice attribute information corresponding to each statement in the audio data.

[0098] In the embodiments of this application, the fourth emotional information and / or third sound attribute information corresponding to each statement in the audio data can be carried in the audio data sent by the client. That is, the fourth emotional information can be the first emotional information or the third emotional information shown in the above embodiments, and the third sound attribute information can be the first sound attribute information or the second sound attribute information shown in the above embodiments. Of course, it can also be obtained by the server parsing the received audio data. The method by which the server parses the audio data to obtain the fourth emotional information and / or third sound attribute information corresponding to each statement can be found in the above embodiments and will not be repeated here.

[0099] Further, step S220 binds each point of the facial expression data to each point in the facial skeleton of the virtual avatar, followed by step S2b (not shown in the attached figure), wherein: In step S2b, the displacement information of each point in the facial skeleton corresponding to the virtual image is determined based on the fourth emotional information and / or the third voice attribute information corresponding to each statement and the binding relationship.

[0100] Specifically, utilizing the fourth emotional information and / or third voice attribute information corresponding to each statement may affect the displacement information of various points in the facial skeleton corresponding to the virtual image. In this embodiment, after obtaining the emotional information and / or voice attribute information corresponding to each statement, the displacement information of various points in the facial skeleton corresponding to the virtual image can be determined by a network model, or by other methods; this embodiment is not limited to any particular method.

[0101] Furthermore, after obtaining the binding relationship and the displacement information of each point in the facial skeleton corresponding to the virtual image through the above embodiments, based on the binding relationship, the client is controlled to drive the facial expressions of the virtual image while playing audio data. Specifically, this may include: Based on the binding relationship and the displacement information of various points in the facial skeleton, the client controls the facial expressions of the virtual character while playing audio data.

[0102] Specifically, for example, the total length of the audio data is 30 seconds, and there are 1,800 facial expression data corresponding to the audio data. Each facial expression data has a corresponding point. 60 facial expression data are played per second. When playing the audio data, the audio data and facial expression data can correspond, so audio-visual synchronization can be achieved when playing the audio data.

[0103] In this embodiment of the application, facial expression data corresponding to the audio data is determined by determining the fourth emotional information and / or the third sound attribute information corresponding to the audio data. Then, by binding the facial expression data to the facial bones of the virtual image, it is easy to determine the facial expression corresponding to the audio data. Furthermore, the virtual image is driven by the audio data so that the corresponding facial expression data can be played simultaneously when the audio data is played. Since the facial expression data is determined by the audio data and the facial expression data is bound to the facial bones of the virtual image, it helps to improve the accuracy of audio-visual correspondence and achieve audio-visual synchronization.

[0104] Furthermore, since different virtual characters may correspond to different skeletal points, in order to improve the accuracy of binding and better match the facial expressions of the virtual characters, the points of the facial expression data are bound to the points of the facial bones of the virtual characters. This may also include: acquiring image information containing the virtual characters; and identifying the points of the facial bones of the virtual characters based on the image information containing the virtual characters.

[0105] Specifically, after acquiring the image of the virtual avatar, the image information containing the virtual avatar is imported into the feature recognition model. Facial feature recognition is performed on the image information containing the virtual avatar, and then the skeletal points in the virtual avatar's face are determined based on the recognized facial features.

[0106] Furthermore, after identifying each point in the facial skeleton of the virtual image, each point of the facial expression data is bound to each point in the facial skeleton of the virtual image. Specifically, this may include binding each point of the facial expression data to each point in the identified facial skeleton of the virtual image.

[0107] For the embodiments of this application, the method of binding each point of facial expression data with each point of the facial bones of the recognized virtual image is detailed in the above embodiments and will not be repeated here.

[0108] The above embodiments describe a speech generation and an expression-driven method from the perspectives of the client and server, respectively. The following embodiments illustrate the speech generation and expression-driven method of this application through a specific scenario of client-server interaction. Figure 3 As shown, after obtaining the text data, the client determines the audio data through the AI ​​voice service, and then determines the facial expression data of the virtual character corresponding to the audio data. The client generates a data stream (custom protocol data) by simulating a custom protocol and sends it to the server. After obtaining the audio data and the facial expression data of the virtual character, the server binds each point of the facial expression data to each point in the facial skeleton of the virtual character. Based on the binding relationship, the server controls the client to drive the facial expression of the virtual character while playing the audio data, so as to achieve audio-visual synchronization.

[0109] The above embodiments describe a method for speech generation and expression-driven processing from the perspective of process flow. The following embodiments describe a device for speech generation and expression-driven processing from the perspective of virtual modules or virtual units. For details, please refer to the following embodiments.

[0110] This application provides a speech generation and expression-driven device, such as... Figure 4 As shown, the device may specifically include an audio data acquisition module 410, an expression data determination module 420, a data transmission module 430, and an execution driver module 440, wherein: Audio data acquisition module 410 is used to acquire audio data; The facial expression data determination module 420 is used to determine the facial expression data of the virtual character corresponding to the audio data. The facial expression data of the virtual character includes multiple points. The data sending module 430 is used to send audio data and facial expression data of the virtual image corresponding to the audio data to the server, so that the server can bind each point of the facial expression data to each point of the facial skeleton of the virtual image. The execution driver module 440 is used to output the facial expressions of the virtual avatar while playing audio data, based on the driver instructions from the server.

[0111] In one possible implementation, the audio data acquisition module 410, when acquiring audio data, specifically performs the following functions: Obtain text information and segment it into sentences to obtain a set of sentences; Each statement is parsed to obtain the first emotional information and / or the first voice attribute information corresponding to each statement; Based on the first emotional information and / or the first sound attribute information corresponding to each statement, audio data corresponding to the text information is generated to obtain audio data.

[0112] In one possible implementation, the device further includes: The information acquisition module is used to acquire information about the virtual avatar and / or secondary emotional information input by the user; Specifically, the audio data acquisition module, when parsing each statement to obtain the first emotional information and / or first sound attribute information corresponding to each statement, is used for: Based on the virtual avatar information and / or the second emotional information input by the user, each statement is parsed to obtain the first emotional information and / or the first voice attribute information corresponding to each statement.

[0113] In one possible implementation, when determining the facial expression data of the virtual character corresponding to the audio data, the expression data determining module 420 is specifically used for: The audio data is segmented into sentences, and the third emotional information and / or second voice attribute information corresponding to each audio sentence are obtained. Based on the third emotional information and / or second voice attribute information corresponding to each audio statement, the facial expression data of the virtual avatar corresponding to each audio statement is determined.

[0114] In one possible implementation, when determining the facial expression data of the virtual character corresponding to the audio data, the expression data determining module 420 is specifically used for: Determine the mouth movement information corresponding to the audio data; Based on the mouth movement information corresponding to the audio data, determine the movement information of other parts of the face; Based on the mouth movement information and other facial movement information corresponding to the audio data, the various points of the facial expression data are determined to obtain the facial expression data of the virtual image corresponding to the audio data.

[0115] In one possible implementation, the device further includes: The data stream generation module is used to generate a data stream from audio data and the facial expression data of the corresponding virtual avatar through a specific network protocol. Specifically, when sending audio data and the facial expression data of the virtual avatar corresponding to the audio data to the server, the data sending module 430 is used for: Send a data stream to the server.

[0116] This application also provides a speech generation and expression-driven device, such as... Figure 5 As shown, the device may specifically include a data acquisition module 510, a point binding module 520, and a control drive module 530, wherein: The data information module 510 is used to acquire audio data and the facial expression data of the virtual image corresponding to the audio data. The facial expression data of the virtual image includes multiple points. The point binding module 520 is used to bind each point of facial expression data to each point in the facial skeleton of the virtual image; The control driver module 530 is used to control the client to drive the facial expressions of the virtual avatar while playing audio data, based on the binding relationship.

[0117] In one possible implementation, the device further includes: The information determination module is used to determine the fourth emotional information and / or third voice attribute information corresponding to each statement in the audio data; The device also includes: The displacement information determination module is used to determine the displacement information of each point in the facial skeleton corresponding to the virtual image based on the fourth emotional information and / or third voice attribute information corresponding to each statement and the binding relationship. Specifically, when the control driver module 530 controls the client to drive the facial expressions of the virtual avatar while playing audio data based on the binding relationship, it is used for: Based on the binding relationship and the displacement information of various points in the facial skeleton, the client controls the facial expressions of the virtual character while playing audio data.

[0118] In one possible implementation, the device further includes: The image information acquisition module is used to acquire image information containing virtual avatars; The identification point module is used to identify various points of the facial bones of the virtual image based on the image information containing the virtual image. Specifically, the point binding module 520, when binding each point of facial expression data to each point in the facial skeleton of the virtual image, is used for: The data points of facial expressions are bound to the data points of the facial skeleton of the identified virtual avatar.

[0119] This application provides a client, such as... Figure 6 As shown, Figure 6 The client 600 shown includes a processor 601 and a memory 603. The processor 601 and the memory 603 are connected, for example, via a bus 602. Optionally, the client 600 may also include a transceiver 604. It should be noted that in practical applications, the transceiver 604 is not limited to one type, and the structure of this client 600 does not constitute a limitation on the embodiments of this application.

[0120] Processor 601 may be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute the various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of this application. Processor 601 may also be a combination that implements computational functions, such as including one or more microprocessor combinations, a combination of a DSP and a microprocessor, etc.

[0121] Bus 602 may include a pathway for transmitting information between the aforementioned components. Bus 602 may be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, etc. Bus 602 can be divided into address bus, data bus, control bus, etc. For ease of representation, Figure 6 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.

[0122] The memory 603 may be a ROM (Read Only Memory) or other type of static storage device capable of storing static information and instructions, RAM (Random Access Memory) or other type of dynamic storage device capable of storing information and instructions, or an EEPROM (Electrically Erasable Programmable Read Only Memory), CD-ROM (Compact Disc Read Only Memory) or other optical disc storage, optical disc storage (including compressed optical discs, laser discs, optical discs, digital universal optical discs, Blu-ray discs, etc.), magnetic disk storage media or other magnetic storage devices, or any other medium capable of carrying or storing desired program code in the form of instructions or data structures and accessible by a computer, but not limited thereto.

[0123] The memory 603 stores application code that executes the scheme of this application, and its execution is controlled by the processor 601. The processor 601 executes the application code stored in the memory 603 to implement the content shown in the foregoing method embodiments.

[0124] Client devices include, but are not limited to: mobile terminals such as mobile phones, laptops, digital radio receivers, PDAs (Personal Digital Assistants), PADs (Tablet PCs), PMPs (Portable Multimedia Players), and in-vehicle terminals (such as in-vehicle navigation terminals), as well as fixed terminals such as digital TVs and desktop computers. Servers can also be included. Figure 6 The client shown is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of this application.

[0125] This application also provides a server in this embodiment. Please refer to the above description of a client for details. The specific content of the server will not be repeated in this application embodiment.

[0126] This application provides a computer-readable storage medium storing a computer program that, when run on a computer, enables the computer to execute the corresponding content in the aforementioned method embodiments.

[0127] It should be understood that although the steps in the flowcharts of the accompanying figures are shown sequentially as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the accompanying figures may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the sub-steps or stages of other steps.

[0128] The above description is only a partial embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of this application, and these improvements and modifications should also be considered within the scope of protection of this application.

Claims

1. A speech generation and expression-driven method, characterized in that, Executed by the client, including: Acquire audio data; Determine the facial expression data of the virtual avatar corresponding to the audio data, wherein the facial expression data of the virtual avatar includes multiple points; The audio data and the facial expression data of the virtual image corresponding to the audio data are sent to the server so that the server can bind each point of the facial expression data to each point of the facial skeleton of the virtual image. Based on the server's driver instructions, the virtual avatar's facial expressions are output while the audio data is being played. The acquisition of audio data includes: Obtain text information and perform sentence segmentation on the text information to obtain a set of sentences; Each statement is parsed to obtain the first emotional information and / or the first voice attribute information corresponding to each statement; Based on the first emotional information and / or first sound attribute information corresponding to each statement, audio data corresponding to the text information is generated to obtain the audio data; The step of parsing each statement to obtain the first emotional information and / or first voice attribute information corresponding to each statement also includes, prior to: To obtain information about virtual avatars and / or secondary emotional information input by users; The step of parsing each statement to obtain the first emotional information and / or first voice attribute information corresponding to each statement includes: Based on the virtual avatar information and / or the second emotional information input by the user, each statement is parsed to obtain the first emotional information and / or the first voice attribute information corresponding to each statement.

2. The speech generation and expression-driven method according to claim 1, characterized in that, The step of determining the facial expression data of the virtual avatar corresponding to the audio data includes: The audio data is segmented into sentences, and the third emotional information and / or second sound attribute information corresponding to each audio sentence are obtained; Based on the third emotional information and / or second voice attribute information corresponding to each audio statement, the facial expression data of the virtual image corresponding to each audio statement is determined.

3. A speech generation and expression-driven method according to any one of claims 1-2, characterized in that, The step of determining the facial expression data of the virtual avatar corresponding to the audio data includes: Determine the mouth movement information corresponding to the audio data; Based on the mouth movement information corresponding to the audio data, determine the movement information of other parts of the face; Based on the mouth movement information corresponding to the audio data and the movement information of other parts of the face, the various points of the facial expression data are determined to obtain the facial expression data of the virtual image corresponding to the audio data.

4. The speech generation and expression-driven method according to claim 1, characterized in that, The step of sending the audio data and the facial expression data of the virtual avatar corresponding to the audio data to the server also includes, prior to: The audio data and the facial expression data of the corresponding virtual avatar are used to generate a data stream via a network protocol; The step of sending the audio data and the facial expression data of the virtual avatar corresponding to the audio data to the server includes: The data stream is sent to the server.

5. A speech generation and expression-driven method, characterized in that, Executed by the server, including: The client receives audio data sent by the client and facial expression data of the virtual avatar corresponding to the audio data. The client executes a voice generation and expression driving method as described in claim 1. The server obtains the audio data and facial expression data of the virtual avatar corresponding to the audio data. The facial expression data of the virtual avatar includes multiple points. The facial expression data points are bound to the facial bones of the virtual image. Based on the binding relationship, the client controls the facial expressions of the virtual avatar while playing audio data.

6. The speech generation and expression-driven method according to claim 5, characterized in that, The process of binding each point of facial expression data to each point in the facial skeleton of the virtual avatar also includes: Determine the fourth emotional information and / or third voice attribute information corresponding to each statement in the audio data; The process of binding each point of facial expression data to each point in the facial skeleton of the virtual avatar further includes: Based on the fourth emotional information and / or third voice attribute information corresponding to each statement and the binding relationship, the displacement information of each point in the facial skeleton corresponding to the virtual image is determined. The step of controlling the client to drive the facial expressions of the virtual avatar while playing audio data, based on the binding relationship, includes: Based on the binding relationship and the displacement information of each point in the facial skeleton, the client is controlled to drive the facial expressions of the virtual character while playing audio data.

7. The speech generation and expression-driven method according to claim 5, characterized in that, The process of binding each point of facial expression data to each point in the facial skeleton of the virtual avatar also includes: Obtain image information containing the virtual image; Based on the image information containing the virtual image, identify the various points of the facial bones of the virtual image; The step of binding each point of facial expression data with each point in the facial skeleton of the virtual image includes: The facial expression data points are bound to the facial bones of the identified virtual image.

8. A client application, characterized in that, The client includes: At least one processor; Memory; At least one application, wherein the at least one application is stored in memory and configured to be executed by at least one processor, the at least one application being configured to: perform the speech generation and expression-driven method of any one of claims 1-4.

9. A server-side component, characterized in that, The server includes: At least one processor; Memory; At least one application, wherein the at least one application is stored in memory and configured to be executed by at least one processor, the at least one application being configured to: perform the speech generation and expression-driven method of any one of claims 5-7.

Citation Information

Patent Citations

  • Expression generation method and device, computing equipment and storage medium

    CN110570499A

  • Text broadcasting method and device, electronic equipment and storage medium

    CN110941954A

  • Electronic book audio generation method, electronic equipment and storage medium

    CN111739509A

  • Facial information generation method and device

    CN114513678A

  • Model processing method and device and emotional speech synthesis method and device

    CN114724540A