Speech synthesis method and device, electronic device and storage medium

By training the speech synthesis model for speech disorder users, using the target speech characteristics and text data, selecting speech data with high tone similarity, the problem of large differences between speech synthesis and users is solved, and accessibility communication between speech disorder users is achieved.

CN115148185BActive Publication Date: 2025-05-06BEIJING XIAOMI MOBILE SOFTWARE CO LTD +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210483322.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-05
Publication Date
2025-05-06
Estimated Expiration
2042-05-05

AI Technical Summary

Technical Problem

In the prior art, voice synthesized voice has the characteristics of mechanical vocalization, and the tone is very different from the user's own tone, and it is difficult to generate voice data that matches the tone for the speech hindered user.

Method used

By obtaining text data of speech disorder users, and using machine learning methods to train the speech synthesis model, the model uses target speech characteristics and corresponding text data for training, and the target speech characteristics are pre-selected speech characteristics of the target user. Among the multiple candidate speech data, voice data with high tone similarity is selected as the target speech data.

Benefits of technology

It enables speech-impaired users to express their content in their own customized tone, improving the barrier-free communication experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115148185B_ABST
    Figure CN115148185B_ABST
Patent Text Reader

Abstract

The embodiments of the present disclosure provide a speech synthesis method, device, electronic device and computer-readable storage medium. The message processing method comprises: obtaining text data input by a speech-impaired user; inputting the text data into a speech synthesis model to obtain synthesized speech data; wherein the synthesized speech data and the text data have the same language content, and the synthesized speech data is audio data; wherein the speech synthesis model is obtained by training a basic model using a sample set using a machine learning method, and the sample set comprises: the target speech feature and text data corresponding to the language content of the target speech feature, wherein the target speech feature is a speech feature of a pre-selected target user.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technology, and in particular to a speech synthesis method, device, electronic device and computer-readable storage medium. Background Art

[0002] Machine learning models can be used for speech synthesis. For example, using audio data as sample data to train a predetermined model will result in subsequent synthesized speech based on text or the like.

[0003] On the one hand, in the related art, there may be problems with synthesized speech, including: the synthesized speech has the characteristics of mechanical sounding, or the timbre of the synthesized speech is very different from the timbre of a specific user.

[0004] On the other hand, in the related technology, some speech-impaired users have pronunciation or voice expression problems. At this time, how to obtain voice data to generate a machine learning model that can replace the speech-impaired users to generate voice is another problem that needs to be solved urgently in the existing technology. Summary of the invention

[0005] Embodiments of the present disclosure provide a speech synthesis method, apparatus, electronic device, and computer-readable storage medium.

[0006] In a first aspect, the present application provides a speech synthesis method, which may include:

[0007] Get text data input by speech-impaired users;

[0008] Inputting the text data into a speech synthesis model to obtain synthesized speech data; wherein the synthesized speech data and the text data have the same language content, and the synthesized speech data is audio data;

[0009] The speech synthesis model is obtained by training a basic model using a sample set through a machine learning method, and the sample set includes: the target speech feature and text data corresponding to the language content of the target speech feature, wherein the target speech feature is a pre-selected speech feature of a target user.

[0010] In some possible implementations, the sample set includes a plurality of samples, each of which includes a speech-text pair;

[0011] Each of the speech-text pairs includes a phoneme and a Mel-spectrogram feature parameter corresponding to the phoneme, wherein the phoneme is obtained by preprocessing the text.

[0012] In some possible implementations, the target speech feature is obtained by extracting features from target speech data.

[0013] The method further includes: selecting from a plurality of candidate voice data to obtain the target voice data; selecting from a plurality of candidate voice data to obtain the target voice data.

[0014] In some possible implementations, selecting from a plurality of candidate voice data to obtain the target voice data includes:

[0015] Inputting the speech data of the speech-impaired user and the plurality of candidate speech data into a voiceprint recognition model to obtain timbre similarities between the plurality of candidate speech data and the speech data of the speech-impaired user;

[0016] According to the timbre similarity, one target voice data is determined from the plurality of candidate voice data.

[0017] In some possible implementations, determining the target voice data from the plurality of candidate voice data according to the timbre similarity includes:

[0018] Acquiring voice data of the speech-impaired user;

[0019] Analyze the speech data of the speech-impaired user to obtain a first speech feature;

[0020] Analyze the plurality of candidate voice data to obtain a second voice feature;

[0021] The step of inputting the speech data of the speech-impaired user and the plurality of candidate speech data into a voiceprint recognition model to obtain timbre similarities between the plurality of candidate speech data and the speech data of the speech-impaired user includes:

[0022] The second voice feature and the first voice feature are selected from the multiple candidate voice data and input into the voiceprint recognition model, and the voice data of the speech-impaired user and the multiple candidate voice data are input into the voiceprint recognition model to obtain the timbre similarity between the multiple candidate voice data and the voice data of the speech-impaired user.

[0023] In some possible implementations, selecting from a plurality of candidate voice data according to the timbre similarity to obtain the target voice data includes:

[0024] According to the timbre similarity, selecting candidate voice data whose similarity is in a first similarity interval from the plurality of candidate voice data as the first voice data;

[0025] According to the similarity between the user attribute corresponding to the first voice data and the user attribute of the speech-impaired user, selecting the first voice data whose similarity is in a second similarity interval from the first voice data;

[0026] The target voice data is obtained according to the first voice data.

[0027] In some possible implementations, selecting target voice data from the first voice data according to the similarity between the user attributes of the pronunciation user of the first voice data and the user attributes of the speech-impaired user includes:

[0028] Based on the speech synthesis request, determining the second speech data; wherein the speech synthesis request is used to indicate the selection result of the user, and the second speech data is the speech data in the first speech data that matches the speech synthesis request;

[0029] The target voice data is obtained according to the second voice data.

[0030] In some possible implementations, obtaining the target voice data according to the second voice data includes:

[0031] detecting an adjustment instruction acting on the second voice data;

[0032] According to the adjustment instruction, the second voice data is adjusted to obtain the target voice data, wherein the adjustment instruction is used to instruct to adjust the proportion of sound waves in different frequency bands in the second voice data.

[0033] In some possible implementations, the speech synthesis model includes:

[0034] A text encoding module, used to extract linguistic information from a first phoneme sequence and obtain a text encoding sequence representing the linguistic information; wherein the first phoneme sequence is obtained by preprocessing text data;

[0035] A phoneme filtering module, used for filtering the text encoding sequence to obtain a second phoneme sequence consisting of initial consonants and / or finals;

[0036] a duration prediction module, configured to obtain a first duration sequence according to the second phoneme sequence, wherein the first duration sequence includes: the duration of each phoneme in the second phoneme sequence; wherein the second phoneme sequence and the first duration sequence are added to obtain a first sequence;

[0037] The attention mechanism module is used to receive a first sequence, align the first sequence with each element of an acoustic feature sequence obtained based on the first sequence, and obtain a second sequence; the second sequence includes: a frame length of an acoustic feature of each phoneme; wherein the elements of the acoustic feature sequence include: a Mel-spectrogram feature parameter corresponding to the phoneme;

[0038] an acoustic decoding module, configured to obtain the acoustic features of the synthesized speech data according to the first sequence and the second sequence; in some possible implementations, when the target speech data is used to train the basic model, a phoneme classification module is used to perform phoneme classification according to the second phoneme sequence output by the phoneme filtering model of the basic model to obtain a phoneme classification result;

[0039] Determine a loss value according to a difference between the phoneme classification result and the label corresponding to the text data;

[0040] According to the loss value, the model parameters of the basic model are adjusted to obtain the speech synthesis model.

[0041] In a second aspect, the present application provides a speech synthesis device, comprising a plurality of functional units for implementing any one of the methods of the first aspect. The speech synthesis device may include:

[0042] An acquisition module, used to acquire text data input by a speech-impaired user;

[0043] A synthesis module, used for inputting the text data into a speech synthesis model to obtain synthesized speech data;

[0044] The synthesized speech data and the text data have the same language content, and the synthesized speech data is audio data; the speech synthesis model is obtained by training a basic model using a sample set through a machine learning method, and the sample set includes: the target speech feature and text data corresponding to the language content of the target speech feature, wherein the target speech feature is a pre-selected speech feature of a target user.

[0045] In some possible implementations, the sample set includes a plurality of samples, each of which includes a speech-text pair;

[0046] Each of the speech-text pairs includes a phoneme and a Mel-spectrogram feature parameter corresponding to the phoneme, wherein the phoneme is obtained by preprocessing the text.

[0047] In some possible implementations, the target speech feature is obtained by extracting features from the target speech data.

[0048] The device also includes:

[0049] The selection module is used to select from a plurality of candidate voice data to obtain the target voice data.

[0050] In some possible implementations, the selection module is specifically used to:

[0051] Inputting the speech data of the speech-impaired user and the plurality of candidate speech data into a voiceprint recognition model to obtain timbre similarities between the plurality of candidate speech data and the speech data of the speech-impaired user;

[0052] According to the timbre similarity, one target voice data is determined from the plurality of candidate voice data.

[0053] In some possible implementations, the selection module is specifically used to:

[0054] Acquiring voice data of the speech-impaired user;

[0055] Analyze the speech data of the speech-impaired user to obtain a first speech feature;

[0056] Analyze the plurality of candidate voice data to obtain a second voice feature;

[0057] The second voice feature and the first voice feature are selected from the multiple candidate voice data and input into the voiceprint recognition model, and the voice data of the speech-impaired user and the multiple candidate voice data are input into the voiceprint recognition model to obtain the timbre similarity between the multiple candidate voice data and the voice data of the speech-impaired user.

[0058] In some possible implementations, the selection module is specifically used to:

[0059] According to the timbre similarity, selecting candidate voice data whose similarity is in a first similarity interval from the plurality of candidate voice data as the first voice data;

[0060] According to the similarity between the user attribute corresponding to the first voice data and the user attribute of the speech-impaired user, selecting the first voice data whose similarity is in a second similarity interval from the first voice data;

[0061] The target voice data is obtained according to the first voice data.

[0062] In some possible implementations, after obtaining the first voice data, a module is selected, specifically configured to:

[0063] Based on the speech synthesis request, determining the second speech data; wherein the speech synthesis request is used to indicate the selection result of the user, and the second speech data is the speech data in the first speech data that matches the speech synthesis request;

[0064] The target voice data is obtained according to the second voice data.

[0065] In some possible implementations, the selection module is specifically used to:

[0066] selecting second voice data from the first voice data according to a selection input reflecting the speech synthesis needs of the speech-impaired user; wherein the second voice data is the first voice data selected by the selection input;

[0067] The target voice data is obtained according to the second voice data.

[0068] In some possible implementations, the speech synthesis model includes:

[0069] A text encoding module, used to extract linguistic information from a first phoneme sequence and obtain a text encoding sequence representing the linguistic information; wherein the first phoneme sequence is obtained by preprocessing text data;

[0070] A phoneme filtering module, used for filtering the text encoding sequence to obtain a second phoneme sequence consisting of initial consonants and / or finals;

[0071] a duration prediction module, configured to obtain a first duration sequence according to the second phoneme sequence, wherein the first duration sequence includes: the duration of each phoneme in the second phoneme sequence; wherein the second phoneme sequence and the first duration sequence are added to obtain a first sequence;

[0072] The attention mechanism module is used to receive a first sequence, align the first sequence with each element of an acoustic feature sequence obtained based on the first sequence, and obtain a second sequence; the second sequence includes: a frame length of an acoustic feature of each phoneme; wherein the elements of the acoustic feature sequence include: a Mel-spectrogram feature parameter corresponding to the phoneme;

[0073] The acoustic decoding module is used to obtain the acoustic features of the synthesized speech data according to the first sequence and the second sequence.

[0074] In some possible implementations, the device further includes:

[0075] A classification module, configured to perform phoneme classification according to a second phoneme sequence output by a phoneme filtering model of the basic model using a phoneme classification module when the basic model is trained using the target speech data, so as to obtain a phoneme classification result;

[0076] A determination module, used to determine a loss value according to a difference between the phoneme classification result and a label corresponding to the text data;

[0077] An adjustment module is used to adjust the model parameters of the basic model according to the loss value to obtain the speech synthesis model.

[0078] In a third aspect, the present application further provides an electronic device, including:

[0079] a memory for storing processor-executable instructions;

[0080] Processor; wherein the processor is configured to: implement the method of the first aspect and its possible implementation methods when executing executable instructions.

[0081] In a fourth aspect, the present application provides a computer-readable storage medium storing an executable program, wherein the executable program, when executed by a processor, implements the method of the first aspect and its possible implementation methods.

[0082] Compared with the prior art, the technical solution provided by the embodiment of the present application has the following beneficial effects:

[0083] In the present application, text data input by a speech-impaired user is obtained, and the text data is input into a speech synthesis model to obtain synthesized speech data, wherein the synthesized speech data and the text data have the same language content, and the synthesized speech data is audio data, wherein the speech synthesis model is obtained by training a basic model using a sample set using a machine learning method, and the sample set includes: the target speech feature and text data corresponding to the language content of the target speech feature, wherein the target speech feature is a pre-selected speech feature of a target user. In this way, speech synthesis of text data input by a person with pronunciation impairment can be completed based on the speech synthesis model, so that the speech-impaired user can express the content to be expressed with his or her own customized timbre, thereby improving the barrier-free communication experience of special users.

[0084] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0085] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention.

[0086] Figure 1 It is a flowchart of an embodiment of a speech synthesis method in an embodiment of the present application.

[0087] Figure 2 It is a flowchart of another speech synthesis method in the embodiments of the present application.

[0088] Figure 3 It is a structural diagram of another speech synthesis model in an embodiment of the present application.

[0089] Figure 4 It is a flowchart of an embodiment of a speech synthesis method of an embodiment of the present application.

[0090] Figure 5 A schematic diagram of the training and use stages of a speech synthesis model in an embodiment of the present application.

[0091] Figure 6 It is a structural schematic diagram of a speech synthesis device in an embodiment of the present application.

[0092] Figure 7 It is a schematic diagram of the structure of an electronic device in an embodiment of the present application. DETAILED DESCRIPTION

[0093] In the following description, reference is made to the drawings that form part of the present application and show specific aspects of the embodiments of the present application or specific aspects of the embodiments of the present application in an illustrative manner. It should be understood that the embodiments of the present application can be used in other aspects and may include structural or logical changes not depicted in the drawings. Therefore, the following detailed description should not be understood in a restrictive sense, and the scope of the present application is defined by the appended claims. For example, it should be understood that the disclosure in conjunction with the described method can be equally applicable to the corresponding device or apparatus for performing the method, and vice versa. For example, if one or more specific method steps are described, the corresponding device may include one or more units such as functional units to perform the one or more method steps described (for example, one unit performs one or more steps, or multiple units, each of which performs one or more of the multiple steps), even if such one or more units are not explicitly described or illustrated in the drawings. On the other hand, for example, if a specific device is described based on one or more units such as functional units, the corresponding method may include a step to perform the functionality of one or more units (for example, one step performs the functionality of one or more units, or multiple steps, each of which performs the functionality of one or more units in multiple units), even if such one or more steps are not explicitly described or illustrated in the drawings. Further, it should be understood that, unless explicitly stated otherwise, features of the various exemplary embodiments and / or aspects described herein may be combined with each other.

[0094] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0095] Text To Speech (TTS), or "from text to speech", is part of human-computer dialogue, allowing machines to speak. Speech synthesis technology and speech recognition technology are two key technologies necessary to achieve human-computer voice communication and build a spoken language system with listening and speaking capabilities. It enables computer devices to have human-like speaking capabilities.

[0096] In the process of speech synthesis, it is necessary to first obtain the audio data of the user's recording and the corresponding text audio pair as training data, and use the machine learning method to input the training data and pre-train the speech synthesis model with a large amount of data. However, for speech-impaired users, speech-impaired users have language barriers, such as damaged vocal cords that make it impossible to make sounds, or they cannot speak. In this way, it is impossible to collect the speech data of speech-impaired users for model training, which makes the synthesized speech generated by the synthetic speech model used by speech-impaired users fail to reflect the voice characteristics of the speech-impaired users themselves.

[0097] In view of this, an embodiment of the present application provides a speech synthesis method to solve the above-mentioned problem.

[0098] See also Figure 1 As shown, a speech synthesis method provided in an embodiment of the present application may include:

[0099] S101, obtaining text data input by a speech-impaired user;

[0100] S102, inputting the text data into a speech synthesis model to obtain synthesized speech data; wherein the synthesized speech data and the text data have the same language content, and the synthesized speech data is audio data;

[0101] Wherein, the speech synthesis model is obtained by training a basic model using a sample set through a machine learning method, and the sample set includes: the target speech feature and text data corresponding to the language content of the target speech feature, wherein the target speech feature is a speech feature of a pre-selected target user.

[0102] The speech synthesis method can be performed by an electronic device such as a terminal device or a server. The terminal device can be a user equipment (UE), a personal digital assistant (PDA), a handheld device, a wearable device, etc. The server can be implemented by an independent server or a server cluster composed of multiple physical servers.

[0103] Receive text data input by a user from a user interface (UI) of an electronic device, for example, text data input through a physical keyboard or a displayed virtual keyboard. The text data may be data input in a language that can be recognized by various speech synthesis models. For example, the text data may be text data input in Chinese characters or text data in English.

[0104] The speech synthesis model includes but is not limited to: a machine learning model. In short, the speech synthesis model can be a model that can convert input data into synthesized speech data.

[0105] The speech synthesis module provided in the embodiments of the present disclosure is a type of speech synthesis (Text To Speech, TTS) technology, that is, "from text to speech", which is a part of human-computer dialogue.

[0106] Exemplarily, when the electronic device obtains input data from a speech-impaired user and activates the speech synthesis function, the text data is converted into an input sequence suitable for a speech synthesis model and input into the speech synthesis model; the speech synthesis model performs a series of calculations on the input sequence based on its own model parameters and outputs the synthesized speech data.

[0107] In the embodiment of the present disclosure, the target user is different from the speech-impaired user, and the speech synthesis model is generated by training using speech data of the target user.

[0108] The speech-impaired user has a speech-impaired function, and may not know how to speak because of hearing loss, or cannot speak due to vocal cord abnormalities, etc. In the disclosed embodiment, a speech synthesis model can be used to generate synthesized speech data simulating the speech of the speech-impaired user after acquiring the text data of the speech-impaired user.

[0109] The target users are those who can speak normally.

[0110] The synthesized speech data is speech data with the pronunciation characteristics of speech-impaired users. After being output through a speaker or earphone, it can produce an effect of simulating the speech of speech-impaired users. In this way, speech synthesis of text data input by people with pronunciation impairments can be completed based on the speech synthesis model, so that speech-impaired users can express what they want to express with their own customized timbre, thereby improving the barrier-free communication experience of special users.

[0111] Before using the speech synthesis model, it is necessary to train the speech synthesis model.

[0112] The basic model can be a general voice model trained using a large amount of user voice data. By continuing the training with the target voice data, the trained voice synthesis model can output a voice synthesis model with the desired timbre of the speech-impaired user.

[0113] Exemplarily, the sample set includes a plurality of samples, each of which includes a speech-text pair;

[0114] Each of the speech-text pairs includes a phoneme and a Mel-spectrogram feature parameter corresponding to the phoneme, wherein the phoneme is obtained by preprocessing the text.

[0115] The Mel spectrum feature parameter is a parameter that reflects the characteristics of the sound.

[0116] The pretreatment includes but is not limited to at least one of the following:

[0117] Processing of spoken text such as word segmentation, part of speech prediction, prosodic word prediction, prosodic phrase prediction, intonation phrase prediction and / or text-to-phoneme conversion.

[0118] In one embodiment, the target speech feature is obtained by extracting features from the target speech data.

[0119] The method further comprises:

[0120] The target voice data is obtained by selecting from a plurality of candidate voice data.

[0121] For example, a plurality of speech databases are preset, and the speech databases contain a lot of candidate speech data.

[0122] In the disclosed embodiment, the target voice data is selected from a candidate voice database.

[0123] Exemplarily, a voice database of various voice data donors is pre-established, and each voice in the voice database can be used as the candidate voice data.

[0124] The step of selecting from a plurality of candidate voice data to obtain the target voice data comprises:

[0125] Inputting the speech features of the speech-impaired user and the speech features of the plurality of candidate speech data into a voiceprint recognition model to obtain timbre similarities between the plurality of candidate speech data and the speech data of the speech-impaired user;

[0126] According to the timbre similarity, one target voice data is determined from the plurality of candidate voice data.

[0127] refer to Figure 2As shown, before using the voiceprint recognition model to obtain the timbre similarity, the method further includes:

[0128] S201: Acquire voice data of the speech-impaired user;

[0129] S202: Analyze the speech data of the speech-impaired user to obtain a first speech feature;

[0130] S203: Analyze the plurality of candidate voice data to obtain a second voice feature;

[0131] S204: Selecting the second voice feature and the first voice feature from the plurality of candidate voice data and inputting them into the voiceprint recognition model to obtain timbre similarities between the plurality of candidate voice data and the voice data of the speech-impaired user;

[0132] S205: Determine the target voice data from the plurality of candidate voice data according to the timbre similarity.

[0133] For example, the speech-impaired user may make simple speech sounds, such as "ah ah", "hmm", "ah eh", etc. After these speech data are collected, they can be used to extract speech features of the speech-impaired user.

[0134] In the embodiment of the present disclosure, the first speech feature and the second speech feature may be features extracted using the same speech feature extraction method, and the features include various acoustic features.

[0135] For example, the first speech feature and the second speech feature can be obtained by performing pre-emphasis, framing, windowing, Fourier transform, energy spectrum calculation, Mel filtering, and logarithm processing on the speech data.

[0136] The extracted first speech feature and the second speech feature are jointly input into the voiceprint recognition model to obtain features such as timbre similarity output by the voiceprint recognition model that reflect the similarity of voice characteristics between the speech-impaired user and the target user, thereby facilitating the selection of target speech data for training the basic model from a plurality of candidate speech data.

[0137] In some embodiments, the first voice feature and the second voice feature may further include:

[0138] The target user is selected based on parameters such as volume loudness, fundamental frequency (pitch), vocal tract length and resonance peak, and the speech data of the target user is selected as the target speech data.

[0139] In some embodiments, selecting the target voice data from the plurality of candidate voice data comprises:

[0140] According to the timbre similarity, selecting candidate voice data whose similarity is in a first similarity interval from the plurality of candidate voice data as the first voice data;

[0141] According to the similarity between the user attribute corresponding to the first voice data and the user attribute of the speech-impaired user, selecting the first voice data whose similarity is in a second similarity interval from the first voice data;

[0142] The target voice data is obtained according to the second voice data.

[0143] In one embodiment, the candidate voice data with the highest timbre similarity may be directly selected as the target voice data.

[0144] In the disclosed embodiment, target voice data is further selected based on user attribute similarity.

[0145] In some embodiments, the first similarity interval may include: an interval greater than a first similarity threshold. For example, the first similarity threshold may be 50%, 60%, 70% or 80%.

[0146] In some other embodiments, the second similarity interval may include: an interval of a second similarity threshold, for example, the second similarity threshold may be 55%, 75%, 80% or 85%.

[0147] In this embodiment, the height and weight of the speaker will determine the voice characteristics. Therefore, candidate voice data of multiple candidate users with the highest voice similarity in descending order can be selected as the first voice data according to the voice similarity.

[0148] Furthermore, according to the user attributes, further selection is made from the first voice data to obtain the target voice data.

[0149] Exemplarily, the user attributes include: User attributes may include but are not limited to: basic attributes such as gender, age, education level and / or region.

[0150] Furthermore, the user attributes also include: non-basic attributes such as user height, weight and / or body shape that may affect the user's pronunciation.

[0151] One or more user attributes may be mapped to a feature space, and the distance between the feature vectors in the feature space mapped with the user attributes of the candidate user and the speech-impaired user is calculated to obtain the user attribute similarity.

[0152] In one embodiment, the target voice data is obtained by selecting the first voice data with the highest user attribute similarity;

[0153] In another embodiment, a weighted sum calculation is performed based on the user attribute similarity and the timbre similarity to obtain a candidate user with the highest weighted sum value as the target user, and then the target voice data is determined based on the voice data of the target user.

[0154] In the disclosed embodiment, the target user is determined not only by combining the timbre but also the user attributes, so that target speech data that better meets the needs of speech-impaired users can be selected, thereby making the speech synthesis effect of the speech synthesis model better.

[0155] In one embodiment, after obtaining the first voice data, the method further includes:

[0156] Based on the speech synthesis request, determining the second speech data; wherein the speech synthesis request is used to indicate the selection result of the user, and the second speech data is the speech data in the first speech data that matches the speech synthesis request;

[0157] The target voice data is obtained according to the second voice data.

[0158] For example, multiple candidate users of the target user are selected by combining the timbre similarity and the user attributes, and a speech synthesis request can be further obtained at this time; the speech synthesis request can be determined according to the selection input of the speech-impaired user or the selection input of the relatives and friends of the speech-impaired user. According to the speech synthesis request, the second speech data of a target user is further selected from the first speech data.

[0159] For example, an audio output device such as a speaker is used to output the first voice data, and the speech-impaired user or the relatives and friends of the speech-impaired user can select a first voice data as the second voice data according to their own preferences. Therefore, the selection input reflects the speech synthesis request of the speech-impaired user. The speech synthesis request reflects the user's speech synthesis-related needs.

[0160] In one embodiment, the second speech data may be directly used as target speech data for training the basic model.

[0161] In this embodiment, in order to consider the privacy protection of the target user, etc., the sound characteristics of the voice data are further adjusted.

[0162] At this time, according to the adjustment instruction, the second voice data is adjusted to obtain adjusted second voice data, and the adjusted second voice data is the target voice data.

[0163] Therefore, in some embodiments, selecting target voice data from the first voice data according to the similarity between the user attributes of the pronunciation user of the first voice data and the user attributes of the speech-impaired user includes:

[0164] detecting an adjustment instruction acting on the second voice data;

[0165] According to the adjustment instruction, the second voice data is adjusted to obtain the target voice data, wherein the adjustment instruction is used to instruct to adjust the proportion of sound waves in different frequency bands in the second voice data.

[0166] For example, an adjustment instruction acting on the UI interface is detected, and the adjustment instruction can be an adjustment instruction for arbitrarily adjusting the timbre of the second voice data. The timbre of the user depends on the proportion of the frequency of the sound, so the adjustment instruction can be used to adjust the proportion of different sound wave frequency bands in the second voice.

[0167] The electronic device can determine the adjustment parameter according to the adjustment instruction. The adjustment parameter can be used to at least adjust the timbre of the second voice data.

[0168] The adjustment parameters may include: one or more adjustments to the frequency proportion of different sound wave frequency bands of the second voice data in the sound, thereby generating a new voice data, and the new voice data is the aforementioned target voice data.

[0169] For example, even for different users with the closest timbre, there will always be subtle differences in tone, emotion, intensity, etc. when speaking, and these differences are reflected in the proportion of different frequency bands in the sound.

[0170] For example, in the 6kHz-16kHz frequency band of sound waves, changes in frequency components affect the timbre expression and resolution. In this frequency band, a weak timbre will appear to be colorless and without personality; a strong timbre will be shrill and harsh.

[0171] Within the 600HZ-6kHz frequency band of sound waves, changes in frequency components affect the brightness and clarity of the timbre; when it is too weak, the timbre will be dim and hazy; when it is too strong, it will appear dull.

[0172] The change in the frequency component of the sound wave between 200Hz and 600Hz affects the strength and solidity of the timbre; when it is too weak, the timbre will appear empty and powerless; when it is too strong, it will be stiff and lifeless.

[0173] At 20Hz-200HZ, changes in frequency components affect the thickness and fullness of the timbre; when it is too weak, the timbre will appear thin and pale; when it is too strong, it will be turbid and unclear.

[0174] Exemplarily, the sound equalizer can be adjusted using existing audio processing tools, such as a sound converter (Sound eXchange, SoX), etc. When adjusting the sound equalizer, the audio processing tool can display adjustment progress bars for each frequency band respectively, and display drag buttons on the progress bars. Then, at this time, the adjustment instruction is based on the drag operation of the button on the adjustment progress bar. In addition, the sound equalizer can also be adjusted by displaying a numerical box on the audio processing tool that can modify specific values, and adjusting the sound equalizer by modifying the specific values. At this time, the adjustment instruction is based on the operation of modifying the specific values ​​of the equalizer.

[0175] refer to Figure 3 As shown, the speech synthesis model trained in the embodiment of the present disclosure includes:

[0176] A text encoding module, used to extract linguistic information from a first phoneme sequence and obtain a text encoding sequence representing the linguistic information; wherein the first phoneme sequence is obtained by preprocessing text data;

[0177] A phoneme filtering module, used for filtering the text encoding sequence to obtain a second phoneme sequence composed of initial consonants and / or finals;

[0178] a duration prediction module, configured to obtain a first duration sequence according to the second phoneme sequence, wherein the first duration sequence includes: the duration of each phoneme in the second phoneme sequence; wherein the second phoneme sequence and the first duration sequence are added to obtain a first sequence;

[0179] The attention mechanism module is used to receive a first sequence, align the first sequence with each element of an acoustic feature sequence obtained based on the first sequence, and obtain a second sequence; the second sequence includes: a frame length of an acoustic feature of each phoneme; wherein the elements of the acoustic feature sequence include: a Mel-spectrogram feature parameter corresponding to the phoneme;

[0180] The acoustic decoding module is used to obtain the acoustic features of the synthesized speech data according to the first sequence and the second sequence.

[0181] The speech synthesis model provided by the embodiment of the present disclosure includes the above modules. After the above modules complete the above operations, the text data can be converted into synthesized speech data.

[0182] Before the text data is input into the text encoding module, it will be pre-processed by the text analysis module to be converted into a first phoneme sequence. The pre-processing at least includes the conversion of text to phonemes. Furthermore, the pre-processing also includes but is not limited to: word segmentation, part of speech prediction, conversion of text to phonemes, etc.

[0183] The phonetic information extracted by the text encoding module includes but is not limited to: syntactic information and / or grammatical information.

[0184] The phoneme filtering module is located at the back end of the text encoding module and is used to filter out the tone and rhythm in the text encoding sequence and retain the second phoneme sequence composed of initials and finals.

[0185] The duration contained in the first duration sequence is the pronunciation duration of a single phoneme.

[0186] The attention mechanism module processes the first sequence based on the attention mechanism and converts the first sequence into the second sequence, thereby obtaining the frame length of the acoustic features of each phoneme, which can ultimately be used by the acoustic decoding module to obtain the acoustic features of the synthesized speech. Finally, the speech synthesis module will use the acoustic features to convert the text data into synthesized speech data.

[0187] Synthesized speech data is a type of audio data.

[0188] In one embodiment, the method further comprises:

[0189] When the target speech data is used to train the basic model, a phoneme classification module is used to perform phoneme classification according to a second phoneme sequence output by a phoneme filtering model of the basic model to obtain a phoneme classification result;

[0190] Determine a loss value according to a difference between the phoneme classification result and the label corresponding to the text data;

[0191] According to the loss value, the model parameters of the basic model are adjusted to obtain the speech synthesis model.

[0192] In one embodiment, the phoneme classification module is not included in the synthetic speech model during the synthetic speech model use stage.

[0193] In another embodiment, the phoneme classification module is connected to the back end of the phoneme filtering module during the training phase, and is used to calculate the loss value of the phoneme processing of the current model based on the phonemes output by the phoneme filtering module.

[0194] Exemplarily, if the loss value is greater than the loss threshold, the training is continued using the sample set, otherwise the training can be stopped.

[0195] As another example, if the loss value has reached the minimum value, the training can also be stopped, otherwise the training continues.

[0196] For example, the phoneme classification result can distinguish the specific categories of each phoneme, which categories include but are not limited to: initial consonants and / or finals, etc. The labels corresponding to the text data can be labels marked by experts, etc.

[0197] The difference between the phoneme classification result and the label corresponding to the text data may include at least one of the following:

[0198] The number of matches between the phoneme classification and the label corresponding to the phoneme classification result;

[0199] The matching rate between the phoneme classification and the label corresponding to the phoneme classification result;

[0200] The number of phoneme classification and label errors corresponding to the phoneme classification results.

[0201] The above can all be used to confirm the difference between the phoneme classification result and the label corresponding to the text data. If the difference is greater, the loss value is greater, and if the difference is smaller, the loss value is smaller.

[0202] The contents involved in the above embodiment are described below in conjunction with a preferred embodiment.

[0203] Figure 4 This is a flow chart of speech synthesis according to an embodiment of the present application. The speech synthesis method may include:

[0204] The electronic device first obtains the voice data and user attribute information of the speech-impaired user.

[0205] For example, the speech-impaired user is a person with a speech barrier, that is, he cannot speak like a normal person, and can only send out voice data such as "ah ba ah ba". The present invention can be applied in a scenario where a speech-impaired user A makes a voice call with a user B (who may be a user of a speech accessibility function) by means of an artificial intelligence (AI) telephone assistant, and the specific steps are as follows:

[0206] When the device receives the voice message from user B, the AI ​​phone assistant performs voice recognition on user B's voice and converts the recognition result into voice text presentation;

[0207] After seeing the voice text of user B, speech-impaired user A enters the text data of what he wants to express into the text input area and clicks send. The text data can be the text reply intelligently generated by the AI ​​telephone assistant based on the voice text of user B, or it can be the text data edited by user A himself through the input keyboard.

[0208] The AI ​​telephone assistant feeds the input text data into the user's pre-customized speech synthesis model, which converts the text data into synthesized speech data. This allows speech-impaired user A to convey the content he wants to express to the other party (user B) in his customized voice, allowing for barrier-free communication.

[0209] The following part of the embodiment of the present disclosure provides a specific implementation scheme for determining the target voice:

[0210] S1, extracting sound parameters. Specifically, if the speech-impaired user can make simple sounds, then obtaining speech data of the speech-impaired user, such as speech data of “ah ba ah ba ah ba”;

[0211] S2, analyze the acquired speech data and extract acoustic parameters (e.g., Fbank features). After the audio is pre-emphasized, framed and windowed, Fourier transformed, energy spectrum calculated, Mel filtered, and logarithmized, the acoustic parameters (e.g., Fbank features) used for training can be obtained.

[0212] S3, obtaining voice parameters of multiple donors (candidate users), and extracting acoustic parameters (eg, Fbank features) using the same method as above.

[0213] S4, respectively calculating the timbre similarities between the speech-impaired user and multiple donors, and selecting the timbre of the donors whose similarity values ​​meet the threshold range according to a preset similarity threshold as the initial screening result.

[0214] Screening out the timbre of the donor corresponding to the similarity value within the threshold range according to a preset similarity threshold as the initial screening result may include: inputting the voice parameters of the donor and the acoustic parameters of the deaf-mute into a pre-trained voiceprint recognition model, and outputting the timbre similarity value of the donor and the deaf-mute.

[0215] The voiceprint recognition model can adopt the X-VECTORS model, which is also obtained through machine learning and is a classification model of a neural network trained with a large amount of voice data from different speakers.

[0216] The classification model may include multiple frame-level time-delay neural network (TDNN) layers, a statistical pooling layer and two sentence-level fully connected layers, and a softmax layer.

[0217] The loss function used in training this classification model is CE cross entropy.

[0218] The timbre similarity score is calculated by inputting the fbank parameters extracted from the speech data of the recipient and the donor into the above-mentioned X-VECTORS model, using the encoding output by the first fully connected layer as its speech feature expression, and calculating the PLDA score of the two speech feature expressions. After the speech data of the recipient (speech-impaired user) and the donor (candidate user) are input into the model, a similarity value will be obtained. Based on the similarity value, the speech data of some donors are selected as further candidates.

[0219] S5, collect user attributes such as age, height and weight of the donors and the deaf-mute, and select several donors whose appearance features are closest to those of the deaf-mute from the initial screening results.

[0220] S6, discuss with the user and select a certain tone from the second screening results as the basic tone. And reach a consensus with the user and his friends on the tuning direction (for example, the user's own imagined voice is clean and steady, not sharp, loud, emotional, clear, weak but powerful; the user's friends imagine the user's voice to be firm, sunny, refreshing, transparent, and gentle)

[0221] S7, based on the feedback from users and their friends, uses the sox tool to adjust the equalizers of different frequency bands, adjust the timbre, obtain the target timbre that the user wants, and finally confirm with the user.

[0222] Multi-dimensional matching from timbre to sound parameters makes the synthesized speech data have sound characteristics such as sunshine, vitality, and the presence or absence of local accent.

[0223] The split-band equalizer includes:

[0224] S8, according to the target timbre determined by the user, training a speech synthesis model for the user that can synthesize the target timbre.

[0225] The speech-impaired user evaluates the timbre of the target user's voice data, determines how to adjust the equalizer of different sound wave frequency bands of the target user's voice data, and obtains the third voice data. For example, if the speech-impaired user believes that his voice is clean and steady, the electronic device adjusts the frequency of the target user's voice data to 6kHz-16kHz through the audio processing tool sox. Alternatively, if the speech-impaired user believes that his voice should be loud and expressive, the electronic device adjusts the frequency of the target user's voice data to 200-600Hz through the audio processing tool sox. Until it meets the requirements of the speech-impaired user.

[0226] Figure 5 This is a schematic diagram of the structural model of the speech synthesis model in the embodiment of the present application, see Figure 5 As shown:

[0227] The speech synthesis model includes: text analysis module, text encoding module, phoneme filtering module, phoneme classification module, duration prediction module, attention mechanism module and acoustic encoding module. Each module is a neural network with different model structures and has different functions.

[0228] The electronic device processes the acquired text data of the speech-impaired user through a text analysis module, and converts the input data into a phoneme sequence suitable for a speech synthesis model, namely, first input data.

[0229] The text analysis module processes the input data in the following ways: word segmentation, part-of-speech prediction, prosodic word prediction, prosodic phrase prediction, intonation phrase prediction, text-to-phoneme conversion, etc.

[0230] The text encoding module uses a neural network to process the phoneme sequence obtained from the input data of the speech-impaired user processed by the text analysis module, and extracts linguistic information from the phoneme sequence.

[0231] Linguistic information refers to linguistic feature information extracted based on pronunciation annotation and prosodic annotation, such as phoneme sequence, tone and pause.

[0232] The phoneme filtering module is used to filter the tone and rhythm marks in the phoneme sequence that extracts linguistic information, leaving only the phonemes composed of initial consonants and finals. After obtaining the filtered phonemes, the phonemes are classified by the phoneme classification module and the duration prediction module analyzes the duration corresponding to each phoneme. Among them, the phoneme classification module is used to classify the phonemes composed of only initial consonants and finals obtained from the phoneme filtering module according to the total phoneme category, cluster the same phonemes into different phoneme categories, and play a role in information enhancement. The duration prediction module and the phoneme classification module can process the phonemes synchronously.

[0233] The attention mechanism module is used to predict which frames of acoustic parameters are extracted from the audio for each phoneme.

[0234] The acoustic decoding module decodes the information comprehensively obtained from the above-mentioned phoneme filtering module, phoneme classification module, duration prediction module, and attention mechanism module into acoustic parameters required for subsequent conversion into audio, such as Mel cepstrum parameters.

[0235] The first output data is the Mel spectrum parameter that can be finally converted into the first audio data, and a loss function needs to be calculated between the predicted Mel spectrum parameter and the actual Mel spectrum parameter. The fourth input data is the offset Mel cepstrum, which is the loss function that needs to be calculated.

[0236] The second output data is the time length corresponding to each phoneme. A loss function needs to be calculated between the predicted time length and the actual time length.

[0237] The third output data is also the time length corresponding to each phoneme. The third output data also includes the acoustic parameters extracted from the audio frames that express each phoneme. A loss function needs to be calculated between the predicted duration and the actual duration to strengthen the learning of duration.

[0238] The fourth output data is the result of phoneme classification. A loss function needs to be calculated between the predicted phoneme classification and the actual phoneme classification.

[0239] The above model converts input data into synthesized speech data by utilizing the flexibility of the attention mechanism and the displayed duration information. At the same time, based on the monotonic attention mechanism, the duration information of each frame is also utilized, so that the synthesized speech data can achieve a higher degree of naturalness.

[0240] exist Figure 5 The dashed boxes or dashed arrows represent those involved in the model training process, while the solid boxes and solid arrows are used after the speech synthesis model is launched.

[0241] At this point, the process of synthesizing speech using the above-mentioned speech synthesis method is completed.

[0242] Based on the same inventive concept, an embodiment of the present application provides a speech synthesis device, which includes several functional units for implementing the above-mentioned speech synthesis method.

[0243] Figure 6 This is a structural diagram of a message processing device in an embodiment of the present application, see Figure 6 As shown, the speech synthesis device 600 may include:

[0244] An acquisition module 601 is used to acquire text data input by a speech-impaired user;

[0245] A synthesis module 602, used for inputting the text data into a speech synthesis model to obtain synthesized speech data;

[0246] The synthesized speech data and the text data have the same language content, and the synthesized speech data is audio data; the speech synthesis model is obtained by training a basic model using a sample set through a machine learning method, and the sample set includes: the target speech feature and text data corresponding to the language content of the target speech feature, wherein the target speech feature is a pre-selected speech feature of a target user.

[0247] In some possible implementations, the sample set includes a plurality of samples, each of which includes a speech-text pair;

[0248] Each of the speech-text pairs includes a phoneme and a Mel-spectrogram feature parameter corresponding to the phoneme, wherein the phoneme is obtained by preprocessing the text.

[0249] In some possible implementations, the target speech feature is obtained by extracting features from target speech data.

[0250] The device also includes:

[0251] The selection module is used to select from a plurality of candidate voice data to obtain the target voice data.

[0252] In some possible implementations, the selection module is specifically used to:

[0253] Inputting the speech data of the speech-impaired user and the plurality of candidate speech data into a voiceprint recognition model to obtain timbre similarities between the plurality of candidate speech data and the speech data of the speech-impaired user;

[0254] According to the timbre similarity, one target voice data is determined from the plurality of candidate voice data.

[0255] In some possible implementations, the selection module is specifically used to:

[0256] Acquiring voice data of the speech-impaired user;

[0257] Analyze the speech data of the speech-impaired user to obtain a first speech feature;

[0258] Analyze the plurality of candidate voice data to obtain a second voice feature;

[0259] The second voice feature and the first voice feature are selected from the multiple candidate voice data and input into the voiceprint recognition model, and the voice data of the speech-impaired user and the multiple candidate voice data are input into the voiceprint recognition model to obtain the timbre similarity between the multiple candidate voice data and the voice data of the speech-impaired user.

[0260] In some possible implementations, the selection module is specifically used to:

[0261] According to the timbre similarity, selecting candidate voice data whose similarity is in a first similarity interval from the plurality of candidate voice data as the first voice data;

[0262] According to the similarity between the user attribute corresponding to the first voice data and the user attribute of the speech-impaired user, selecting the first voice data whose similarity is in a second similarity interval from the first voice data;

[0263] The target voice data is obtained according to the first voice data.

[0264] In some possible implementations, after obtaining the first voice data, the selection module is specifically configured to:

[0265] Based on the speech synthesis request, determining the second speech data; wherein the speech synthesis request is used to indicate the selection result of the user, and the second speech data is the speech data in the first speech data that matches the speech synthesis request;

[0266] The target voice data is obtained according to the second voice data.

[0267] In some possible implementations, the selection module is specifically used to:

[0268] detecting an adjustment instruction acting on the second voice data;

[0269] According to the adjustment instruction, the second voice data is adjusted to obtain the target voice data, wherein the adjustment instruction is used to instruct to adjust the proportion of sound waves in different frequency bands in the second voice data.

[0270] In some possible implementations, the speech synthesis model includes:

[0271] A text encoding module, used to extract linguistic information from a first phoneme sequence and obtain a text encoding sequence representing the linguistic information; wherein the first phoneme sequence is obtained by preprocessing text data;

[0272] A phoneme filtering module, used for filtering the text encoding sequence to obtain a second phoneme sequence composed of initial consonants and / or finals;

[0273] a duration prediction module, configured to obtain a first duration sequence according to the second phoneme sequence, wherein the first duration sequence includes: the duration of each phoneme in the second phoneme sequence; wherein the second phoneme sequence and the first duration sequence are added to obtain a first sequence;

[0274] The attention mechanism module is used to receive a first sequence, align the first sequence with each element of an acoustic feature sequence obtained based on the first sequence, and obtain a second sequence; the second sequence includes: a frame length of an acoustic feature of each phoneme; wherein the elements of the acoustic feature sequence include: a Mel-spectrogram feature parameter corresponding to the phoneme;

[0275] The acoustic decoding module is used to obtain the acoustic features of the synthesized speech data according to the first sequence and the second sequence.

[0276] In some possible implementations, the device further includes:

[0277] A classification module, configured to perform phoneme classification according to a second phoneme sequence output by a phoneme filtering model of the basic model using a phoneme classification module when the basic model is trained using the target speech data, so as to obtain a phoneme classification result;

[0278] A determination module, used to determine a loss value according to a difference between the phoneme classification result and a label corresponding to the text data;

[0279] An adjustment module is used to adjust the model parameters of the basic model according to the loss value to obtain the speech synthesis model.

[0280] It should be noted that the specific implementation process of the first acquisition module 601 and the generation module 602 can be referred to Figures 1 to 4 For the sake of brevity, the detailed description of the embodiments will not be repeated here.

[0281] Based on the same inventive concept, an embodiment of the present application provides an electronic device, which may be consistent with the speech synthesis method described in one or more of the above embodiments.

[0282] Figure 7 This is a schematic diagram of the structure of an electronic device in an embodiment of the present application, referring to Figure 7 , the electronic device 800 may include one or more of the following components: a processing component 802 , a memory 804 , a power component 806 , a multimedia component 808 , a multimedia data component 810 , an input / output (I / O) interface 812 , a sensor component 814 , and a communication component 816 .

[0283] The processing component 802 generally controls the overall operation of the electronic device 800, such as operations associated with display, phone calls, data communications, camera operations, and recording operations. The processing component 802 may include one or more processors 820 to execute instructions to complete all or part of the steps of the above-mentioned method. In addition, the processing component 802 may include one or more modules to facilitate the interaction between the processing component 802 and other components. For example, the processing component 802 may include a multimedia module to facilitate the interaction between the multimedia component 808 and the processing component 802.

[0284] The memory 804 is configured to store various types of data to support operations on the device 800. Examples of such data include instructions for any application or method operating on the electronic device 800, contact data, phone book data, messages, pictures, videos, etc. The memory 804 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.

[0285] The power component 806 provides power to the various components of the electronic device 800. The power component 806 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the electronic device 800.

[0286] The multimedia component 808 includes a screen that provides an output interface between the electronic device 800 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touch, slide, and gestures on the touch panel. The touch sensor may not only sense the boundaries of the touch or slide action, but also detect the duration and pressure associated with the touch or slide operation. In some embodiments, the multimedia component 808 includes a front camera and / or a rear camera. When the device 800 is in an operating state, such as a shooting state or a video state, the front camera and / or the rear camera may receive external multimedia data. Each front camera and rear camera may be a fixed optical lens system or have a focal length and optical zoom capability.

[0287] The multimedia data component 810 is configured to output and / or input multimedia data signals. For example, the multimedia data component 810 includes a microphone (MIC), and when the electronic device 800 is in an operating state, such as a call state, a recording state, and a voice recognition state, the microphone is configured to receive external multimedia data signals. The received multimedia data signal can be further stored in the memory 804 or sent via the communication component 816. In some embodiments, the multimedia data component 810 also includes a speaker for outputting multimedia data signals.

[0288] I / O interface 812 provides an interface between processing component 802 and peripheral interface modules, such as keyboards, click wheels, buttons, etc. These buttons may include but are not limited to: home button, volume button, start button, and lock button.

[0289] The sensor assembly 814 includes one or more sensors for providing various aspects of status assessment for the electronic device 800. For example, the sensor assembly 814 can detect the open / closed state of the device 800, the relative positioning of the components, such as the display and keypad of the electronic device 800, and the sensor assembly 814 can also detect the position change of the electronic device 800 or a component of the electronic device 800, the presence or absence of contact between the user and the electronic device 800, the orientation or acceleration / deceleration of the electronic device 800 and the temperature change of the electronic device 800. The sensor assembly 814 can include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor assembly 814 can also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor assembly 814 can also include an acceleration sensor, a gyroscope sensor, a magnetic sensor, a pressure sensor or a temperature sensor.

[0290] The communication component 816 is configured to facilitate wired or wireless communication between the electronic device 800 and other devices. The electronic device 800 can access a wireless network based on a communication standard, such as Wi-Fi, 2G or 3G, or a combination thereof. In an exemplary embodiment, the communication component 816 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 816 also includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology and other technologies.

[0291] In an exemplary embodiment, the electronic device 800 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the above methods.

[0292] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 804 including instructions, and the above instructions can be executed by the processor 820 of the electronic device 800 to complete the above method. For example, the non-transitory computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.

[0293] An embodiment of the present disclosure provides a computer storage medium, which may be a non-temporary computer-readable storage medium. When the instructions in the storage medium are executed by a processor of a server or a terminal, the electronic device can execute the speech synthesis method provided by any of the aforementioned technical solutions.

[0294] Get text data input by speech-impaired users;

[0295] Inputting the text data into a speech synthesis model to obtain synthesized speech data; wherein the synthesized speech data and the text data have the same language content, and the synthesized speech data is audio data;

[0296] The speech synthesis model is obtained by training a basic model using a sample set through a machine learning method, and the sample set includes: the target speech feature and text data corresponding to the language content of the target speech feature, wherein the target speech feature is a pre-selected speech feature of a target user.

[0297] In one embodiment, the sample set includes a plurality of samples, each of which includes a speech-text pair;

[0298] Each of the speech-text pairs includes a phoneme and a Mel-spectrogram feature parameter corresponding to the phoneme, wherein the phoneme is obtained by preprocessing the text.

[0299] In one embodiment, the target speech feature is obtained by extracting features from the target speech data.

[0300] The method further includes: selecting from a plurality of candidate voice data to obtain the target voice data.

[0301] In one embodiment, the step of selecting from a plurality of candidate voice data to obtain the target voice data includes:

[0302] Inputting the speech data of the speech-impaired user and the plurality of candidate speech data into a voiceprint recognition model to obtain timbre similarities between the plurality of candidate speech data and the speech data of the speech-impaired user;

[0303] According to the timbre similarity, one target voice data is determined from the plurality of candidate voice data.

[0304] In one embodiment, determining the target voice data from the plurality of candidate voice data according to the timbre similarity comprises:

[0305] Acquiring voice data of the speech-impaired user;

[0306] Analyzing the speech data of the speech-impaired user to obtain a first speech feature;

[0307] Analyze the plurality of candidate voice data to obtain a second voice feature;

[0308] The step of inputting the speech data of the speech-impaired user and the plurality of candidate speech data into a voiceprint recognition model to obtain timbre similarities between the plurality of candidate speech data and the speech data of the speech-impaired user includes:

[0309] The second voice feature and the first voice feature are selected from the multiple candidate voice data and input into the voiceprint recognition model, and the voice data of the speech-impaired user and the multiple candidate voice data are input into the voiceprint recognition model to obtain the timbre similarity between the multiple candidate voice data and the voice data of the speech-impaired user.

[0310] In one embodiment, selecting the target voice data from the plurality of candidate voice data according to the timbre similarity comprises:

[0311] According to the timbre similarity, selecting candidate voice data whose similarity is in a first similarity interval from the plurality of candidate voice data as the first voice data;

[0312] According to the similarity between the user attribute corresponding to the first voice data and the user attribute of the speech-impaired user, selecting the first voice data whose similarity is in a second similarity interval from the first voice data;

[0313] The target voice data is obtained according to the first voice data.

[0314] Understandably,

[0315] After obtaining the first voice data, the method further includes:

[0316] Based on the speech synthesis request, determining the second speech data; wherein the speech synthesis request is used to indicate the selection result of the user, and the second speech data is the speech data in the first speech data that matches the speech synthesis request;

[0317] The target voice data is obtained according to the second voice data.

[0318] It can be understood that obtaining the target voice data according to the second voice data includes:

[0319] detecting an adjustment instruction acting on the second voice data;

[0320] According to the adjustment instruction, the second voice data is adjusted to obtain the target voice data, wherein the adjustment instruction is used to instruct to adjust the proportion of sound waves in different frequency bands in the second voice data.

[0321] It can be understood that the speech synthesis model includes:

[0322] A text encoding module, used to extract linguistic information from a first phoneme sequence and obtain a text encoding sequence representing the linguistic information; wherein the first phoneme sequence is obtained by preprocessing text data;

[0323] A phoneme filtering module, used for filtering the text encoding sequence to obtain a second phoneme sequence consisting of initial consonants and / or finals;

[0324] a duration prediction module, configured to obtain a first duration sequence according to the second phoneme sequence, wherein the first duration sequence includes: the duration of each phoneme in the second phoneme sequence; wherein the second phoneme sequence and the first duration sequence are added to obtain a first sequence;

[0325] The attention mechanism module is used to receive a first sequence, align the first sequence with each element of an acoustic feature sequence obtained based on the first sequence, and obtain a second sequence; the second sequence includes: a frame length of an acoustic feature of each phoneme; wherein the elements of the acoustic feature sequence include: a Mel-spectrogram feature parameter corresponding to the phoneme;

[0326] The acoustic decoding module is used to obtain the acoustic features of the synthesized speech data according to the first sequence and the second sequence.

[0327] In one embodiment, the method further comprises:

[0328] When the target speech data is used to train the basic model, a phoneme classification module is used to perform phoneme classification according to a second phoneme sequence output by a phoneme filtering model of the basic model to obtain a phoneme classification result;

[0329] Determine a loss value according to a difference between the phoneme classification result and the label corresponding to the text data;

[0330] According to the loss value, the model parameters of the basic model are adjusted to obtain the speech synthesis model.

[0331] Those skilled in the art will appreciate that the sequence numbers of the steps in the above embodiments do not imply a sequence of execution, and the execution sequence of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0332] The above is only an exemplary embodiment of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by a person skilled in the art within the technical scope disclosed in the present application should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.

Claims

1. A speech synthesis method, characterized in that: The method comprises: Get text data input by speech-impaired users; Inputting the text data into a speech synthesis model to obtain synthesized speech data; wherein the synthesized speech data and the text data have the same language content, and the synthesized speech data is audio data; The speech synthesis model is obtained by training a basic model using a sample set by a machine learning method, wherein the sample set includes: a target speech feature and text data corresponding to the language content of the target speech feature, wherein the target speech feature is a speech feature of target speech data; the target speech data includes speech data of a pre-selected target user; The target voice data is determined in the following manner: the voice data of the speech-impaired user and multiple candidate voice data are input into a voiceprint recognition model to obtain the timbre similarity between the multiple candidate voice data and the voice data of the speech-impaired user; and one of the target voice data is determined from the multiple candidate voice data based on the timbre similarity.

2. The method according to claim 1, characterized in that The sample set includes a plurality of samples, each of which includes a speech-text pair; Each of the speech-text pairs includes a phoneme and a Mel-spectrogram feature parameter corresponding to the phoneme, wherein the phoneme is obtained by preprocessing the text.

3. The method according to claim 1, characterized in that The step of determining the target voice data from the plurality of candidate voice data according to the timbre similarity comprises: Acquiring voice data of the speech-impaired user; Analyzing the speech data of the speech-impaired user to obtain a first speech feature; Analyze the plurality of candidate voice data to obtain a second voice feature; The step of inputting the speech data of the speech-impaired user and the plurality of candidate speech data into a voiceprint recognition model to obtain timbre similarities between the plurality of candidate speech data and the speech data of the speech-impaired user includes: The second voice feature and the first voice feature are selected from the plurality of candidate voice data and input into the voiceprint recognition model to obtain timbre similarity between the plurality of candidate voice data and the voice data of the speech-impaired user.

4. The method according to claim 1 or 3, characterized in that: The step of selecting, according to the timbre similarity, from the plurality of candidate voice data to obtain the target voice data comprises: According to the timbre similarity, selecting candidate voice data whose similarity is in a first similarity interval from the plurality of candidate voice data as the first voice data; According to the similarity between the user attribute corresponding to the first voice data and the user attribute of the speech-impaired user, selecting the first voice data whose similarity is in a second similarity interval from the first voice data; The target voice data is obtained according to the first voice data.

5. The method according to claim 4, characterized in that After obtaining the first voice data, the method further includes: Based on the speech synthesis request, determining the second speech data; wherein the speech synthesis request is used to indicate the selection result of the user, and the second speech data is the speech data in the first speech data that matches the speech synthesis request; The target voice data is obtained according to the second voice data.

6. The method according to claim 5, wherein: The step of obtaining the target voice data according to the second voice data includes: detecting an adjustment instruction acting on the second voice data; According to the adjustment instruction, the second voice data is adjusted to obtain the target voice data, wherein the adjustment instruction is used to instruct to adjust the proportion of sound waves in different frequency bands in the second voice data.

7. The method according to claim 1, characterized in that The speech synthesis model comprises: A text encoding module, used to extract linguistic information from a first phoneme sequence and obtain a text encoding sequence representing the linguistic information; wherein the first phoneme sequence is obtained by preprocessing text data; A phoneme filtering module, used for filtering the text encoding sequence to obtain a second phoneme sequence consisting of initial consonants and / or finals; a duration prediction module, configured to obtain a first duration sequence according to the second phoneme sequence, wherein the first duration sequence includes: the duration of each phoneme in the second phoneme sequence; wherein the second phoneme sequence and the first duration sequence are added to obtain a first sequence; The attention mechanism module is used to receive a first sequence, align the first sequence with each element of an acoustic feature sequence obtained based on the first sequence, and obtain a second sequence; the second sequence includes: a frame length of an acoustic feature of each phoneme; wherein the elements of the acoustic feature sequence include: a Mel-spectrogram feature parameter corresponding to the phoneme; The acoustic decoding module is used to obtain the acoustic features of the synthesized speech data according to the first sequence and the second sequence.

8. The method according to claim 7, characterized in that The method further comprises: When the target speech data is used to train the basic model, a phoneme classification module is used to perform phoneme classification according to a second phoneme sequence output by a phoneme filtering model of the basic model to obtain a phoneme classification result; Determine a loss value according to a difference between the phoneme classification result and the label corresponding to the text data; According to the loss value, the model parameters of the basic model are adjusted to obtain the speech synthesis model.

9. A speech synthesis device, characterized in that: The device comprises: An acquisition module, used to acquire text data input by a speech-impaired user; A synthesis module, used for inputting the text data into a speech synthesis model to obtain synthesized speech data; wherein the synthesized speech data and the text data have the same language content, and the synthesized speech data is audio data; The speech synthesis model is obtained by training a basic model using a sample set by a machine learning method, wherein the sample set includes: a target speech feature and text data corresponding to the language content of the target speech feature, wherein the target speech feature is a speech feature of target speech data; the target speech data includes speech data of a pre-selected target user; The selection module is used to input the speech data of the speech-impaired user and multiple candidate speech data into a voiceprint recognition model to obtain the timbre similarity between the multiple candidate speech data and the speech data of the speech-impaired user; and determine one target speech data from the multiple candidate speech data according to the timbre similarity.

10. An electronic device, characterized in that: include: a memory for storing processor-executable instructions; A processor connected to the memory; wherein the processor is configured to execute the method according to any one of claims 1 to 8.

11. A non-transitory computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, implement the method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Speech synthesis model training method, speech synthesis method and device thereof

    CN113393828A