Audio synthesis method, device, electronic device and readable storage medium
By obtaining and utilizing the pronunciation characteristics parameters of the target pronunciation person, individually synthesize the pronunciation, the problem that pronunciation synthesis in the prior art cannot reflect the user's voice characteristics is solved, and the pronunciation synthesis effect is achieved that is closer to the user's pronunciation characteristics.
Patent Information
- Application Number
- CN202111148956.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-09-28
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2041-09-28
AI Technical Summary
The existing voice synthesis technology cannot effectively reflect the voice characteristics of different users, resulting in poor synthesized voice effects.
By obtaining the target information and the pronunciation characteristic parameters of the target pronunciation person, including the speech speed reference vector, the pause length reference vector and the style vector, the acoustic feature information is determined and converted to generate the target audio data corresponding to the target information.
The target audio data is personalized according to the speaking style and rhythm characteristics of different pronunciators, so that the generated audio data is closer to the pronunciation characteristics of the target pronunciation, and the effect of pronunciation synthesis is improved.
Smart Images

Figure CN113870828B_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the technical field of speech synthesis, and specifically relates to an audio synthesis method, device, electronic device, and readable storage medium. Background Art
[0002] Text to Speech (TTS) technology refers to the technology of converting text information into speech information. Personal Text to Speech refers to the technology of synthesizing speech that conforms to the speaking style of a specific person based on TTS speech technology after recording certain speech segments of a person through a recording device.
[0003] However, in current speech synthesis technology, the synthesized speech cannot reflect the vocal characteristics of different users, and the synthesis effect is poor. Summary of the Invention
[0004] The purpose of the embodiments of this application is to provide an audio synthesis method, device, electronic device, and readable storage medium, which can solve the problem that in speech synthesis technology, the synthesized speech cannot reflect the vocal characteristics of different users and the synthesis effect is poor.
[0005] In a first aspect, the embodiments of this application provide an audio synthesis method, which includes:
[0006] Obtain target information;
[0007] Obtain the prosodic characteristic parameters of the target speaker, where the prosodic characteristic parameters include a speech rate reference vector, a pause length reference vector, and a style vector;
[0008] Determine acoustic feature information according to the target information and the prosodic characteristic parameters;
[0009] Convert the acoustic feature information to generate target audio data corresponding to the target information.
[0010] In a second aspect, the embodiments of this application provide an audio synthesis device, which includes:
[0011] A first acquisition module, configured to acquire target information;
[0012] A second acquisition module, configured to acquire the prosodic characteristic parameters of the target speaker, where the prosodic characteristic parameters include a speech rate reference vector, a pause length reference vector, and a style vector;
[0013] A first determination module, configured to determine acoustic feature information according to the target information and the prosodic characteristic parameters;
[0014] A generation module, configured to convert the acoustic feature information to generate target audio data corresponding to the target information.
[0015] In a third aspect, an embodiment of the present application provides an electronic device, which includes a processor, a memory, and a program or instruction stored on the memory and executable on the processor. When the program or instruction is executed by the processor, the steps of the method described in the first aspect are implemented.
[0016] In a fourth aspect, an embodiment of the present application provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, the steps of the method described in the first aspect are implemented.
[0017] In a fifth aspect, an embodiment of the present application provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor, and the processor is configured to run a program or instruction to implement the method described in the first aspect.
[0018] In the embodiment of the present application, the target information and the prosody characteristic parameters of the target speaker are obtained. According to the target information and the prosody characteristic parameters of the target speaker, the acoustic feature information is determined, and the acoustic feature information is converted to generate target audio data corresponding to the target information. In this way, the prosody characteristic parameters of the target speaker can be obtained, and the prosody characteristic parameters of the target speaker can be used to affect the generation of the acoustic feature information. The target audio data can be personalized synthesized according to the speaking styles and prosody characteristics of different speakers, so that the generated target audio data is closer to the pronunciation characteristics of the target speaker. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 is a schematic flowchart of an audio synthesis method provided by an embodiment of the present application;
[0020] Figure 2 is a schematic diagram of a style vector encoding and decoding model provided by an embodiment of the present application;
[0021] Figure 3 is a schematic diagram of the synthesis process of target audio data provided by an embodiment of the present application;
[0022] Figure 4 is a schematic structural diagram of an audio synthesis device provided by an embodiment of the present application;
[0023] Figure 5 is a schematic structural diagram of an electronic device provided by an embodiment of the present application;
[0024] Figure 6 is a schematic hardware structure diagram of an electronic device provided by an embodiment of the present application to implement the present application. Detailed implementation manners
[0025] Next, the technical solutions in the embodiments of the present application will be clearly described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art belong to the scope of protection of the present application.
[0026] The terms "first", "second", etc. in the specification and claims of the present application are used to distinguish similar objects, rather than to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first", "second", etc. are usually of the same type, and the number of objects is not limited. For example, the first object can be one or multiple. In addition, "and / or" in the specification and claims means at least one of the connected objects, and the character " / " generally indicates an "or" relationship between the associated objects before and after.
[0027] Next, in conjunction with the accompanying drawings, the audio synthesis method provided in the embodiments of the present application will be described in detail through specific embodiments and their application scenarios.
[0028] Please refer to Figure 1 , which is an audio synthesis method provided in the embodiments of the present application. This method is applied to an electronic device, and this method may include step 1100-step 1400, which will be described in detail below.
[0029] Step 1100, obtain target information.
[0030] In this embodiment, the target information may be text information input by the user that needs to be converted into audio data. The target information may be text information. For example, a sentence input by the user through the text input method. The target information may also be voice information. For example, a sentence recorded by the user.
[0031] Step 1200, obtain the prosody characteristic parameters of the target speaker, where the prosody characteristic parameters include a speech rate reference vector, a pause length reference vector, and a style vector.
[0032] In this embodiment, the prosody characteristic parameters of the target speaker can reflect the prosody characteristics of the target speaker's speech. The prosody characteristics may be the rhyme and rhythm in reading aloud. Specifically, they may include long pauses, short pauses, and breathing positions in sentence reading, and may also include the speed, stress, etc. of sentence reading.
[0033] The prosody characteristic parameters may include a speech rate reference vector, a pause length reference vector, and a style vector.
[0034] The speech rate reference vector can represent the speech rate of the target speaker. For different speakers, the speed of reading sentences is different, and the speech rate reference vectors are different.
[0035] In some alternative embodiments, obtaining the speech rate reference vector may further include: step 2100 - step 2300.
[0036] Step 2100, obtain the historical audio data of the target speaker.
[0037] In this embodiment, the historical audio data of the target speaker may be the audio data of the target speaker stored in the electronic device. For example, the pre-recorded audio data of the target speaker. Also for example, the voice chat records of the target speaker stored in the instant messaging application.
[0038] In specific implementation, receive a first input of the target speaker, and in response to the first input, obtain the historical audio data of the target speaker. The first input may be an input to the storage directory of the historical audio data. It should be noted that after obtaining the historical audio data of the target speaker, the historical audio data may be screened to screen the historical audio data that meets the requirements. In this way, the audio data with a signal-to-noise ratio that does not meet the requirements and the audio data of multiple people speaking can be screened out, improving the acquisition efficiency of prosodic features.
[0039] Step 2200, determine the first average speech rate of the target speaker according to the historical audio data.
[0040] The first average speech rate may be the average pronunciation duration of each phoneme. Among them, for Chinese, a phoneme may be, for example, an initial consonant or a final vowel in a sentence.
[0041] The first average speech rate may be determined according to the duration of the target sentence and the number of phonemes included in the target sentence. In specific implementation, obtain the target sentence from the historical audio data, and use the ratio of the duration of the target sentence to the number of phonemes included in the target sentence as the first average speech rate.
[0042] Step 2300, determine the speech rate reference vector according to the first average speech rate and the preset average speech rate.
[0043] The preset average speech rate may reflect the reading rhythm that conforms to most users. Exemplarily, the preset average speech rate may be an average speech rate determined based on big data. The big data may include, for example, the audio data of multiple users. In specific implementation, according to the first average speech rate S and the preset average speech rate S', determining the speech rate reference vector A may be to determine the speech rate reference vector according to the ratio of the first average speech rate S to the preset average speech rate S'.
[0044] In this embodiment, by using the historical audio data of the target speaker stored in the electronic device to obtain the speech rate reference vector of the target speaker, more audio data can be used to obtain the personalized prosody characteristic parameters of the target speaker. Combining with the subsequent steps, audio data more in line with the speaking habits of the target speaker can be generated. In addition, by extracting the speech rate reference vector of the target speaker through the electronic device, user data leakage can be avoided, and the security of the interaction can be improved.
[0045] In this embodiment, the pause length reference vector can represent the pause habit of the target speaker when reading a sentence. For different speakers, the pause positions during the sentence reading process are different, and the pause length reference vectors are different. For example, some speakers are used to pausing and taking a breath every two or three characters when reading a sentence, while some speakers pause and take a breath after reading the whole sentence.
[0046] In some alternative embodiments, obtaining the pause length reference vector may further include: step 3100 - step 3300.
[0047] Step 3100, obtain the historical audio data of the target speaker.
[0048] In this embodiment, the historical audio data of the target speaker may be the audio data of the target speaker stored in the electronic device. For example, the pre-recorded audio data of the target speaker. Also for example, the voice chat records of the target speaker stored in the instant messaging application.
[0049] Step 3200, determine the pause probabilities corresponding to different syllable lengths according to the historical audio data.
[0050] In this embodiment, for Chinese, the syllable length may be the number of Chinese characters read. The pause probabilities corresponding to different syllable lengths are shown in the following table.
[0051]
[0052]
[0053] Step 3300, determine the pause length reference vector according to the pause probabilities corresponding to different syllable lengths.
[0054] In this embodiment, by using the historical audio data of the target speaker stored in the electronic device to obtain the pause length reference vector of the target speaker, more audio data can be used to obtain the personalized prosody characteristic parameters of the target speaker. Combining with the subsequent steps, audio data more in line with the speaking habits of the target speaker can be generated. In addition, by extracting the pause length reference vector of the target speaker through the electronic device, user data leakage can be avoided, and the security of the interaction can be improved.
[0055] In this embodiment, the style vector may represent the prosodic style of the speaker. For example, it can be a natural communication style, a broadcasting style, or a style of reading a novel aloud. The style vector can be obtained through clustering analysis of multiple speakers. Speakers with similar style vector distances have similar prosodic styles in their pronunciations.
[0056] In this embodiment, the style vector of the target speaker can be obtained based on an encoding and decoding model. Taking Figure 2 the shown style vector encoding and decoding model as an example, from the historical audio data of the target speaker, audio features X and text feature parameters are extracted. The audio features X are input into the encoder (Encoder) 401 to obtain the style vector C. Then, the style vector C and the text feature parameters are input into the decoder (Decoder) 402 to output the audio features X'. After that, the parameters of each module are optimized so that the difference between the output audio features X' and the input audio features X is less than a preset threshold. Based on the optimized style vector encoding and decoding model, the style vector C of each piece of audio data is obtained i , and the average value of the style vectors C of each piece of audio data i is used as the style vector C of the target speaker.
[0057] In this embodiment, by using the historical audio data of the target speaker stored in the electronic device to obtain the style vector of the target speaker, personalized prosodic characteristic parameters of the target speaker can be obtained using a relatively large amount of audio data. Combining with subsequent steps, audio data more in line with the pronunciation characteristics of the target speaker can be generated. In addition, in this way, the prosodic characteristic parameters of the target speaker can be extracted offline, and the extraction process of the prosodic characteristic parameters and the training process of the acoustic model can be carried out independently, which can improve the training efficiency.
[0058] After step 1200, step 1300 is executed to determine the acoustic feature information according to the target information and the prosodic characteristic parameters.
[0059] The acoustic feature information can be the feature information input into the vocoder to generate audio data. Different types of acoustic feature information can be selected according to the requirements of the vocoder. For example, mel spectrum, pitch, mgc, etc.
[0060] In this embodiment, the acoustic feature information is related to both the text content and the reading habits and styles of the speaker. That is to say, the acoustic feature information is related to the text prosody of the text content itself and also related to the prosodic characteristics of the speaker. Based on this, the acoustic feature information is determined according to the target information and the prosodic characteristic parameters of the target speaker, so as to generate the target audio according to the acoustic feature information, making the target audio closer to the pronunciation characteristics of the target speaker.
[0061] In some embodiments of the present application, determining the acoustic feature information according to the target information and the prosody characteristic parameters includes: step 4100 - step 4500.
[0062] Step 4100, analyze the target information to obtain text feature parameters, where the text feature parameters include a first phoneme sequence and text prosody.
[0063] In this embodiment, since the acoustic feature information is related to the text content, based on this, it is necessary to obtain the text feature parameters of the target information, so as to combine the text feature parameters of the target information and the prosody characteristic parameters of the target speaker to generate acoustic feature information that conforms to the speaking characteristics of the target speaker.
[0064] The text feature parameters may include a first phoneme sequence and text prosody. The first phoneme sequence may be determined according to the word boundaries of the target information. The first phoneme sequence is determined based on the relevance between the text contents of the target information. The text feature parameters may also include a tone sequence, stress, etc.
[0065] In specific implementation, when the target information is text information, the text information is input into the text analysis module, and text feature parameters are output. The text analysis module may use traditional mode classification algorithms such as decision trees and ME, or may use algorithms such as BiLstm, Bert, and TCN of neural networks to perform sequence labeling tasks to obtain the final labeling result, that is, the text feature parameters.
[0066] It should be noted that when the target information is voice information, the voice information may be recognized to obtain the text information corresponding to the target information, and further text analysis is performed on the text information to obtain text feature parameters.
[0067] Step 4200, generate a second phoneme sequence according to the text prosody, the pause length reference vector, and the first phoneme sequence.
[0068] In some optional embodiments, generating the second phoneme sequence according to the text prosody, the pause length reference vector, and the first phoneme sequence may further include: generating corrected prosody information according to the text prosody and the pause length reference vector; generating a second phoneme sequence according to the corrected prosody information and the first phoneme sequence.
[0069] In specific implementation, taking the prosody probability corresponding to the text prosody as the node probability and the pause length reference vector as the path probability, using the dynamic programming algorithm to find the optimal path, that is, the corrected prosody information. Then, the corrected prosody information is merged with the first phoneme sequence to generate a second phoneme sequence, that is, a phoneme sequence containing prosody information.
[0070] Step 4300, determine a first audio feature according to the second phoneme sequence and the speech rate reference vector.
[0071] In some optional embodiments, the determining of the first audio feature according to the second phoneme sequence and the speech rate reference vector may further include: performing duration prediction based on the second phoneme sequence to obtain a first phoneme duration; adjusting the first phoneme duration according to the speech rate reference vector to obtain a second phoneme duration; expanding the second phoneme sequence according to the second phoneme duration to obtain a first audio feature.
[0072] The first phoneme duration may be the pronunciation duration of each phoneme predicted according to the second phoneme sequence. The second phoneme duration may be the pronunciation duration of each phoneme considering the speech rate of the target speaker.
[0073] In specific implementation, input the second phoneme sequence into a duration prediction module to predict the first phoneme duration. Then, the speech rate reference vector adjusts the first phoneme duration to obtain the second phoneme duration. Then, expand the second phoneme sequence according to the second phoneme duration to obtain the first audio feature after expanding the number of frames.
[0074] Step 4400, determine a second audio feature according to the first audio feature and the style vector.
[0075] The second audio feature may be the audio feature obtained after the style vector affects the first audio feature. The second audio feature is more in line with the speaking style of the target speaker.
[0076] Step 4500, based on an acoustic prediction model, determine the acoustic feature information according to the second audio feature. Wherein, the acoustic prediction model is used to obtain acoustic feature information according to the second audio feature.
[0077] In some embodiments of the present application, before the determining of the acoustic feature information based on the acoustic prediction model according to the second audio feature, the method further includes: obtaining first audio data of the target speaker, where the first audio data is audio data of the target speaker reading a preset text; performing model training based on the first audio data to obtain the acoustic prediction model.
[0078] After step 1300, execute step 1400 to convert the acoustic feature information to generate target audio data corresponding to the target information.
[0079] In specific implementation, the acoustic feature information is input into a vocoder, and after conversion, target audio data is obtained. Among them, different vocoders can be selected according to different deployment scenarios and service requirements. Exemplarily, it can be a traditional vocoder, for example, an LPC vocoder, a WORLD vocoder, etc. Exemplarily, it can also be a neural network vocoder, for example, an LPCNet vocoder, a WaveNet vocoder, a WaveRNN vocoder, a HiFiGAN vocoder, a MelGAN vocoder, etc.
[0080] Please refer to Figure 3 , which is a schematic diagram of the synthesis process of a kind of target audio data in an embodiment of the present application. Taking the target information as text information as an example, specifically, the text information is input into the text analysis module 301, and the first phoneme sequence and text prosody are output; the text prosody and the first phoneme sequence are input into the personalized user acoustic model 302; then, the corrected prosody information is generated by using the text prosody and the pause length reference vector B; the corrected prosody information is merged with the first phoneme sequence to generate a phoneme sequence (the second phoneme sequence) containing prosody information; then, duration prediction is performed according to the second phoneme sequence to predict the pronunciation duration of each phoneme (the first phoneme duration L), then, the speaking speed reference vector A adjusts the first phoneme duration L to obtain the adjusted second phoneme duration L'; then, the second phoneme sequence is expanded according to the adjusted second phoneme duration L' to obtain the first audio feature X after expanding the number of frames; then, the style vector C of the target speaker is superimposed or spliced onto the first audio feature X to output the second audio feature X'; finally, the second audio feature X' is input into the acoustic prediction model for acoustic prediction to output the acoustic feature Y, and the acoustic feature Y is input into the vocoder 303 to output the target audio data.
[0081] In the embodiment of the present application, the target information and the prosody characteristic parameters of the target speaker are obtained. According to the target information and the prosody characteristic parameters of the target speaker, the acoustic feature information is determined, and the acoustic feature information is converted to generate the target audio data corresponding to the target information. In this way, the prosody characteristic parameters of the target speaker can be obtained, and the prosody characteristic parameters of the target speaker can be used to affect the generation of the acoustic feature information. The target audio data can be personalized synthesized according to the speaking styles and prosody characteristics of different speakers, so that the generated target audio data is closer to the pronunciation characteristics of the target speaker. In addition, the audio synthesis method and device provided in this embodiment can be applied to screen reading on electronic devices, the timbre of voice assistants, the timbre of speakers, etc., with wide applicability and good user experience.
[0082] It should be noted that for the audio synthesis method provided in the embodiments of the present application, the execution subject can be an audio synthesis device, or a control module in the audio synthesis device for executing the audio synthesis method. In the embodiments of the present application, the case where the audio synthesis device executes the audio synthesis method is taken as an example to illustrate the audio synthesis device provided in the embodiments of the present application.
[0083] See Figure 4 , the embodiments of the present application also provide an audio synthesis device 400, and the audio synthesis device 400 includes a first acquisition module 401, a second acquisition module 402, a first determination module 403, and a generation module 404.
[0084] The first acquisition module 401 is configured to acquire target information;
[0085] The second acquisition module 402 is configured to acquire prosody characteristic parameters of a target speaker, and the prosody characteristic parameters include a speech rate reference vector, a pause length reference vector, and a style vector;
[0086] The first determination module 403 is configured to determine acoustic feature information according to the target information and the prosody characteristic parameters;
[0087] The generation module 404 is configured to perform conversion on the acoustic feature information to generate target audio data corresponding to the target information.
[0088] Optionally, the first determination module includes: a text analysis unit configured to analyze the target information to obtain text feature parameters, where the text feature parameters include a first phoneme sequence and text prosody; a first generation unit configured to generate a second phoneme sequence according to the text prosody, the pause length reference vector, and the first phoneme sequence; a first determination unit configured to determine a first audio feature according to the second phoneme sequence and the speech rate reference vector; a second determination unit configured to determine a second audio feature according to the first audio feature and the style vector; a third determination unit configured to determine the acoustic feature information based on an acoustic prediction model according to the second audio feature.
[0089] Optionally, the first determination unit is specifically configured to: generate corrected prosody information according to the text prosody and the pause length reference vector; generate a second phoneme sequence according to the corrected prosody information and the first phoneme sequence.
[0090] Optionally, the second determination unit is specifically configured to: perform duration prediction based on the second phoneme sequence to obtain a first phoneme duration; adjust the first phoneme duration according to the speech rate reference vector to obtain a second phoneme duration; expand the second phoneme sequence according to the second phoneme duration to obtain a first audio feature.
[0091] Optionally, the device further includes: a third acquisition module, configured to acquire first audio data of the target speaker, where the first audio data is audio data of the target speaker reading a preset text; and a training module, configured to perform model training based on the first audio data to obtain the acoustic prediction model, where the acoustic prediction model is used to obtain acoustic feature information according to second audio features.
[0092] Optionally, the prosody characteristic parameter includes a speech rate reference vector, and the second acquisition module includes: a first acquisition unit, configured to acquire historical audio data of the target speaker; a fourth determination unit, configured to determine a first average speech rate of the target speaker according to the historical audio data; and a fifth determination unit, configured to determine the speech rate reference vector according to the first average speech rate and a preset average speech rate.
[0093] Optionally, the prosody characteristic parameter includes a pause length reference vector, and the second acquisition module includes: a second acquisition unit, configured to acquire historical audio data of the target speaker; a sixth determination unit, configured to determine a pause probability corresponding to different syllable lengths according to the historical audio data; and a seventh determination unit, configured to determine the pause length reference vector according to the pause probability corresponding to different syllable lengths.
[0094] In the embodiment of the present application, target information and prosody characteristic parameters of a target speaker are acquired, acoustic feature information is determined according to the target information and the prosody characteristic parameters of the target speaker, and the acoustic feature information is converted to generate target audio data corresponding to the target information. In this way, the prosody characteristic parameters of the target speaker can be acquired, and the prosody characteristic parameters of the target speaker can be used to affect the generation of the acoustic feature information. Target audio data can be personalized synthesized according to the speaking styles and prosody characteristics of different speakers, so that the generated target audio data is closer to the pronunciation characteristics of the target speaker.
[0095] The audio synthesis device in the embodiments of this application can be a device, or a component, integrated circuit, or chip in a terminal. The device can be a mobile electronic device or a non-mobile electronic device. Exemplarily, the mobile electronic device can be a mobile phone, a tablet computer, a laptop computer, a handheld computer, a vehicle-mounted electronic device, a wearable device, an ultra-mobile personal computer (UMPC), a netbook, or a personal digital assistant (PDA), etc., and the non-mobile electronic device can be a server, a Network Attached Storage (NAS), a personal computer (PC), a television (TV), a teller machine, or a self-service machine, etc. The embodiments of this application do not make specific limitations.
[0096] The audio synthesis device in the embodiments of this application can be a device with an operating system. The operating system can be the Android operating system, the iOS operating system, or other possible operating systems. The embodiments of this application do not make specific limitations.
[0097] The audio synthesis device provided in the embodiments of this application can implement Figure 1 each process implemented by the method embodiments. To avoid repetition, it will not be elaborated here.
[0098] Optionally, as Figure 5 shown, the embodiments of this application also provide an electronic device 500, including a processor 501, a memory 502, a program or instruction stored on the memory 502 and executable on the processor 501. When the program or instruction is executed by the processor 501, it implements each process of the above-mentioned audio synthesis method embodiment and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.
[0099] It should be noted that the electronic device in the embodiments of this application includes the above-mentioned mobile electronic devices and non-mobile electronic devices.
[0100] Figure 6 It is a schematic diagram of the hardware structure of an electronic device for implementing the embodiments of this application.
[0101] The electronic device 600 includes but is not limited to: a radio frequency unit 601, a network module 602, an audio output unit 603, an input unit 604, a sensor 605, a display unit 606, a user input unit 607, an interface unit 608, a memory 609, and a processor 610, etc.
[0102] Those skilled in the art can understand that the electronic device 600 may further include a power source (such as a battery) for supplying power to each component. The power source can be logically connected to the processor 610 through a power management system, so as to manage functions such as charging, discharging, and power consumption management through the power management system. Figure 6 The structure of the electronic device shown in Figure 6 does not limit the electronic device. The electronic device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements, which will not be elaborated here.
[0103] Among them, the processor 610 is configured to: obtain target information; obtain prosodic characteristic parameters of a target speaker, where the prosodic characteristic parameters include a speech rate reference vector, a pause length reference vector, and a style vector; determine acoustic feature information according to the target information and the prosodic characteristic parameters; and convert the acoustic feature information to generate target audio data corresponding to the target information.
[0104] Optionally, when the processor 610 determines the acoustic feature information according to the target information and the prosodic characteristic parameters, it is configured to: analyze the target information to obtain text feature parameters, where the text feature parameters include a first phoneme sequence and text prosody; generate a second phoneme sequence according to the text prosody, the pause length reference vector, and the first phoneme sequence; determine a first audio feature according to the second phoneme sequence and the speech rate reference vector; determine a second audio feature according to the first audio feature and the style vector; and determine the acoustic feature information according to the second audio feature based on an acoustic prediction model.
[0105] Optionally, when the processor 610 generates the second phoneme sequence according to the text prosody, the pause length reference vector, and the first phoneme sequence, it is configured to: generate corrected prosody information according to the text prosody and the pause length reference vector; and generate the second phoneme sequence according to the corrected prosody information and the first phoneme sequence.
[0106] Optionally, when the processor 610 determines the first audio feature according to the second phoneme sequence and the speech rate reference vector, it is configured to: perform duration prediction based on the second phoneme sequence to obtain a first phoneme duration; adjust the first phoneme duration according to the speech rate reference vector to obtain a second phoneme duration; and expand the second phoneme sequence according to the second phoneme duration to obtain the first audio feature.
[0107] Optionally, before determining the acoustic feature information according to the second audio feature based on the acoustic prediction model, the processor 610 is further configured to: obtain first audio data of the target speaker, where the first audio data is audio data of the target speaker reading a preset text; perform model training based on the first audio data to obtain the acoustic prediction model, where the acoustic prediction model is used to obtain the acoustic feature information according to the second audio feature.
[0108] Optionally, the prosody characteristic parameter includes a speech rate reference vector. When obtaining the prosody characteristic parameter of the target speaker, the processor 610 includes: obtaining historical audio data of the target speaker; determining a first average speech rate of the target speaker according to the historical audio data; and determining the speech rate reference vector according to the first average speech rate and a preset average speech rate.
[0109] Optionally, the prosody characteristic parameter includes a pause length reference vector. When obtaining the prosody characteristic parameter of the target speaker, the processor 710 is configured to: obtain historical audio data of the target speaker; determine pause probabilities corresponding to different syllable lengths according to the historical audio data; and determine the pause length reference vector according to the pause probabilities corresponding to different syllable lengths.
[0110] In the embodiment of the present application, the target information and the prosody characteristic parameter of the target speaker are obtained, the acoustic feature information is determined according to the target information and the prosody characteristic parameter of the target speaker, and the acoustic feature information is converted to generate target audio data corresponding to the target information. In this way, the prosody characteristic parameter of the target speaker can be obtained, and the prosody characteristic parameter of the target speaker can be used to affect the generation of the acoustic feature information. The target audio can be personalized synthesized according to the speaking styles and prosody characteristics of different speakers, so that the generated target audio data is closer to the pronunciation characteristics of the target speaker.
[0111] It should be understood that in the embodiments of the present application, the input unit 604 may include a Graphics Processing Unit (GPU) 6041 and a microphone 6042. The GPU 6041 processes the image data of static pictures or videos obtained by an image capture device (such as a camera) in the video capture mode or the image capture mode. The display unit 606 may include a display panel 6061, and the display panel 6061 may be configured in the form of a liquid crystal display, an organic light-emitting diode, etc. The user input unit 607 includes a touch panel 6071 and other input devices 6072. The touch panel 6071 is also called a touch screen. The touch panel 6071 may include two parts: a touch detection device and a touch controller. The other input devices 7072 may include, but are not limited to, a physical keyboard, function keys (such as volume control keys, power on / off keys, etc.), a trackball, a mouse, and a joystick, which will not be elaborated here. The memory 609 may be used to store software programs and various data, including but not limited to application programs and operating systems. The processor 610 may integrate an application processor and a modem processor. Among them, the application processor mainly processes the operating system, user interface, application programs, etc., and the modem processor mainly processes wireless communication. It can be understood that the above-mentioned modem processor may not be integrated into the processor 610.
[0112] The embodiments of the present application further provide a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, it implements each process of the above-mentioned embodiment of the audio synthesis method and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.
[0113] Among them, the processor is the processor in the electronic device described in the above embodiment. The readable storage medium includes a computer-readable storage medium, such as a computer Read-Only Memory (ROM), a Random Access Memory (RAM), a magnetic disk, or an optical disc, etc.
[0114] The embodiments of the present application further provide a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run a program or instruction to implement each process of the above-mentioned embodiment of the audio synthesis method and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.
[0115] It should be understood that the chip mentioned in the embodiments of the present application may also be referred to as a system-on-chip, a system chip, a chip system, or a system-on-a-chip, etc.
[0116] It should be noted that, in this text, the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or apparatus including a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or apparatus. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article or apparatus including such element. In addition, it should be pointed out that the scope of the methods and apparatuses in the embodiments of the present application is not limited to performing functions in the order shown or discussed, but may also include performing functions in a substantially simultaneous manner or in a reverse order according to the functions involved. For example, the described methods may be performed in an order different from that described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.
[0117] Through the description of the above embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a computer software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions for causing a terminal (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of the present application.
[0118] The embodiments of the present application have been described above with reference to the accompanying drawings. However, the present application is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Those of ordinary skill in the art, under the inspiration of the present application and without departing from the spirit and scope protected by the claims of the present application, can also make many forms, all of which fall within the protection scope of the present application.
Claims
1. An audio synthesis method, characterized in that, the method includes: obtaining target information; using the historical audio data of the target speaker to obtain the prosody characteristic parameters of the target speaker, where the prosody characteristic parameters include a speech rate reference vector, a pause length reference vector, and a style vector; determining acoustic feature information according to the target information and the prosody characteristic parameters; converting the acoustic feature information to generate target audio data corresponding to the target information; wherein, the determining the acoustic feature information according to the target information and the prosody characteristic parameters includes: analyzing the target information to obtain text feature parameters, where the text feature parameters include a first phoneme sequence and text prosody; generating a second phoneme sequence according to the text prosody, the pause length reference vector, and the first phoneme sequence; determining a first audio feature according to the second phoneme sequence and the speech rate reference vector; determining a second audio feature according to the first audio feature and the style vector; determining the acoustic feature information according to the second audio feature based on an acoustic prediction model; wherein, the generating a second phoneme sequence according to the text prosody, the pause length reference vector, and the first phoneme sequence includes: using the prosody probability corresponding to the text prosody as a node probability, using the pause length reference vector as a path probability, and using a dynamic programming algorithm to find an optimal path to obtain corrected prosody information; merging the corrected prosody information with the first phoneme sequence to generate a second phoneme sequence.
2. The method according to claim 1, characterized in that, the determining a first audio feature according to the second phoneme sequence and the speech rate reference vector includes: performing duration prediction based on the second phoneme sequence to obtain a first phoneme duration; adjusting the first phoneme duration according to the speech rate reference vector to obtain a second phoneme duration; expanding the second phoneme sequence according to the second phoneme duration to obtain a first audio feature.
3. The method according to claim 1, characterized in that, before the determining the acoustic feature information according to the second audio feature based on an acoustic prediction model, the method further includes: obtaining first audio data of the target speaker, where the first audio data is audio data of the target speaker reading a preset text; performing model training based on the first audio data to obtain the acoustic prediction model; wherein, the acoustic prediction model is used to obtain acoustic feature information according to the second audio feature.
4. The method according to claim 1, characterized in that, the prosody characteristic parameters include a speech rate reference vector, and the obtaining the prosody characteristic parameters of the target speaker includes: obtaining the historical audio data of the target speaker; determining a first average speech rate of the target speaker according to the historical audio data; determining the speech rate reference vector according to the first average speech rate and a preset average speech rate.
5. The method according to claim 1, characterized in that, The prosody characteristic parameters include a pause length reference vector, and obtaining the prosody characteristic parameters of the target speaker includes: Obtaining the historical audio data of the target speaker; Determining the pause probabilities corresponding to different syllable lengths according to the historical audio data; Determining a pause length reference vector according to the pause probabilities corresponding to different syllable lengths.
6. An audio synthesis device, characterized in that, the device includes: A first acquisition module for acquiring target information; A second acquisition module for using the historical audio data of the target speaker to acquire the prosody characteristic parameters of the target speaker, where the prosody characteristic parameters include a speech rate reference vector, a pause length reference vector, and a style vector; A first determination module for determining acoustic feature information according to the target information and the prosody characteristic parameters; A generation module for converting the acoustic feature information to generate target audio data corresponding to the target information; wherein, the first determination module includes: A text analysis unit for analyzing the target information to obtain text feature parameters, where the text feature parameters include a first phoneme sequence and text prosody; A first generation unit for generating a second phoneme sequence according to the text prosody, the pause length reference vector, and the first phoneme sequence; A first determination unit for determining a first audio feature according to the second phoneme sequence and the speech rate reference vector; A second determination unit for determining a second audio feature according to the first audio feature and the style vector; A third determination unit for determining the acoustic feature information according to the second audio feature based on an acoustic prediction model; wherein, the first generation unit is specifically used for: Taking the prosody probability corresponding to the text prosody as the node probability, taking the pause length reference vector as the path probability, and using a dynamic programming algorithm to find the optimal path to obtain the corrected prosody information; Merging the corrected prosody information with the first phoneme sequence to generate a second phoneme sequence.
7. The device according to claim 6, characterized in that, the second determination unit is specifically used for: Performing duration prediction based on the second phoneme sequence to obtain a first phoneme duration; Adjusting the first phoneme duration according to the speech rate reference vector to obtain a second phoneme duration; Expanding the second phoneme sequence according to the second phoneme duration to obtain a first audio feature.
8. The device according to claim 6, characterized in that, the device further includes: A third acquisition module for acquiring the first audio data of the target speaker, where the first audio data is the audio data of the target speaker reading a preset text; A training module for performing model training based on the first audio data to obtain the acoustic prediction model, where the acoustic prediction model is used to obtain acoustic feature information according to the second audio feature.
9. The device according to claim 6, characterized in that, the prosody characteristic parameters include a speech rate reference vector, and the second acquisition module includes: A first acquisition unit for acquiring the historical audio data of the target speaker; A fourth determination unit, configured to determine a first average speech rate of the target speaker according to the historical audio data; A fifth determination unit, configured to determine the speech rate reference vector according to the first average speech rate and a preset average speech rate.
10. The apparatus according to claim 6, wherein, the prosody feature parameter includes a pause length reference vector, and the second acquisition module includes: A second acquisition unit, configured to acquire the historical audio data of the target speaker; A sixth determination unit, configured to determine a pause probability corresponding to different syllable lengths according to the historical audio data; A seventh determination unit, configured to determine the pause length reference vector according to the pause probability corresponding to different syllable lengths.
11. An electronic device, wherein, it includes a processor, a memory, and a program or instruction stored on the memory and executable on the processor, and when the program or instruction is executed by the processor, the steps of the audio synthesis method according to any one of claims 1 to 5 are implemented.
12. A readable storage medium, wherein, a program or instruction is stored on the readable storage medium, and when the program or instruction is executed by a processor, the steps of the audio synthesis method according to any one of claims 1 to 5 are implemented.
Citation Information
Patent Citations
Speech synthesis method and related equipment
CN108962217A
Speech synthesis method and device, equipment and storage medium
CN112786009A