Speech synthesis and broadcast method, teaching method, live broadcast method and device
Through the switching mechanism of online and local voice synthesis services, the problem of high voice synthesis cost on end-side devices with limited resources is solved, and a natural and smooth voice synthesis experience is achieved.
Patent Information
- Application Number
- CN202110352807.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-03-31
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2041-03-31
AI Technical Summary
How to reduce the cost of speech synthesis while ensuring the effect of speech synthesis, especially to achieve natural and smooth speech synthesis on end-side devices with limited resources.
The switching mechanism between online voice synthesis services and local voice synthesis services is adopted. When the online voice synthesis service is unavailable, the local voice synthesis service is used to generate audio data similar to that of online voice synthesis services, ensuring the continuity and nature of voice synthesis.
When the online voice synthesis service is unavailable, audio data with similar tones is generated through the local voice synthesis service, providing a seamless voice synthesis experience and reducing the cost of voice synthesis.
Smart Images

Figure CN115148184B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of data processing, and particularly to a method and apparatus for speech synthesis and broadcasting, a teaching method, and a live broadcast method. Background Art
[0002] Text-to-speech (TTS) technology is a technology that converts text into voice output. With the development of the AI wave, the application scenarios of speech synthesis are becoming more and more extensive, such as smart speakers, virtual assistants, audiobooks, etc.
[0003] Limited by the limited resources of end-side devices (including mobile phones and other embedded devices, etc.), when applying speech synthesis to end-side devices, how to minimize the speech synthesis cost while ensuring the speech synthesis effect is an urgent problem to be solved currently.
[0004] That is, a speech synthesis scheme is needed that can minimize the speech synthesis cost while ensuring the speech synthesis effect. Summary of the Invention
[0005] One technical problem to be solved by the present disclosure is to provide a speech synthesis scheme that can minimize the speech synthesis cost while ensuring the speech synthesis effect.
[0006] According to the first aspect of the present disclosure, a speech synthesis method is provided, including: synthesizing first audio data for a first text based on an online speech synthesis service; in response to the unavailability of the online speech synthesis service, synthesizing second audio data for a second text based on a local speech synthesis service, and the similarity between the timbre of the second audio data and the timbre of at least part of the first audio data is greater than or equal to a first threshold.
[0007] According to the second aspect of the present disclosure, a speech synthesis method is provided, including: synthesizing third audio data for a third text based on a local speech synthesis service; in response to the availability of the online speech synthesis service, synthesizing fourth audio data for a fourth text based on the online speech synthesis service, and the similarity between the timbre of the fourth audio data and the timbre of at least part of the third audio data is greater than or equal to a first threshold.
[0008] According to the third aspect of the present disclosure, a speech broadcast method is provided, including: obtaining first audio data synthesized for a first text based on an online speech synthesis service; in response to the unavailability of the online speech synthesis service, obtaining second audio data synthesized for a second text based on a local speech synthesis service, and the similarity between the timbre of the second audio data and the timbre of at least part of the first audio data is greater than or equal to a first threshold; and broadcasting the first audio data and / or the second audio data.
[0009] According to a fourth aspect of the present disclosure, there is provided a voice broadcast method, including: obtaining third audio data synthesized for a third text based on a local voice synthesis service; in response to the availability of an online voice synthesis service, obtaining fourth audio data synthesized for a fourth text based on the online voice synthesis service, where the similarity between the timbre of the fourth audio data and the timbre of at least part of the third audio data is greater than or equal to a first threshold; and broadcasting the third audio data and / or the fourth audio data.
[0010] According to a fifth aspect of the present disclosure, there is provided a teaching method, including: determining an answer corresponding to a question raised by a student; obtaining first audio data synthesized for a first part of the answer based on an online voice synthesis service; in response to the unavailability of the online voice synthesis service, obtaining second audio data synthesized for a second part of the answer based on a local voice synthesis service, where the similarity between the timbre of the second audio data and the timbre of at least part of the first audio data is greater than or equal to a first threshold; and broadcasting the first audio data and / or the second audio data.
[0011] According to a sixth aspect of the present disclosure, there is provided a live broadcast method, including: obtaining first audio data synthesized for a first part of the content to be broadcast based on an online voice synthesis service; in response to the unavailability of the online voice synthesis service, obtaining second audio data synthesized for a second part of the content to be broadcast based on a local voice synthesis service, where the similarity between the timbre of the second audio data and the timbre of at least part of the first audio data is greater than or equal to a first threshold; and broadcasting the first audio data and / or the second audio data during the presentation of the live broadcast screen.
[0012] According to a seventh aspect of the present disclosure, there is provided a voice synthesis device, including: a first synthesis module for synthesizing first audio data for a first text based on an online voice synthesis service; and a second synthesis module for, in response to the unavailability of the online voice synthesis service, synthesizing second audio data for a second text based on a local voice synthesis service, where the similarity between the timbre of the second audio data and the timbre of at least part of the first audio data is greater than or equal to a first threshold.
[0013] According to an eighth aspect of the present disclosure, there is provided a voice synthesis device, including: a first synthesis module for synthesizing third audio data for a third text based on a local voice synthesis service; a second synthesis module for, in response to the availability of the online voice synthesis service, synthesizing fourth audio data for a fourth text based on the online voice synthesis service, where the similarity between the timbre of the fourth audio data and the timbre of at least part of the third audio data is greater than or equal to a first threshold.
[0014] According to a ninth aspect of the present disclosure, there is provided a voice broadcast device, including: a first acquisition module configured to acquire first audio data synthesized for a first text based on an online voice synthesis service; a second acquisition module configured to, in response to the unavailability of the online voice synthesis service, acquire second audio data synthesized for a second text based on a local voice synthesis service, the similarity between the timbre of the second audio data and the timbre of at least a part of the first audio data being greater than or equal to a first threshold; and a broadcast module configured to broadcast the first audio data and / or the second audio data.
[0015] According to a tenth aspect of the present disclosure, there is provided a voice broadcast device, including: a first acquisition module configured to acquire third audio data synthesized for a third text based on a local voice synthesis service; a second acquisition module configured to, in response to the availability of the online voice synthesis service, acquire fourth audio data synthesized for a fourth text based on the online voice synthesis service, the similarity between the timbre of the fourth audio data and the timbre of at least a part of the third audio data being greater than or equal to a first threshold; and a broadcast module configured to broadcast the third audio data and / or the fourth audio data.
[0016] According to an eleventh aspect of the present disclosure, there is provided a teaching device, including: a determination module configured to determine an answer corresponding to a question raised by a student; a first acquisition module configured to acquire first audio data synthesized for a first part of the answer based on an online voice synthesis service; a second acquisition module configured to, in response to the unavailability of the online voice synthesis service, acquire second audio data synthesized for a second part of the answer based on a local voice synthesis service, the similarity between the timbre of the second audio data and the timbre of at least a part of the first audio data being greater than or equal to a first threshold; and a broadcast module configured to broadcast the first audio data and / or the second audio data.
[0017] According to a twelfth aspect of the present disclosure, there is provided a live broadcast device, including: a first acquisition module configured to acquire first audio data synthesized for a first part of the content to be broadcast based on an online voice synthesis service; a first acquisition module configured to, in response to the unavailability of the online voice synthesis service, acquire second audio data synthesized for a second part of the content to be broadcast based on a local voice synthesis service, the similarity between the timbre of the second audio data and the timbre of at least a part of the first audio data being greater than or equal to a first threshold; and a broadcast module configured to broadcast the first audio data and / or the second audio data during the presentation of a live broadcast screen.
[0018] According to a thirteenth aspect of the present disclosure, there is provided a computing device, including: a processor; and a memory storing executable code thereon, which, when executed by the processor, causes the processor to execute the method according to any one of the first aspect to the sixth aspect as described above.
[0019] According to the fourteenth aspect of the present disclosure, there is provided a non-transitory machine-readable storage medium having executable code stored thereon, which when executed by a processor of an electronic device, causes the processor to execute the method described in any one of the first aspect to the sixth aspect as described above.
[0020] Thus, when switching from an online speech synthesis service to a local speech synthesis service and performing speech synthesis based on the local speech synthesis service, by synthesizing second audio data whose timbre has a similarity greater than or equal to a first threshold with the timbre of at least part of the first audio data synthesized based on the online speech synthesis service, a user can obtain a natural and smooth speech synthesis experience auditorily. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] By describing the exemplary embodiments of the present disclosure in more detail in conjunction with the accompanying drawings, the above and other objects, features, and advantages of the present disclosure will become more apparent. Among them, in the exemplary embodiments of the present disclosure, the same reference numerals generally represent the same components.
[0022] Figure 1 FIG. shows a schematic flowchart of a speech synthesis method according to an embodiment of the present disclosure.
[0023] Figure 2 FIG. shows a schematic flowchart of a voice broadcast method according to an embodiment of the present disclosure.
[0024] Figure 3 FIG. shows a flowchart of an application scenario according to an embodiment of the present disclosure.
[0025] Figure 4 FIG. shows a schematic structural diagram of a voice interaction device according to an embodiment of the present disclosure.
[0026] Figure 5 FIG. shows a schematic structural diagram of a voice broadcast device according to an embodiment of the present disclosure.
[0027] Figure 6 FIG. shows a schematic structural diagram of a teaching device according to an embodiment of the present disclosure.
[0028] Figure 7 FIG. shows a schematic structural diagram of a live broadcast device according to an embodiment of the present disclosure.
[0029] Figure 8 FIG. shows a schematic structural diagram of a computing device according to an embodiment of the present disclosure. DETAILED DESCRIPTION
[0030] The preferred embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although the preferred embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided to make the present disclosure more thorough and complete, and to fully convey the scope of the present disclosure to those skilled in the art.
[0031] Limited by the limited resources of end-side devices (including mobile phones and other embedded devices), there is still a large gap between the effects of end-side speech synthesis (i.e., local speech synthesis) and cloud speech synthesis (i.e., online speech synthesis) in terms of cleanliness, prosody, etc.
[0032] In order to better attract users, end-side devices or end-side applications usually choose the online speech synthesis method with better sound quality. An end-side application is an application program installed and running on an end-side device.
[0033] Although the cost of online speech synthesis has decreased significantly with the iterative upgrade of speech synthesis algorithms, for scenarios such as news broadcasting, audiobook listening, teaching, and live streaming, due to the very large amount of speech synthesis data, high costs still need to be borne. At the same time, the innovation of mobile communication technology has not completely changed the current situation that the real-time performance and success rate of online speech synthesis are worse than those of local speech synthesis.
[0034] The present disclosure focuses on how to provide a natural and fluent synthesis effect for end-side devices or end-side applications to improve the user experience.
[0035] Figure 1 Fig. shows a schematic flowchart of a speech synthesis method according to an embodiment of the present disclosure.
[0036] Figure 1 The method shown can be implemented in software by a computer program and can also be executed by a computing device with a specific configuration. Figure 1 The method shown.
[0037] As an example, Figure 1 The method shown can be executed by a speech synthesis system, which can be installed in an end-side device in the form of an application program or a plug-in to provide speech synthesis services for the end-side device or end-side application.
[0038] See Figure 1 , in step S110, first audio data is synthesized for the first text based on the online speech synthesis service. The first text is the text for which speech needs to be synthesized.
[0039] The online speech synthesis service is used to perform speech synthesis for the first text using the online speech synthesis method, and the online speech synthesis service can be provided by an online speech synthesis system located in the cloud.
[0040] The online speech synthesis method can be, but is not limited to, a speech synthesis method based on waveform splicing, a speech synthesis method based on parameters, or a speech synthesis method based on deep learning.
[0041] The speech synthesis method based on waveform splicing means that for the text of the speech to be synthesized, the audio data corresponding to the words and phrases included in the text are selected from the speech database, and then these audio data are spliced in the order of the words and phrases in the text to obtain the speech synthesis result of the text.
[0042] The speech synthesis method based on parameters means that for the text of the speech to be synthesized, the speech data of the text are reconstructed by using acoustic parameters and a vocoder. For example, an HMM (Hidden Markov Model) can be pre-trained based on the speech database, and the HMM predicts the acoustic parameters and duration for the text, and then the vocoder is used to complete the speech synthesis.
[0043] The speech synthesis method based on deep learning means that a speech synthesis model for outputting the speech synthesis result for the text can be pre-trained by using deep learning technology, and the speech synthesis model is used to generate the speech data corresponding to the text to be synthesized.
[0044] In step S120, in response to the unavailability of the online speech synthesis service, the second audio data is synthesized for the second text based on the local speech synthesis service.
[0045] In the process of synthesizing audio data based on the online speech synthesis service, communication with the cloud providing the online speech synthesis service is required to obtain the audio data synthesized based on the online speech synthesis service. When the online speech synthesis service is unavailable due to poor communication quality, network disconnection, etc., the second audio data can be synthesized for the second text based on the local speech synthesis service.
[0046] The second text is the text for which speech needs to be synthesized. The second text and the first text can form the complete text of the speech to be synthesized. That is, the second text can be the part of the text of the speech to be synthesized that has not been synthesized yet from the start of speech synthesis based on the online speech synthesis service until the online speech synthesis service becomes unavailable.
[0047] The local speech synthesis service is used to perform speech synthesis using the local speech synthesis method. The local speech synthesis service can be provided by a local speech synthesis system located at the end side. The local speech synthesis method can be, but is not limited to, a speech synthesis method based on waveform splicing, a speech synthesis method based on parameters, or a speech synthesis method based on deep learning.
[0048] Generally speaking, the online speech synthesis method can be a waveform splicing-based speech synthesis method with better speech synthesis effect. The local speech synthesis method can be a parameter-based speech synthesis method or a deep learning-based speech synthesis method that has low requirements for end-side resources.
[0049] Of course, the online speech synthesis service and the local speech synthesis service can also adopt speech synthesis methods based on the same speech synthesis principle, that is, the online speech synthesis method and the local speech synthesis method can also be speech synthesis methods based on the same speech synthesis principle.
[0050] Considering the limited resources of the end-side device, when the local speech synthesis service and the online speech synthesis service adopt the same principle of speech synthesis method, the speech synthesis method adopted by the local speech synthesis service can be a simplified speech synthesis method obtained by reducing the requirements for speech synthesis quality, and this speech synthesis method can reduce the requirements for device resources when used.
[0051] Taking the example that both the local speech synthesis service and the online speech synthesis service adopt the deep learning-based speech synthesis method, compared with the online speech synthesis service, the local speech synthesis service can use a speech synthesis model with a simpler structure but relatively poorer speech synthesis effect for speech synthesis.
[0052] The local speech synthesis service can provide speech synthesis services with multiple timbres, that is, the local speech synthesis service has the function of speech synthesis with multiple timbres and can be used to synthesize speech with multiple timbres.
[0053] In order to obtain a natural and smooth speech synthesis effect, when synthesizing the second audio data for the second text based on the local speech synthesis service, the second audio data with a timbre similarity greater than or equal to the first threshold to the timbre of at least part of the first audio data can be synthesized according to the timbre of at least part of the first audio data. Among them, the first threshold can be a relatively high value set in advance, such as 90%. By setting the first threshold to a relatively high value, when the timbre similarity of the second audio data to the timbre of at least part of the first audio data is greater than or equal to the first threshold, it can be basically considered that the timbre of the second audio data is the same as or similar to the timbre of at least part of the first audio data.
[0054] Thus, when switching from the online speech synthesis service to the local speech synthesis service, due to the timbre of the second audio data being the same as or similar to the timbre of at least part of the first audio data, the user can obtain a seamless speech synthesis experience in terms of audition.
[0055] As an example, the local speech synthesis service can provide multiple acoustic models. When synthesizing the second audio data for the second text based on the local speech synthesis service, an acoustic model whose synthesized voice timbre is the same as or similar to at least part of the first audio data can be selected from the multiple acoustic models, and the selected acoustic model can be used to perform speech synthesis on the second text to obtain the second audio data.
[0056] The acoustic model can be a speaker model trained using a machine learning algorithm. Each acoustic model corresponds to a timbre. Specifically, the acoustic model can be a model for predicting acoustic parameters (such as an HMM model), or a model for outputting a speech synthesis result for a text (such as a speech synthesis model trained based on a deep learning algorithm). When the acoustic model is used to predict acoustic parameters, for the text of the speech to be synthesized, the acoustic model can be used to predict the acoustic parameters of the text, and then a vocoder can be used to synthesize speech data based on the acoustic parameters. When the acoustic model is a speech synthesis model, for the text of the speech to be synthesized, the speech synthesis model can be directly used to obtain the speech synthesis result of the text, that is, speech data.
[0057] When selecting a suitable acoustic model from multiple acoustic models, third audio data synthesized for the text corresponding to at least part of the first audio data using different acoustic models can be obtained, the similarity between each third audio data and the at least part of the first audio data can be calculated, and the acoustic model corresponding to the third audio data with a higher similarity ranking (such as the highest similarity) or a similarity greater than or equal to a second threshold can be selected. Among them, the calculation method of similarity can adopt but is not limited to the Dynamic Time Warping (DTW) algorithm. The DTW algorithm is suitable for measuring the similarity of sequence signals that do not match in time or speed. The specific calculation principle of the DTW algorithm can refer to the prior art. The second threshold can be the same as or different from the first threshold. For example, the second threshold can be greater than or equal to the first threshold.
[0058] When synthesizing the second audio data for the second text based on the local speech synthesis service, the acoustic model used for speech synthesis can also be set or replaced based on the user's interaction instructions. For example, in response to the user's first interaction instruction, an acoustic model for performing speech synthesis on the second text can be selected from multiple acoustic models; and / or, in response to the user's second interaction instruction, the selected acoustic model can be replaced. The first interaction instruction and the second interaction instruction can be made by the user through speech, text input, touch, etc.
[0059] The online speech synthesis service can also provide speech synthesis services with multiple timbres, that is, the online speech synthesis service can also have the function of speech synthesis with multiple timbres for synthesizing speech with multiple timbres.
[0060] The speech synthesis functions with multiple voices provided by the online speech synthesis service can be regarded as multiple online speakers, and the speech synthesis functions with multiple voices provided by the local speech synthesis service can be regarded as multiple local speakers. The present disclosure can pre - establish a correspondence relationship between the online speaker and the local speaker with consistent voices, so that when the online speech synthesis service is unavailable, it can be directly switched to the local speaker with the highest similarity.
[0061] Taking the online speech synthesis service adopting the waveform - splicing - based speech synthesis method and the local speech synthesis service using an acoustic model to synthesize speech as an example, the online speech synthesis service can synthesize speech based on a speech database. The speech database includes one or more audio material sets, and each audio material set corresponds to a speaker (corresponding to the online speaker mentioned above). Each audio material set can include the pronunciation data of a certain real speaker or virtual speaker for multiple pieces of corpus.
[0062] The local speech synthesis service provides multiple acoustic models, and the local speech synthesis service can select an acoustic model from multiple acoustic models to synthesize speech. The present disclosure can pre - determine the acoustic model corresponding to the audio material set. The timbre of the audio data synthesized based on the acoustic model corresponding to the audio material set is the same as or basically the same as that of the audio material set.
[0063] The audio material set, that is, the speech package. The speech database used by the online speech synthesis service can also be called the cloud speech database, and the speech package (i.e., the audio material set) in the cloud speech database can also be called the cloud speech package. The acoustic model used by the local speech synthesis service can also be called the local speech package. Establishing the correspondence relationship between the audio material set and the acoustic model, that is, establishing the correspondence relationship between the local speech package and the cloud speech package, that is, the mapping relationship between the local speech package and the cloud speech package.
[0064] Thus, in response to the unavailability of the online speech synthesis service, the acoustic model corresponding to the audio material set used when synthesizing the first audio data can be directly used to synthesize the second audio data for the second text. The process of synthesizing speech data using the acoustic model can refer to the relevant description above and will not be elaborated here. Among them, the audio material set used when synthesizing the first audio data can be determined by the online speech synthesis service based on the text to be synthesized or specified by the user.
[0065] Optionally, both the cloud speech synthesis service and the local speech synthesis service can adopt the waveform - splicing - based speech synthesis method. At this time, the speech database used by the local speech synthesis service can be called the local speech database. The speech package (i.e., the audio material set) in the local speech database can also be called the local speech package.
[0066] The present disclosure can also save the synthesized first audio data and / or second audio data so that when the same text needs to be synthesized by voice in the future, the saved audio data can be directly used.
[0067] In response to the online voice synthesis service transitioning from an unavailable state to an available state, voice synthesis can be re-performed based on the online voice synthesis service. Thus, the local voice synthesis service can be used as an alternative when the online voice synthesis service is unavailable. And based on the voice synthesis solution of the present disclosure, when applying the alternative, voice data with the same or similar timbre as the voice synthesized by the online voice synthesis service can be obtained, so that when switching between the local voice synthesis service and the online voice synthesis service, users can obtain a seamless voice synthesis experience sensually.
[0068] The present disclosure also proposes a voice synthesis method, including: synthesizing third audio data for a third text based on a local voice synthesis service; in response to the online voice synthesis service being available, synthesizing fourth audio data for a fourth text based on the online voice synthesis service, and the similarity between the timbre of the fourth audio data and the timbre of at least part of the third audio data is greater than or equal to a first threshold. The first threshold can be a relatively high value set in advance, such as 90%. By setting the first threshold to a relatively high value, when the similarity between the timbre of the fourth audio data and the timbre of at least part of the third audio data is greater than or equal to the first threshold, it can be basically considered that the timbre of the fourth audio data is the same or similar to the timbre of at least part of the third audio data.
[0069] In this embodiment, if the online voice synthesis service is initially in an unavailable state, the local voice synthesis service can be used to synthesize third audio data for the third text based on the local voice synthesis service. In response to the online voice synthesis service changing from an unavailable state to an available state, fourth audio data can be synthesized for the fourth text based on the online voice synthesis service.
[0070] Both the third text and the fourth text are texts for which voice synthesis is to be performed. The third text and the fourth text can form a complete text for which voice synthesis is to be performed. That is, the third text can be the part of the text for which voice synthesis has not been performed yet in the text for which voice synthesis is to be performed from the start of voice synthesis based on the local voice synthesis service to the use of the online voice synthesis service for voice synthesis.
[0071] Regarding the local voice synthesis service and the online voice synthesis service, reference can be made to the relevant descriptions above.
[0072] The online voice synthesis service can provide voice synthesis services with multiple timbres, that is, the online voice synthesis service has the function of voice synthesis with multiple timbres and can be used to synthesize voices with multiple timbres.
[0073] To obtain a natural and fluent speech synthesis effect, when synthesizing the fourth audio data for the fourth text based on an online speech synthesis service, the fourth audio data with a timbre similarity greater than or equal to a first threshold (i.e., the same or similar) to the timbre of at least part of the third audio data can be synthesized according to the timbre of at least part of the third audio data.
[0074] Thus, when switching from a local speech synthesis service to an online speech synthesis service, since the timbre of the fourth audio data is the same or similar to the timbre of at least part of the third audio data, the user can obtain a seamless speech synthesis experience auditorily.
[0075] As an example, an online speech synthesis service can synthesize speech based on a speech library, which includes one or more audio material sets, and each audio material set corresponds to a speaker (corresponding to the online speaker mentioned above). Each audio material set may include pronunciation data of a certain real or virtual speaker for multiple pieces of corpus.
[0076] When synthesizing the fourth audio data for the fourth text based on an online speech synthesis service, an audio material set with a timbre similarity greater than or equal to the first threshold (i.e., the same or similar) to at least part of the third audio data can be selected from multiple audio material sets, and the selected audio material set can be used to perform speech synthesis on the fourth text to obtain the fourth audio data.
[0077] Figure 2 FIG. shows a schematic flowchart of a speech broadcast method according to an embodiment of the present disclosure. Figure 2 The method shown can be implemented in software by a computer program and can also be executed by a computing device with a specific configuration Figure 2 for the method shown. As an example, Figure 2 the method shown can be executed by an end-side device or an end-side application.
[0078] See Figure 2 , in step S210, obtain the first audio data synthesized for the first text based on an online speech synthesis service.
[0079] Regarding the implementation of the online speech synthesis service and synthesizing the first audio data for the first text based on the online speech synthesis service, reference can be made to the relevant description above in conjunction with Figure 1 and will not be elaborated here.
[0080] In step S220, in response to the unavailability of the online speech synthesis service, obtain the second audio data synthesized for the second text based on a local speech synthesis service, and the timbre of the second audio data is greater than or equal to the first threshold in similarity to the timbre of at least part of the first audio data.
[0081] For the local speech synthesis service, the implementation method of synthesizing the second audio data for the second text based on the local speech synthesis service can refer to the relevant description in the above text in combination with Figure 1 and will not be elaborated here.
[0082] In step S230, the first audio data and / or the second audio data is / are broadcast.
[0083] The present disclosure does not limit the execution order of step S230. That is, after obtaining the first audio data and the second audio data, the first audio data and the second audio data can be broadcast. It is also possible to first broadcast the first audio data after obtaining the first audio data, and then broadcast the second audio data after obtaining the second audio data subsequently. Optionally, the broadcast of the audio data can be performed in real time, that is, in response to obtaining the audio data synthesized based on the speech synthesis service, the audio data can be broadcast in real time.
[0084] For the user, when switching from the online speech synthesis service to the local speech synthesis service, since the timbre of the second audio data is the same as or similar to at least part of the timbre of the first audio data, when the first audio data and the second audio data are broadcast, the user can obtain a seamless speech synthesis experience aurally.
[0085] The present disclosure can also save the first audio data and / or the second audio data. For example, the first audio data and / or the second audio data can be saved locally in association with the corresponding text.
[0086] Thus, when it is necessary to synthesize speech data for a text, it can first be determined whether there is audio data corresponding to the text locally. If so, the audio data can be directly broadcast without obtaining the audio data through the online speech synthesis service or the local speech synthesis service.
[0087] As an example, the text for which speech is to be synthesized can be the content provided by a client application (i.e., the end-side application), and the text for which speech is to be synthesized can include the above-mentioned first text and second text.
[0088] For the text for which speech is to be synthesized, it can first be determined whether there is audio data corresponding to the text for which speech is to be synthesized on the server corresponding to the client application. In the case where it is determined that there is audio data corresponding to the text for which speech is to be synthesized on the server, the audio data is obtained from the server and broadcast, without obtaining the audio data through the online speech synthesis service or the local speech synthesis service, thereby reducing the number of requests for the speech synthesis service and further reducing the speech synthesis cost. Among them, the server corresponding to the client application can save the speech synthesis data of some common texts or the speech synthesis data of texts frequently used by the user.
[0089] In the case where it is determined that the server does not have audio data corresponding to the text of the speech to be synthesized, it is possible to continue to determine whether there is audio data corresponding to the text of the speech to be synthesized locally. In the case where it is determined that there is also no audio data corresponding to the text of the speech to be synthesized locally, step S210 is then executed. Thus, through the "local cache + server cache" two-level cache mechanism, the number of requests to the third-party speech synthesis service can be greatly reduced, and the speech synthesis cost can be reduced.
[0090] In this embodiment, the local speech synthesis service or the online speech synthesis service can be an in-built function of the client application, that is, the client application itself can provide the local speech synthesis service or the online speech synthesis service. Alternatively, the local speech synthesis service and / or the online speech synthesis service can also be used as a third-party speech synthesis service to provide a speech synthesis service for the client application.
[0091] The present disclosure can also obtain the content to be broadcast, parse the content to be broadcast, and determine the text of the speech to be synthesized and / or the parameter information related to speech broadcast existing in the content to be broadcast. When executing step S230, the first audio data and / or the second audio data can be broadcast based on the parameter information. Among them, the parameter information can include but is not limited to the broadcast speed, broadcast volume, etc. The content to be broadcast can be specified by the user. For example, the content to be broadcast can be determined according to operations such as selection and / or input performed by the user on the client application.
[0092] The present disclosure also proposes a speech broadcast method, including: obtaining third audio data synthesized based on a local speech synthesis service for a third text; in response to the availability of an online speech synthesis service, obtaining fourth audio data synthesized based on the online speech synthesis service for a fourth text, where the similarity between the timbre of the fourth audio data and the timbre of at least part of the third audio data is greater than or equal to a first threshold; and broadcasting the third audio data and / or the fourth audio data.
[0093] In this embodiment, if the online speech synthesis service is initially in an unavailable state, it is possible to first obtain the third audio data synthesized based on the local speech synthesis service for the third text. In response to the online speech synthesis service changing from an unavailable state to an available state, the fourth audio data synthesized based on the online speech synthesis service for the fourth text is then obtained.
[0094] For the acquisition processes of the local speech synthesis service, the online speech synthesis service, the third audio data, and the fourth audio data, reference can be made to the relevant descriptions above, and details are not repeated here.
[0095] After obtaining the third audio data and the fourth audio data, the third audio data and the fourth audio data can be broadcast. It is also possible to first broadcast the third audio data after obtaining the third audio data, and then broadcast the fourth audio data after obtaining the fourth audio data subsequently.
[0096] For the user, when switching from the local speech synthesis service to the online speech synthesis service, since the similarity between the timbre of the fourth audio data and the timbre of at least part of the third audio data is greater than or equal to the first threshold, that is, the timbres of the two are the same or similar, when the third audio data and the fourth audio data are broadcast, the user can obtain a seamless speech synthesis experience auditorily.
[0097] The present disclosure can also save the third audio data and / or the fourth audio data. For example, the third audio data and / or the fourth audio data can be saved locally in association with the corresponding text.
[0098] Thus, when it is necessary to synthesize speech data for a text, it can first be determined whether there is audio data corresponding to the text locally. If so, the audio data can be directly broadcast without obtaining the audio data through the online speech synthesis service or the local speech synthesis service again.
[0099] Figure 3 Shows a flowchart of an application scenario according to an embodiment of the present disclosure.
[0100] As Figure 3 shown, the user can open the application running on the end-side device and select the content to be broadcast. The application can be an application that requires speech broadcast capabilities.
[0101] In response to receiving a broadcast instruction from the user for the content to be broadcast, the application can first determine whether the network is normal. If the network is normal, the content to be broadcast can be filtered. For example, the content or the content number can be sent back to the server to check whether the server has cached the speech data corresponding to the content. If it has been cached, then the application can directly access the server to obtain the corresponding speech data and broadcast it.
[0102] If the network is not normal, or the server does not cache the speech data corresponding to the content. Then the content to be broadcast can be handed over to the speech synthesis system. Among them, the speech synthesis system can be a third-party speech synthesis system that exists independently of the application, or the speech synthesis system can also be provided by the application, such as a plugin embedded in the application for providing speech synthesis services.
[0103] The speech synthesis system can process the text sent into the speech synthesis system. The specific processing can include determining whether the text contains special instructions (such as adjusting the volume, speech rate, etc.) and verifying whether there is a cached audio corresponding to the text locally.
[0104] If the audio corresponding to the text is cached locally, the local audio can be read and handed over to the application for playback.
[0105] If there is no cached audio corresponding to the text locally, it can be determined whether the network is available, that is, whether there is a network disconnection. If the network is available, an online speech synthesis service can be requested to synthesize speech for the text to be broadcast. If the online speech synthesis service becomes unavailable during the process (i.e., the request fails), the local speech synthesis service can be switched to, and speech can be synthesized based on the local speech synthesis service. Among them, when switching to the local speech synthesis service, a fallback pronunciation selection operation can be performed to select a speaker with a pronunciation similar to the voice color synthesized by the cloud for local speech synthesis.
[0106] For the audio data synthesized based on the online speech synthesis service and the local speech synthesis service, they can all be cached locally so that the local file can be directly read during the next broadcast. Optionally, before caching, it can be determined whether the local caching function is enabled. If it is enabled, the synthesized speech is cached.
[0107] Optionally, to reduce the transmission cost, the audio stream obtained from the server or the audio stream obtained based on the online speech synthesis service is usually a compressed format (such as mp3) file. Therefore, it is necessary to parse the received audio stream to convert it into a format that the player can recognize (such as PCM, Pulse Code Modulation). The parsed audio stream can be sent to the player for playback.
[0108] This disclosure can be applied to, but is not limited to, application scenarios such as news broadcast, audiobook listening, teaching, live broadcast, etc.
[0109] Taking the application in the teaching scenario as an example, this disclosure can also be implemented as a teaching method, which can be executed by devices used by students (such as smartphones, smart bracelets, smart watches, etc.).
[0110] The teaching method includes the following steps: determining an answer corresponding to the question raised by the student; obtaining first audio data synthesized based on the online speech synthesis service for the first part of the answer; in response to the unavailability of the online speech synthesis service, obtaining second audio data synthesized based on the local speech synthesis service for the second part of the answer, where the similarity between the voice color of the second audio data and at least part of the voice color of the first audio data is greater than or equal to a first threshold; and broadcasting the first audio data and / or the second audio data.
[0111] The first part of the answer and the second part of the answer can form a complete answer corresponding to the question. After obtaining the first audio data and the second audio data corresponding to the complete answer, the first audio data and the second audio data can be broadcast; alternatively, the obtained first audio data and second audio data can be broadcast in real time.
[0112] Taking the application to a live broadcast scenario as an example, the present disclosure can also be implemented as a live broadcast method, which can be executed by a device capable of video live broadcast (such as a smart phone, a computer).
[0113] The live broadcast method includes the following steps: obtaining first audio data synthesized based on an online speech synthesis service for a first part of the content to be broadcast; in response to the unavailability of the online speech synthesis service, obtaining second audio data synthesized based on a local speech synthesis service for a second part of the content to be broadcast, where the similarity between the timbre of the second audio data and the timbre of at least part of the first audio data is greater than or equal to a first threshold; and broadcasting the first audio data and / or the second audio data during the process of presenting the live broadcast screen.
[0114] The live broadcast screen can be image data captured in real time, or a screen image displayed in real time on the screen of the device. The first part of the content to be broadcast and the second part of the content to be broadcast constitute the complete content to be broadcast. When broadcasting the first audio data and / or the second audio data during the process of presenting the live broadcast screen, the live broadcast screen and the audio data (the first audio data, the second audio data) can be output in association after being aligned in time.
[0115] The speech synthesis method of the present disclosure can also be implemented as a speech synthesis device. Figure 4 FIG. shows a schematic structural diagram of a voice interaction device according to an embodiment of the present disclosure. Among them, the functional modules of the speech synthesis device can be implemented by hardware, software, or a combination of hardware and software that implements the principles of the present disclosure. Those skilled in the art can understand that Figure 4 the described functional modules can be combined or divided into sub-modules, so as to implement the principles of the above invention. Therefore, the description herein can support any possible combination, or division, or further limitation of the functional modules described herein.
[0116] Brief descriptions will be given below of the functional modules that the speech synthesis device may have and the operations that each functional module can perform. For the detailed parts involved, reference can be made to the relevant descriptions above, and details will not be repeated here.
[0117] See Figure 4 , the speech synthesis device 400 includes a first synthesis module 410 and a second synthesis module 420.
[0118] In one embodiment of the present disclosure, the first synthesis module 410 may synthesize first audio data for the first text based on an online speech synthesis service. The second synthesis module 420 may, in response to the unavailability of the online speech synthesis service, synthesize second audio data for the second text based on a local speech synthesis service, and the similarity between the timbre of the second audio data and the timbre of at least part of the first audio data is greater than or equal to a first threshold.
[0119] The second synthesis module 420 may select an acoustic model from multiple acoustic models whose synthesized speech timbre has a similarity greater than or equal to the first threshold with the timbre of at least part of the first audio data, and use the selected acoustic model to perform speech synthesis on the second text to obtain the second audio data.
[0120] Specifically, the second synthesis module 420 may obtain third audio data synthesized for the text corresponding to at least part of the first audio data using different acoustic models respectively, calculate the similarity between each of the third audio data and at least part of the first audio data, and select the acoustic model corresponding to the third audio data with a higher similarity ranking or a similarity greater than or equal to a second threshold.
[0121] The second synthesis module 420 may also, in response to a first interaction instruction from the user, select an acoustic model for performing speech synthesis on the second text. And / or the second synthesis module 420 may further, in response to a second interaction instruction from the user, replace the selected acoustic model.
[0122] As an example, the online speech synthesis service synthesizes speech based on a speech library, the speech library includes one or more audio material sets, each audio material set corresponds to a speaker, and the local speech synthesis selects an acoustic model from multiple acoustic models to synthesize speech. The speech synthesis device 400 may further include a determination module for pre-determining an acoustic model corresponding to the audio material set. Among them, the timbre of the audio data synthesized based on the acoustic model corresponding to the audio material set is the same as or substantially the same as the audio material set. The second synthesis module 420 may, in response to the unavailability of the online speech synthesis service, use the acoustic model corresponding to the audio material set used when synthesizing the first audio data to synthesize the second audio data for the second text.
[0123] The speech synthesis device 400 may further include a storage module for storing the first audio data and / or the second audio data.
[0124] In another embodiment of the present disclosure, the first synthesis module 410 may synthesize third audio data for a third text based on a local speech synthesis service; the second synthesis module 420 may, in response to the availability of an online speech synthesis service, synthesize fourth audio data for a fourth text based on the online speech synthesis service, and the similarity between the timbre of the fourth audio data and the timbre of at least part of the third audio data is greater than or equal to a first threshold. For the operations that the second synthesis module 420 may perform, reference may be made to the relevant description above, and details will not be elaborated here.
[0125] The speech broadcast method of the present disclosure may also be implemented as a speech broadcast device. Figure 5 FIG. shows a schematic structural diagram of a speech broadcast device according to an embodiment of the present disclosure. Among them, the functional modules of the speech broadcast device may be implemented by hardware, software, or a combination of hardware and software that implements the principles of the present disclosure. Those skilled in the art can understand that Figure 5 the described functional modules may be combined or divided into sub-modules to implement the above-mentioned principles of the invention. Therefore, the description herein can support any possible combination, or division, or further limitation of the functional modules described herein.
[0126] Brief descriptions will be given below of the functional modules that the speech broadcast device may have and the operations that each functional module may perform. For the detailed parts involved, reference may be made to the relevant description above, and details will not be elaborated here.
[0127] See Figure 5 , the speech broadcast device 500 includes a first acquisition module 510, a second acquisition module 520, and a broadcast module 530.
[0128] In an embodiment of the present disclosure, the first acquisition module 510 is configured to acquire first audio data synthesized for a first text based on an online speech synthesis service. The second acquisition module 520 is configured to acquire second audio data synthesized for a second text based on a local speech synthesis service in response to the unavailability of the online speech synthesis service, and the similarity between the timbre of the second audio data and the timbre of at least part of the first audio data is greater than or equal to a first threshold. The broadcast module 530 is configured to broadcast the first audio data and / or the second audio data.
[0129] The speech broadcast device 500 may further include a storage module for storing the first audio data and / or the second audio data.
[0130] As an example, the text to be synthesized into speech is the content provided by the client application. The text to be synthesized into speech includes the first text and the second text. The voice broadcast device 500 may further include a judgment module and a third acquisition module. The judgment module is used to judge whether there is audio data corresponding to the text to be synthesized into speech in the server corresponding to the client application; the third acquisition module is used to acquire the audio data from the server when the judgment module determines that there is audio data corresponding to the text to be synthesized into speech in the server.
[0131] The judgment module may further judge whether there is audio data corresponding to the text to be synthesized into speech locally when it determines that there is no audio data corresponding to the text to be synthesized into speech in the server. The first acquisition module 510 may perform an operation of acquiring first audio data synthesized for the first text based on an online speech synthesis service when the judgment module determines that there is no audio data corresponding to the text to be synthesized into speech locally.
[0132] The voice broadcast device 500 may further include a fourth acquisition module and an analysis module. The fourth acquisition module is used to acquire the content to be broadcast, and the analysis module is used to analyze the content to be broadcast to determine the text to be synthesized into speech and / or parameter information related to voice broadcast existing in the content to be broadcast. The broadcast module 530 may broadcast the first audio data and / or the second audio data based on the parameter information.
[0133] In another embodiment of the present disclosure, the first acquisition module 510 is used to acquire third audio data synthesized for the third text based on a local speech synthesis service. The second acquisition module 520 is used to acquire fourth audio data synthesized for the fourth text based on an online speech synthesis service in response to the availability of the online speech synthesis service, and the similarity between the timbre of the fourth audio data and the timbre of at least part of the third audio data is greater than or equal to a first threshold. The broadcast module 530 is used to broadcast the third audio data and / or the fourth audio data.
[0134] The voice broadcast device 500 may further include a storage module for storing the third audio data and / or the fourth audio data.
[0135] As an example, the text to be synthesized into speech is the content provided by the client application. The text to be synthesized into speech includes the third text and the fourth text. The voice broadcast device 500 may further include a judgment module and a third acquisition module. The judgment module is used to judge whether there is audio data corresponding to the text to be synthesized into speech in the server corresponding to the client application; the third acquisition module is used to acquire the audio data from the server when the judgment module determines that there is audio data corresponding to the text to be synthesized into speech in the server.
[0136] The determination module can also determine whether there is audio data corresponding to the text of the speech to be synthesized on the server, and if not, determine whether there is audio data corresponding to the text of the speech to be synthesized locally. The first acquisition module 510 can, when the determination module determines that there is no audio data corresponding to the text of the speech to be synthesized locally, perform an operation of acquiring third audio data synthesized for a third text based on a local speech synthesis service.
[0137] The voice broadcast device 500 may further include a fourth acquisition module and a parsing module. The fourth acquisition module is used to acquire the content to be broadcast, and the parsing module is used to parse the content to be broadcast to determine the text of the speech to be synthesized and / or parameter information related to voice broadcast existing in the content to be broadcast. The broadcast module 530 can broadcast the third audio data and / or the fourth audio data based on the parameter information.
[0138] The teaching method of the present disclosure can also be implemented as a teaching device. Figure 6 FIG. shows a schematic structural diagram of a teaching device according to an embodiment of the present disclosure. Among them, the functional modules of the teaching device can be implemented by hardware, software, or a combination of hardware and software that implements the principles of the present disclosure. Those skilled in the art can understand that Figure 6 the described functional modules can be combined or divided into sub-modules to implement the principles of the above invention. Therefore, the description herein can support any possible combination, or division, or further limitation of the functional modules described herein.
[0139] Brief descriptions will be given below of the functional modules that the teaching device may have and the operations that each functional module can perform. For the detailed parts involved, reference can be made to the relevant descriptions above, and details will not be repeated here.
[0140] See Figure 6 , the teaching device 600 includes a determination module 610, a first acquisition module 620, a second acquisition module 630, and a broadcast module 640.
[0141] The determination module 610 is used to determine an answer corresponding to a question raised by a student. The first acquisition module 620 is used to acquire first audio data synthesized for a first part of the answer based on an online speech synthesis service. The second acquisition module 630 is used to acquire second audio data synthesized for a second part of the answer based on a local speech synthesis service in response to the unavailability of the online speech synthesis service, and the timbre of the second audio data has a similarity greater than or equal to a first threshold with at least part of the timbre of the first audio data. The broadcast module 640 is used to broadcast the first audio data and / or the second audio data.
[0142] The live broadcast method of the present disclosure can also be implemented as a live broadcast device. Figure 7The structural schematic diagram of a live broadcast device according to an embodiment of the present disclosure is shown. Among them, the functional modules of the live broadcast device can be implemented by hardware, software, or a combination of hardware and software that implements the principles of the present disclosure. Those skilled in the art can understand that Figure 7 the described functional modules can be combined or divided into sub-modules to implement the principles of the above invention. Therefore, the description herein can support any possible combination, division, or further limitation of the functional modules described herein.
[0143] A brief description of the functional modules that the live broadcast device may have and the operations that each functional module can perform will be given below. For the detailed parts involved, reference can be made to the relevant descriptions above, and details will not be repeated here.
[0144] Referring to Figure 7 , the live broadcast device 700 includes a first acquisition module 710, a second acquisition module 720, and a broadcast module 730.
[0145] The first acquisition module 710 is used to acquire first audio data synthesized for the first part of the content to be broadcast based on an online speech synthesis service. The second acquisition module 720 is used to acquire second audio data synthesized for the second part of the content to be broadcast based on a local speech synthesis service in response to the unavailability of the online speech synthesis service. The similarity between the timbre of the second audio data and the timbre of at least part of the first audio data is greater than or equal to a first threshold. The broadcast module 730 is used to broadcast the first audio data and / or the second audio data during the process of presenting the live broadcast screen.
[0146] Figure 8 The structural schematic diagram of a computing device that can be used to implement any one of the above speech synthesis method, speech broadcast method, teaching method, and live broadcast method according to an embodiment of the present invention is shown.
[0147] Referring to Figure 8 , the computing device 800 includes a memory 810 and a processor 820.
[0148] The processor 820 can be a multi-core processor or can include multiple processors. In some embodiments, the processor 820 can include a general main processor and one or more special coprocessors, such as a graphics processing unit (GPU), a digital signal processor (DSP), etc. In some embodiments, the processor 820 can be implemented using custom circuits, such as an application specific integrated circuit (ASIC) or a field programmable gate array (FPGA).
[0149] The memory 810 may include various types of storage units, such as system memory, read-only memory (ROM), and permanent storage devices. Among them, the ROM can store static data or instructions required by the processor 820 or other modules of the computer. The permanent storage device can be a read-write storage device. The permanent storage device can be a non-volatile storage device that does not lose the stored instructions and data even when the computer is powered off. In some embodiments, the permanent storage device uses a mass storage device (such as a magnetic or optical disk, flash memory) as the permanent storage device. In some other embodiments, the permanent storage device can be a removable storage device (such as a floppy disk, optical drive). The system memory can be a read-write storage device or a volatile read-write storage device, such as dynamic random access memory. The system memory can store some or all of the instructions and data required by the processor during operation. In addition, the memory 810 can include any combination of computer-readable storage media, including various types of semiconductor storage chips (DRAM, SRAM, SDRAM, flash memory, programmable read-only memory), and magnetic disks and / or optical disks can also be used. In some embodiments, the memory 810 can include a removable storage device that is readable and / or writable, such as a compact disc (CD), read-only digital versatile disc (such as DVD-ROM, dual-layer DVD-ROM), read-only Blu-ray disc, super density disc, flash memory card (such as SD card, min SD card, Micro-SD card, etc.), magnetic floppy disk, etc. Computer-readable storage media do not include carrier waves and instantaneous electronic signals transmitted wirelessly or wired.
[0150] Executable code is stored on the memory 810, and when the executable code is processed by the processor 820, it can cause the processor 820 to execute any one of the above-mentioned speech synthesis method, speech broadcast method, teaching method, and live broadcast method.
[0151] The speech synthesis method, speech broadcast method, teaching method, live broadcast method, corresponding devices and equipment according to the present invention have been described in detail above with reference to the accompanying drawings.
[0152] In addition, the method according to the present invention can also be implemented as a computer program or computer program product, which includes computer program code instructions for executing the above steps defined in the above method of the present invention.
[0153] Alternatively, the present invention can also be implemented as a non-transitory machine-readable storage medium (or computer-readable storage medium, or machine-readable storage medium) having executable code (or computer program, or computer instruction code) stored thereon, which when executed by a processor of an electronic device (or computing device, server, etc.), causes the processor to perform the various steps of the above-described method according to the present invention.
[0154] Those skilled in the art will also understand that the various exemplary logical blocks, modules, circuits, and algorithm steps described in connection with the disclosure herein can be implemented as electronic hardware, computer software, or a combination of both.
[0155] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems and methods according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a portion of code, which contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the figures. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or by a combination of dedicated hardware and computer instructions.
[0156] The embodiments of the present invention have been described above. The above description is exemplary and not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The selection of the terms used herein is intended to best explain the principles of the embodiments, the practical application, or the improvement of the technology in the market, or to enable other ordinary skilled persons in the art to understand the embodiments disclosed herein.
Claims
1. A voice synthesis method, comprising: Synthesizing first audio data for a first text based on an online voice synthesis service; In response to the unavailability of the online voice synthesis service, synthesizing second audio data for a second text based on a local voice synthesis service, wherein the similarity between the timbre of the second audio data and the timbre of at least part of the first audio data is greater than or equal to a first threshold; Wherein, the step of synthesizing the second audio data for the second text based on the local voice synthesis service includes: obtaining third audio data synthesized for the text corresponding to at least part of the first audio data using different acoustic models respectively; calculating the similarity between each of the third audio data and the at least part of the first audio data; selecting an acoustic model based on the similarity to obtain the second audio data, and the acoustic model is used to represent a speaker model corresponding to a timbre trained using a machine learning algorithm, and each acoustic model corresponds to a speaker with a certain timbre.
2. The method according to claim 1, wherein The step of selecting an acoustic model based on the similarity to obtain the second audio data includes: Selecting an acoustic model based on the similarity, and performing voice synthesis on the second text using the selected acoustic model to obtain the second audio data.
3. The method according to claim 2, wherein, The step of selecting an acoustic model based on the similarity includes: Selecting the acoustic model corresponding to the third audio data with a relatively high similarity ranking or a similarity greater than or equal to a second threshold.
4. The method according to claim 2, further comprising: In response to a first interaction instruction of a user, selecting an acoustic model for performing voice synthesis on the second text from multiple acoustic models, and / or In response to a second interaction instruction of the user, replacing the selected acoustic model.
5. The method according to claim 2, wherein The acoustic model is a speaker model trained using a machine learning algorithm, and the speaker model is used to output acoustic parameters or audio data.
6. The method according to claim 1, wherein, The online voice synthesis service synthesizes voices based on a voice library, the voice library includes one or more audio material sets, each audio material set corresponds to a speaker, and the local voice synthesis service selects an acoustic model from multiple acoustic models to synthesize voices. The method further includes: Pre-determining an acoustic model corresponding to the audio material set, wherein the timbre of the audio data synthesized based on the acoustic model corresponding to the audio material set is the same as or substantially the same as the audio material set, The step of synthesizing the second audio data for the second text based on the local voice synthesis service includes: using the acoustic model corresponding to the audio material set used when synthesizing the first audio data to synthesize the second audio data for the second text.
7. The method according to claim 1, further comprising: Saving the first audio data and / or the second audio data.
8. A voice synthesis method, comprising: Synthesizing third audio data for a third text based on a local voice synthesis service; In response to the availability of the online voice synthesis service, synthesizing fourth audio data for a fourth text based on the online voice synthesis service, wherein the similarity between the timbre of the fourth audio data and the timbre of at least part of the third audio data is greater than or equal to a first threshold; The step of synthesizing fourth audio data for a fourth text based on the online speech synthesis service includes: obtaining fifth audio data synthesized for the text corresponding to at least part of the third audio data using different acoustic models respectively; calculating the similarity between each of the fifth audio data and the at least part of the third audio data; selecting an acoustic model based on the similarity to obtain the fourth audio data, where the acoustic model is obtained by training using a machine learning algorithm, and each acoustic model corresponds to a speaker model with a specific timbre.
9. A voice broadcast method, comprising: obtaining first audio data synthesized for a first text based on an online speech synthesis service; in response to the unavailability of the online speech synthesis service, obtaining second audio data synthesized for a second text based on a local speech synthesis service, where the similarity between the timbre of the second audio data and the timbre of at least part of the first audio data is greater than or equal to a first threshold; and broadcasting the first audio data and / or the second audio data; where the step of obtaining second audio data synthesized for a second text based on a local speech synthesis service includes: obtaining third audio data synthesized for the text corresponding to at least part of the first audio data using different acoustic models respectively; calculating the similarity between each of the third audio data and the at least part of the first audio data; selecting an acoustic model based on the similarity to obtain the second audio data, where the acoustic model is obtained by training using a machine learning algorithm, and each acoustic model corresponds to a speaker model with a specific timbre.
10. The method according to claim 9, further comprising: saving the first audio data and / or the second audio data.
11. The method according to claim 9, wherein The text of the speech to be synthesized is the content provided by the client application, and the text of the speech to be synthesized includes the first text and the second text. The method further comprises: judging whether there is audio data corresponding to the text of the speech to be synthesized on the server corresponding to the client application; in the case of determining that the server has audio data corresponding to the text of the speech to be synthesized, obtaining the audio data from the server.
12. The method according to claim 11, further comprising: in the case of determining that the server does not have audio data corresponding to the text of the speech to be synthesized, judging whether there is audio data corresponding to the text of the speech to be synthesized locally; in the case of determining that there is no audio data corresponding to the text of the speech to be synthesized locally, performing the step of obtaining first audio data synthesized for a first text based on an online speech synthesis service.
13. The method according to claim 9, further comprising: obtaining the content to be broadcast; parsing the content to be broadcast to determine the text of the speech to be synthesized and / or parameter information related to voice broadcast existing in the content to be broadcast, where the step of broadcasting the first audio data and / or the second audio data includes: broadcasting the first audio data and / or the second audio data based on the parameter information.
14. A voice broadcast method, comprising: Obtain third audio data synthesized for a third text based on a local speech synthesis service; In response to the availability of an online speech synthesis service, obtain fourth audio data synthesized for a fourth text based on the online speech synthesis service, where the similarity between the timbre of the fourth audio data and the timbre of at least part of the third audio data is greater than or equal to a first threshold; And Broadcast the third audio data and / or the fourth audio data; Wherein, the step of obtaining the fourth audio data synthesized for the fourth text based on the online speech synthesis service includes: obtaining fifth audio data synthesized for the text corresponding to at least part of the third audio data using different acoustic models respectively; calculating the similarity between each of the fifth audio data and the at least part of the third audio data; selecting an acoustic model based on the similarity to obtain the fourth audio data, where the acoustic model is obtained by training using a machine learning algorithm, and each acoustic model corresponds to a speaker model with a specific timbre.
15. A teaching method, including: Determine an answer corresponding to a question raised by a student; Obtain first audio data synthesized for a first part of the answer based on an online speech synthesis service; In response to the unavailability of the online speech synthesis service, obtain second audio data synthesized for a second part of the answer based on a local speech synthesis service, where the similarity between the timbre of the second audio data and the timbre of at least part of the first audio data is greater than or equal to a first threshold; And Broadcast the first audio data and / or the second audio data; Wherein, the step of obtaining the second audio data synthesized for the second part of the answer based on the local speech synthesis service includes: obtaining third audio data synthesized for the text corresponding to at least part of the first audio data using different acoustic models respectively; calculating the similarity between each of the third audio data and the at least part of the first audio data; selecting an acoustic model based on the similarity to obtain the second audio data, where the acoustic model is obtained by training using a machine learning algorithm, and each acoustic model corresponds to a speaker model with a specific timbre.
16. A live broadcast method, including: Obtain first audio data synthesized for a first part of the content to be broadcast based on an online speech synthesis service; In response to the unavailability of the online speech synthesis service, obtain second audio data synthesized for a second part of the content to be broadcast based on a local speech synthesis service, where the similarity between the timbre of the second audio data and the timbre of at least part of the first audio data is greater than or equal to a first threshold; And During the presentation of the live broadcast screen, broadcast the first audio data and / or the second audio data; The step of obtaining the second audio data synthesized by the local speech synthesis service for the second part of the content to be broadcast includes: obtaining third audio data synthesized for the text corresponding to at least part of the first audio data by using different acoustic models respectively; calculating the similarity between each of the third audio data and the at least part of the first audio data; selecting an acoustic model based on the similarity to obtain the second audio data, where the acoustic model is obtained by training using a machine learning algorithm, and each acoustic model corresponds to a speaker model with a specific timbre.
17. A speech synthesis device, comprising: A first synthesis module, configured to synthesize first audio data for a first text based on an online speech synthesis service; And A second synthesis module, configured to, in response to the unavailability of the online speech synthesis service, synthesize second audio data for a second text based on a local speech synthesis service, where the similarity between the timbre of the second audio data and the timbre of at least part of the first audio data is greater than or equal to a first threshold; Wherein, the second synthesis module is further configured to obtain third audio data synthesized for the text corresponding to at least part of the first audio data by using different acoustic models respectively; calculate the similarity between each of the third audio data and the at least part of the first audio data; select an acoustic model based on the similarity to obtain the second audio data, where the acoustic model is obtained by training using a machine learning algorithm, and each acoustic model corresponds to a speaker model with a specific timbre.
18. A speech synthesis device, comprising: A first synthesis module, configured to synthesize third audio data for a third text based on a local speech synthesis service; A second synthesis module, configured to, in response to the availability of the online speech synthesis service, synthesize fourth audio data for a fourth text based on the online speech synthesis service, where the similarity between the timbre of the fourth audio data and the timbre of at least part of the third audio data is greater than or equal to a first threshold; Wherein, the second synthesis module is further configured to obtain fifth audio data synthesized for the text corresponding to at least part of the third audio data by using different acoustic models respectively; calculate the similarity between each of the fifth audio data and the at least part of the third audio data; select an acoustic model based on the similarity to obtain the fourth audio data, where the acoustic model is obtained by training using a machine learning algorithm, and each acoustic model corresponds to a speaker model with a specific timbre.
19. A speech broadcast device, comprising: A first acquisition module, configured to acquire first audio data synthesized for a first text based on an online speech synthesis service; A second acquisition module, configured to, in response to the unavailability of the online speech synthesis service, acquire second audio data synthesized for a second text based on a local speech synthesis service, where the similarity between the timbre of the second audio data and the timbre of at least part of the first audio data is greater than or equal to a first threshold; And A broadcast module, configured to broadcast the first audio data and / or the second audio data; The second acquisition module is further configured to acquire third audio data synthesized for the text corresponding to at least part of the first audio data by using different acoustic models respectively; calculate the similarity between each of the third audio data and the at least part of the first audio data; select an acoustic model based on the similarity to obtain the second audio data, where the acoustic model is obtained by training using a machine learning algorithm, and each acoustic model corresponds to a speaker model with a specific voice quality.
20. A voice broadcast device, comprising: A first acquisition module, configured to acquire third audio data synthesized for a third text based on a local voice synthesis service; A second acquisition module, configured to, in response to the availability of an online voice synthesis service, acquire fourth audio data synthesized for a fourth text based on the online voice synthesis service, where the similarity between the voice quality of the fourth audio data and the voice quality of at least part of the third audio data is greater than or equal to a first threshold; And A broadcast module, configured to broadcast the third audio data and / or the fourth audio data; The second acquisition module is further configured to acquire fifth audio data synthesized for the text corresponding to at least part of the third audio data by using different acoustic models respectively; calculate the similarity between each of the fifth audio data and the at least part of the third audio data; select an acoustic model based on the similarity to obtain the fourth audio data, where the acoustic model is obtained by training using a machine learning algorithm, and each acoustic model corresponds to a speaker model with a specific voice quality.
21. An instructional device, comprising: A determination module, configured to determine an answer corresponding to a question raised by a student; A first acquisition module, configured to acquire first audio data synthesized for a first part of the answer based on an online voice synthesis service; A second acquisition module, configured to, in response to the unavailability of the online voice synthesis service, acquire second audio data synthesized for a second part of the answer based on a local voice synthesis service, where the similarity between the voice quality of the second audio data and the voice quality of at least part of the first audio data is greater than or equal to a first threshold; And A broadcast module, configured to broadcast the first audio data and / or the second audio data; The second acquisition module is further configured to acquire third audio data synthesized for the text corresponding to at least part of the first audio data by using different acoustic models respectively; calculate the similarity between each of the third audio data and the at least part of the first audio data; select an acoustic model based on the similarity to obtain the second audio data, where the acoustic model is obtained by training using a machine learning algorithm, and each acoustic model corresponds to a speaker model with a specific voice quality.
22. A live broadcast device, comprising: A first acquisition module, configured to acquire first audio data synthesized for a first part of the content to be broadcast based on an online voice synthesis service; A second acquisition module, configured to, in response to the unavailability of the online speech synthesis service, acquire second audio data synthesized based on a local speech synthesis service for a second part of the content to be broadcast, where the similarity between the timbre of the second audio data and the timbre of at least part of the first audio data is greater than or equal to a first threshold; And A broadcast module, configured to broadcast the first audio data and / or the second audio data during the presentation of the live video; Wherein, the second acquisition module is further configured to acquire third audio data synthesized for the text corresponding to at least part of the first audio data respectively using different acoustic models; calculate the similarity between each of the third audio data and the at least part of the first audio data; select an acoustic model based on the similarity to obtain the second audio data, where the acoustic model is obtained by training using a machine learning algorithm, and each acoustic model corresponds to a speaker model with a specific timbre.
23. A computing device, comprising: A processor; And A memory, storing executable code thereon, which, when executed by the processor, causes the processor to execute the method according to any one of claims 1 to 16.
24. A non-transitory machine-readable storage medium, storing executable code thereon, which, when executed by a processor of an electronic device, causes the processor to execute the method according to any one of claims 1 to 16.
Citation Information
Patent Citations
Voice prompt generation combining native and remotely generated speech data
CN106575501A
Voice synthesis processing method and device
CN107039032A
Text speech synthesis method after speaker emotion simulated optimization translation
CN108831436A