Synthetic voice switching method and related apparatus, device and storage medium

By maintaining the playback of the first synthesized voice and matching the progress of the second synthesized voice when a configuration parameter change is detected, the problem of unsmooth switching of synthesized voices is solved, and seamless voice playback is achieved.

CN118800213BActive Publication Date: 2025-10-14ANHUI IFLYREC TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410930266.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-11
Publication Date
2025-10-14
Estimated Expiration
2044-07-11

AI Technical Summary

Technical Problem

In the prior art, there is a problem of low fluency when switching synthesized voices, resulting in unsmooth voice switching.

Method used

When a configuration parameter change is detected, the first synthesized speech is continued to be played, and the second synthesized speech is synthesized based on the second configuration parameter. The playback progress of the first synthesized speech is matched with the playback progress of the second synthesized speech to achieve smooth switching.

Benefits of technology

The smoothness of synthesized voice switching has been improved, seamless playback of synthesized voice has been achieved, and the continuity of voice switching has been improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118800213B_ABST
    Figure CN118800213B_ABST
Patent Text Reader

Abstract

The application discloses a synthesized speech switching method and related device, equipment and storage medium, wherein the synthesized speech switching method comprises the following steps: playing first synthesized speech synthesized based on first configuration parameters and to-be-synthesized text; in response to detecting a control instruction representing re-synthesizing speech based on second configuration parameters, synthesizing speech based on the second configuration parameters and the to-be-synthesized text to obtain second synthesized speech; determining a second playing progress in the second synthesized speech matched with a first playing progress of the first synthesized speech at a second synthesized speech synthesis completion time; and switching to play the second synthesized speech from the second playing progress of the second synthesized speech. The above scheme can improve the fluency of synthesized speech switching and realize seamless connection playing of synthesized speech switching as much as possible.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of speech processing, in particular to a synthesized speech switching method and related device, equipment and storage medium. BACKGROUND

[0002] With the continuous development of deep learning technology, speech synthesis technology has been developed to automatically convert text into speech.

[0003] In the prior art, when playing synthesized speech, in order to support the user to perform new parameter configuration on the synthesized speech during the playing process, the current playing is first paused, and after the new speech synthesis is completed, the playing of the new synthesized speech is switched. However, waiting for new speech synthesis will cause the synthesized speech switching to be not smooth, and thus the fluency of the synthesized speech switching is low. Therefore, how to improve the fluency of the synthesized speech switching and realize seamless connection playing of the synthesized speech switching as much as possible has become a problem to be solved. SUMMARY

[0004] The technical problem solved by the present application is to provide a synthesized speech switching method and related device, equipment and storage medium, which can improve the fluency of the synthesized speech switching and realize seamless connection playing of the synthesized speech switching as much as possible.

[0005] To solve the above technical problem, the first aspect of the present application provides a synthesized speech switching method, comprising: playing a first synthesized speech obtained by synthesizing a to-be-synthesized text based on first configuration parameters; in response to detecting a control instruction representing re-synthesizing speech based on second configuration parameters, synthesizing speech for the to-be-synthesized text based on the second configuration parameters to obtain a second synthesized speech; determining a second playing progress in the second synthesized speech that matches a first playing progress of the first synthesized speech at a time when the second synthesized speech is synthesized; and switching to play the second synthesized speech from the second playing progress of the second synthesized speech.

[0006] To solve the above technical problem, the second aspect of the present application provides a synthesized speech switching device, comprising: a playing module, a synthesizing module, a determining module and a switching module, the playing module is configured to play a first synthesized speech obtained by synthesizing a to-be-synthesized text based on first configuration parameters; the synthesizing module is configured to, in response to detecting a control instruction representing re-synthesizing speech based on second configuration parameters, synthesize speech for the to-be-synthesized text based on the second configuration parameters to obtain a second synthesized speech; the determining module is configured to determine a second playing progress in the second synthesized speech that matches a first playing progress of the first synthesized speech at a time when the second synthesized speech is synthesized; and the switching module is configured to switch to play the second synthesized speech from the second playing progress of the second synthesized speech.

[0007] In order to solve the above technical problems, the third aspect of the present application provides an electronic device, including a memory and a processor coupled to each other, wherein the memory stores program instructions, and the processor is used to execute the program instructions to implement the synthetic voice switching method in the above first aspect.

[0008] In order to solve the above technical problems, the fourth aspect of the present application provides a computer-readable storage medium storing program instructions that can be executed by a processor, and the program instructions are used to implement the synthetic speech switching method described in the first aspect above.

[0009] The above scheme plays the first synthesized speech synthesized based on the first configuration parameters for the synthesized text. When a control instruction representing re-speech synthesis based on the second configuration parameters is detected, the first synthesized speech continues to be played, and speech synthesis is performed on the synthesized text based on the second configuration parameters to obtain the second synthesized speech. Then, based on the first playback progress of the first synthesized speech at the moment when the synthesis of the second synthesized speech is completed, the second playback progress in the second synthesized speech that matches the first playback progress is determined, and the second synthesized speech can be switched to play starting from the second playback progress of the second synthesized speech. Therefore, during the synthesis of the second synthesized speech, the playback of the first synthesized speech is maintained, and based on the moment when the synthesis of the second synthesized speech is completed, the second playback progress in the second synthesized speech that matches the first playback progress is switched to achieve smooth switching of the synthesized speech. Therefore, the fluency of the switching of the synthesized speech can be improved, and the seamless playback of the switching of the synthesized speech can be achieved as much as possible. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] Figure 1 This is a flow chart of an embodiment of the method for switching synthesized speech according to the present invention;

[0011] Figure 2 This is a schematic diagram of the framework of an embodiment of the synthetic voice switching device of the present application;

[0012] Figure 3 This is a schematic diagram of the framework of an embodiment of the electronic device of the present application;

[0013] Figure 4 It is a schematic diagram of a framework of an embodiment of a computer-readable storage medium of the present application. DETAILED DESCRIPTION

[0014] The following describes the embodiments of the present application in detail with reference to the accompanying drawings.

[0015] In the following description, for the purpose of explanation rather than limitation, specific details such as specific system structures, interfaces, and technologies are provided to facilitate a thorough understanding of the present application.

[0016] The terms "system" and "network" are often used interchangeably in this document. The term "and / or" is simply a description of an association between related objects, indicating that three possible relationships exist. For example, "A and / or B" can mean: A exists alone, A and B exist simultaneously, or B exists alone. Furthermore, the fragment " / " generally indicates an "or" relationship between the related objects. Furthermore, "multiple" in this document refers to two or more than two.

[0017] See also Figure 1 , Figure 1 This is a flow chart of an embodiment of the method for switching synthesized speech according to the present invention. Specifically, the method may include the following steps:

[0018] Step S10: Play the first synthesized speech synthesized from the text to be synthesized based on the first configuration parameters.

[0019] In the disclosed embodiments, the text structure of the text to be synthesized is not limited in this application. For example, the text to be synthesized may include multiple sentences, each of which is described based on a natural language text and can be used for communication, expressing opinions, conveying information, etc. It should be noted that the language system used in the text to be synthesized is not limited in this application and can be, for example, Chinese, English, French, etc., or a combination of multiple language systems. For the sake of brevity, we will not elaborate on each of them here.

[0020] In one implementation scenario, speech synthesis can be performed on the synthesized text based on speech synthesis technology to obtain a first synthesized speech. Speech synthesis technology may include, but is not limited to, an encoder-decoder network model, etc., and is not limited here. Furthermore, the specific process of speech synthesis can be found in the technical details of speech synthesis technologies such as the encoder-decoder network model, and will not be further described here.

[0021] In a specific implementation scenario, as a possible implementation manner, a speech synthesis model can be pre-trained, which can include but is not limited to a network model of an Encoder-Decoder architecture, and the like. The speech synthesis model can be input with the text to be synthesized, and an output result of the speech synthesis model can be used as the first synthesized speech of the text to be synthesized. In order to ensure the synthesis accuracy of the speech synthesis model as much as possible, sample texts and corresponding sample configuration parameters can be collected, and the sample texts are labeled with expected synthesized speech. On this basis, the sample texts and the sample configuration parameters can be processed based on the speech synthesis model to obtain sample synthesized speech of the sample texts. Thus, the network parameters of the speech synthesis model can be adjusted based on the difference between the expected synthesized speech and the sample synthesized speech until the speech synthesis model converges. The text to be synthesized can be processed based on the speech synthesis model that converges to obtain the first synthesized speech. It should be noted that the specific processing process of the speech synthesis model can refer to technical details of a network model such as an Encoder-Decoder architecture, and the like, which will not be described herein.

[0022] In a specific implementation scenario, the first configuration parameter can include an initialization parameter, that is, the first configuration parameter consistent with the preset parameter can be generated without additional manual configuration of the user to generate the initialized first synthesized speech. It should be noted that the first configuration parameter can also be manually configured by the user, which is not limited herein.

[0023] In an implementation scenario, the first configuration parameter includes a synthesized tone, an emotion coefficient, an audio sampling rate, a playback speed, and the like. The first configuration parameter can be used to determine the acoustic characteristics and prosodic characteristics of the first synthesized speech, which is not limited herein.

[0024] In a specific implementation scenario, the prosodic characteristics of the first synthesized speech can be extracted based on a prosodic extraction network based on the first configuration parameter. Specifically, the prosodic extraction network is constructed based on the network structure of a large model. For example, a chapter prosodic extraction network is constructed based on the network structure of a Transformer model, a GPT (Generative Pre-trained Transformer) model, and the like. The large model refers to a machine learning model with large-scale parameters and complex computing structure. These models are usually constructed by deep neural networks, have tens of billions or even hundreds of billions of parameters, and can improve the expression ability and prediction performance of the model to process more complex tasks and data.

[0025] In a specific implementation scenario, based on the large model-based network structure initialization, an initial structure of the prosody extraction network is obtained, target sample texts and target configuration parameters used for training of the prosody extraction network are obtained, target prosodic features are taken as output targets of the prosody extraction network, the target sample texts and the target configuration parameters are taken as input targets of the discourse prosody extraction network, and the prosody extraction network is trained in a self-recurrent manner until convergence, so as to update network parameters of the prosody extraction network. The above method can improve the prosodic quality of the synthesized speech.

[0026] In an implementation scenario, as a possible implementation, before the speech synthesis of the text to be synthesized based on the first configuration parameters, the text to be synthesized is split, for example, according to the end punctuation marks in the text to be synthesized, such as exclamation marks, periods, question marks, etc., for initial splitting, and then according to the text length of each subtext after the initial splitting, for secondary splitting, for example, setting a preset maximum number of words and a preset minimum number of words, and merging the subtexts with insufficient word numbers in sequence or splitting the subtexts with excessive word numbers again to form a plurality of text segments, so that the number of characters contained in each text segment is between the preset maximum number of words and the preset minimum number of words. Then, according to the order of the text segments in the text to be synthesized, the first synthesized speech is synthesized in sequence. Of course, considering that under the premise of sufficient network environment and hardware resources, the generation process of the first synthesized speech corresponding to multiple texts can usually be performed simultaneously, so the first synthesized speech can also be generated without being based on the description order of each text segment. The above examples are only several typical examples in the actual application process, and do not limit the generation timing of the first synthesized speech.

[0027] In an implementation scenario, after the text to be synthesized is synthesized based on the first configuration parameters to obtain the first synthesized speech, the first synthesized speech is stored in a playlist, and the first synthesized speech in the playlist is selected for playing in response to a playing instruction of the first synthesized speech.

[0028] In a specific implementation scenario, the playing instruction of the first synthesized speech can be automatically triggered after a preset time length, or a prompt message is generated to accept a user's confirmation playing instruction, and the specific implementation is not limited in the present application.

[0029] In an implementation scenario, the method displays an interactive interface, and the interactive interface displays the text to be synthesized and is provided with a parameter configuration control for triggering a control instruction.

[0030] In a specific implementation scenario, a subtext in the text to be synthesized that matches the playing progress of the first synthesized speech is highlighted in the interactive interface with a preset marker, specifically, the preset marker is highlighting, underlining, special color text, etc.

[0031] In one specific implementation scenario, the user can select the part of the to-be-composed text displayed in the interactive interface and switch to play the first synthesized speech corresponding to the selected to-be-composed text.

[0032] In one specific implementation scenario, in response to the user's confirmation trigger on the parameter configuration control, the configuration parameter is adjusted to obtain a second configuration parameter. Specifically, the parameter configuration control is a digital input, a sliding adjustment, etc. The specific configuration method is not limited in the present application.

[0033] Step S20: In response to detecting a control instruction representing re-composition of speech based on the second configuration parameter, the to-be-composed text is composed into speech based on the second configuration parameter to obtain a second synthesized speech.

[0034] In one implementation scenario, after detecting the control instruction representing re-composition of speech based on the second configuration parameter, the to-be-composed text is composed into speech based on the second configuration parameter to obtain a second synthesized speech. The synthesis method of the second synthesized speech can refer to the synthesis method of the first synthesized speech described above. For brevity, it will not be described here.

[0035] In one implementation scenario, when the to-be-composed text is composed into speech based on the second configuration parameter, prompt information for prompting that the second synthesized speech is being composed is output. Specifically, the prompt information can be in the form of a pop-up window, etc., and the prompt information is stopped when the second synthesized speech is composed.

[0036] In one specific implementation scenario, as one possible implementation, the prompt information for prompting that the second synthesized speech is being composed is output in another area in the interactive interface that is different from the area where the to-be-composed text is displayed, i.e., the prompt information does not affect the playing of the first synthesized speech and the display of the corresponding subtext, so as to improve the fluency of the synthesized speech switching.

[0037] In one implementation scenario, the to-be-composed text includes a plurality of subtexts, the first synthesized speech includes a plurality of first sub-speeches of the subtexts, each subtext is composed into speech based on the second configuration parameter in sequence to obtain a second sub-speech, and the second synthesized speech includes the second sub-speeches. The above method realizes the generation of the second synthesized speech based on the subtexts, which can reduce the required computing power and improve the efficiency of composition.

[0038] In one specific implementation scenario, the subtext corresponding to the target sub-speech and the subsequent subtexts are selected as target subtexts, respectively, and the target sub-speech is the first sub-speech played at the trigger time of the control instruction. Each target subtext is composed into speech based on the second configuration parameter in sequence to obtain a second sub-speech corresponding to the target subtext, and the second synthesized speech includes at least the second sub-speeches of the target subtexts. The above method can reduce the required computing power and improve the practicality of the synthesized speech switching.

[0039] In one specific implementation scenario, the historical sub-texts are sequentially subjected to speech synthesis based on the second configuration parameter to obtain second sub-audio corresponding to the historical sub-texts, the historical sub-texts being the sub-texts corresponding to the first sub-audio that have been played to the end at the triggering time of the control instruction, the second sub-audio of the historical sub-texts being sequentially added to the second sub-audio of the target sub-texts to obtain second synthesized audio, the target sub-texts being the sub-texts corresponding to the first sub-audio that are being played and have not been played at the triggering time of the control instruction, the historical sub-texts being subjected to speech synthesis after the target sub-texts. For example, the text to be synthesized includes sequentially composed sub-text A, sub-text B, sub-text C, sub-text D, and sub-text E, the sub-texts corresponding to the first sub-audio that have been played to the end at the triggering time of the control instruction being sub-text A and sub-text B, sub-text C, sub-text D, and sub-text E being taken as target sub-texts, the target sub-texts being sequentially subjected to speech synthesis based on the second configuration parameter, and then sub-text A and sub-text B being sequentially synthesized. Specifically, the historical sub-texts adjacent to the target sub-texts can be taken as the first synthesis order, i.e., sub-text B is preferentially synthesized, and the second sub-audio of the historical sub-texts is sequentially added to the second sub-audio of the target sub-texts to obtain second synthesized audio.

[0040] In one implementation scenario, the synthesized second synthesized audio is stored in a play list, and the first synthesized audio is also stored in the play list, facilitating the user to subsequently view historical synthesized audio.

[0041] Step S30: determining a second play progress in the second synthesized audio that matches the first play progress based on the first play progress of the first synthesized audio at the time when the second synthesized audio is synthesized.

[0042] In one implementation scenario, the first synthesized audio synthesized based on the first configuration parameter is played, when a control instruction indicating that speech synthesis is to be performed again based on the second configuration parameter is detected, the text to be synthesized is subjected to speech synthesis based on the second configuration parameter to obtain second synthesized audio, a second play progress in the second synthesized audio that matches a first play progress of the first synthesized audio at the time when the second synthesized audio is synthesized is determined, and the second synthesized audio is switched to be played from the second play progress. Thus, during the synthesis of the second synthesized audio, the playing of the first synthesized audio is maintained, and the second synthesized audio is switched to be played at the second play progress that matches the first play progress based on the time when the second synthesized audio is synthesized, thereby achieving smooth switching of synthesized audio. Therefore, the fluency of switching of synthesized audio is improved, and seamless connection of playing of synthesized audio is achieved as much as possible.

[0043] In an implementation scenario, a first audio duration of the first synthesized speech is obtained, and a second audio duration of the second synthesized speech is obtained. A second playback progress that matches the first playback progress is obtained based on a duration ratio between the second audio duration and the first audio duration and the first playback progress. For example, the first playback progress indicates that the first synthesized speech has been played for 50% at the moment when the second synthesized speech synthesis is completed. The first audio duration is 60 seconds, i.e., the first synthesized speech is played to the 30th second. The second audio duration is 90 seconds, and the duration ratio between the second audio duration and the first audio duration is 1.5. Then, the second playback progress that matches the first playback progress is the 45th second of the second synthesized speech.

[0044] In an implementation scenario, the text to be synthesized includes a plurality of subtexts. The first synthesized speech includes a plurality of first sub-speeches of the subtexts. The second configuration parameter is used to sequentially synthesize speech for each subtext to obtain a plurality of second sub-speeches. The first sub-speech being played is selected as a current sub-speech. During playing of the current sub-speech, at least a synthesis progress of an expected sub-speech is detected. The expected sub-speech is a second sub-speech of a same subtext as the current sub-speech. In response to the at least expected sub-speech being synthesized, a second playback progress of the expected sub-speech that matches a first playback progress of the current sub-speech at a moment when the expected sub-speech is synthesized is determined. The above method determines the expected sub-speech based on the current sub-speech, to synthesize the sub-speech corresponding to the subsequent subtext of the second synthesized speech while the second synthesized speech is being played, reduce the synthesis time of the second synthesized speech, and increase the fluency of synthesized speech switching.

[0045] In a specific implementation scenario, a third audio duration of the current sub-speech is obtained, and a fourth audio duration of the expected sub-speech is obtained. A second playback progress of the expected sub-speech that matches a first playback progress of the current sub-speech at a moment when the expected sub-speech is synthesized is obtained based on a duration ratio between the fourth audio duration and the third audio duration and the first playback progress. The second playback progress can be referred to the signed embodiments, and is not described herein again for brevity.

[0046] Step S40: Switch to playing the second synthesized speech from the second playback progress of the second synthesized speech.

[0047] In one implementation scenario, the first synthesized speech is played based on the first configuration parameter, when the control instruction indicating re-synthesis of speech based on the second configuration parameter is detected, the playing of the first synthesized speech is kept, speech synthesis is performed on the text to be synthesized based on the second configuration parameter to obtain a second synthesized speech, and the second playing progress in the second synthesized speech that matches the first playing progress is determined based on the first playing progress of the first synthesized speech at the time when the second synthesized speech is synthesized, so that the second synthesized speech is switched to be played from the second playing progress. Therefore, during the synthesis of the second synthesized speech, the playing of the first synthesized speech is kept, and the second synthesized speech is switched to the second playing progress that matches the first playing progress based on the time when the second synthesized speech is synthesized, so that smooth switching of the synthesized speech is realized. Therefore, the fluency of switching of the synthesized speech is improved, and seamless connection of the synthesized speech is realized as much as possible.

[0048] In one implementation scenario, after the second synthesized speech is switched to be played, the subtext highlighted by the preset mark in the interactive interface matches the playing progress of the second synthesized speech. The related description of the highlighting of the preset mark can be referred to the foregoing embodiments, and will not be described herein again for brevity.

[0049] In one specific implementation scenario, the user can select part of the text to be synthesized displayed in the interactive interface, and switch to the second synthesized speech corresponding to the selected text to be synthesized to be played.

[0050] In one specific implementation scenario, in response to re-confirmation of the parameter configuration control by the user, the second synthesized speech currently played is taken as the first synthesized speech, and the steps of performing speech synthesis on the text to be synthesized based on the second configuration parameter to obtain the second synthesized speech and the subsequent steps in response to detection of the control instruction indicating re-synthesis of speech based on the second configuration parameter are returned to be executed, so that multiple parameter configurations on the text to be synthesized are realized.

[0051] In one specific implementation scenario, the interactive interface displays the second synthesized speech and a plurality of first synthesized speeches, and the configuration parameters corresponding to the synthesized speeches, so as to facilitate the user to view the historical synthesized speeches.

[0052] The above scheme plays the first synthesized speech synthesized based on the first configuration parameters for the synthesized text. When a control instruction representing re-speech synthesis based on the second configuration parameters is detected, the first synthesized speech continues to be played, and speech synthesis is performed on the synthesized text based on the second configuration parameters to obtain the second synthesized speech. Then, based on the first playback progress of the first synthesized speech at the moment when the synthesis of the second synthesized speech is completed, the second playback progress in the second synthesized speech that matches the first playback progress is determined, and the second synthesized speech can be switched to play starting from the second playback progress of the second synthesized speech. Therefore, during the synthesis of the second synthesized speech, the playback of the first synthesized speech is maintained, and based on the moment when the synthesis of the second synthesized speech is completed, the second playback progress in the second synthesized speech that matches the first playback progress is switched to achieve smooth switching of the synthesized speech. Therefore, the fluency of the switching of the synthesized speech can be improved, and the seamless playback of the switching of the synthesized speech can be achieved as much as possible.

[0053] See also Figure 2 , Figure 2 1 is a schematic diagram of a framework of an embodiment of the synthetic speech switching device 20 of the present application. Figure 2 As shown, the synthetic speech switching device 20 includes: a playback module 21, a synthesis module 22, a determination module 23 and a switching module 24, the playback module 21 is used to play the first synthetic speech synthesized based on the first configuration parameters for the text to be synthesized; the synthesis module 22 is used to respond to the detection of a control instruction representing the re-speech synthesis based on the second configuration parameters, perform speech synthesis on the text to be synthesized based on the second configuration parameters to obtain a second synthetic speech; the determination module 23 is used to determine the second playback progress of the second synthetic speech that matches the first playback progress based on the first playback progress of the first synthetic speech at the time when the synthesis of the second synthetic speech is completed; the switching module 24 is used to switch to playing the second synthetic speech starting from the second playback progress of the second synthetic speech.

[0054] The above scheme is that the synthesized speech switching device 20 plays the first synthesized speech synthesized based on the first configuration parameter, when the control instruction indicating that the speech synthesis based on the second configuration parameter is re-performed is detected, the playing of the first synthesized speech is continuously maintained, the speech synthesis based on the second configuration parameter is performed on the to-be-synthesized text, the second synthesized speech is obtained, and the second playing progress in the second synthesized speech that matches the first playing progress is determined based on the first playing progress of the first synthesized speech at the time when the second synthesized speech is synthesized, so that the second synthesized speech can be played starting from the second playing progress. Therefore, during the synthesis of the second synthesized speech, the playing of the first synthesized speech is maintained, and the second playing progress in the second synthesized speech that matches the first playing progress is switched to based on the time when the synthesis of the second synthesized speech is completed, so that the smooth switching of the synthesized speech is realized. Therefore, the fluency of the synthesized speech switching can be improved, and the seamless connection of the synthesized speech switching is realized as much as possible.

[0055] In some disclosed embodiments, the determining module 23 further includes a first progress matching module (not shown) configured to obtain a first audio duration of the first synthesized speech, and obtain a second audio duration of the second synthesized speech; and obtain the second playing progress that matches the first playing progress based on a duration ratio between the second audio duration and the first audio duration and the first playing progress.

[0056] In some disclosed embodiments, the synthesizing module 22 further includes a sub-speech synthesizing module (not shown) configured to sequentially perform speech synthesis on each sub-text based on the second configuration parameter to obtain a second sub-speech; wherein the second synthesized speech includes the second sub-speech; and the determining module 23 further includes a determining sub-module (not shown) configured to select the first sub-speech being played as a current sub-speech, and detect at least a synthesis progress of an expected sub-speech during the playing of the current sub-speech; wherein the expected sub-speech is the second sub-speech corresponding to the same sub-text as the current sub-speech; and in response to that at least the expected sub-speech has been synthesized, determine the second playing progress in the expected sub-speech that matches the first playing progress of the current sub-speech at the time when the expected sub-speech is synthesized.

[0057] In some disclosed embodiments, the sub-speech synthesizing module (not shown) further includes a subsequent synthesizing sub-module (not shown) configured to select a sub-text corresponding to a target sub-speech and a subsequent sub-text thereof as target sub-texts; wherein the target sub-speech is the first sub-speech played at the triggering time of the control instruction; and sequentially perform speech synthesis on each target sub-text based on the second configuration parameter to obtain a second sub-speech corresponding to each target sub-text; wherein the second synthesized speech includes at least the second sub-speech of each target sub-text.

[0058] In some disclosed embodiments, the sub-voice synthesis module (not shown) further comprises a historical synthesis sub-module (not shown) configured to sequentially perform voice synthesis on each historical sub-text based on the second configuration parameter to obtain a second sub-voice corresponding to the historical sub-text; wherein the historical sub-text is a sub-text corresponding to the first sub-voice that has been played at the triggering time of the control instruction; and the second sub-voice of the historical sub-text is sequentially added to the second sub-voice of the target sub-text to obtain the second synthesized voice; wherein the target sub-text is a sub-text corresponding to the first sub-voice that is being played and has not been played at the triggering time of the control instruction, and the historical sub-text is synthesized after the target sub-text.

[0059] In some disclosed embodiments, the determination sub-module (not shown) further comprises a second progress matching module (not shown) configured to obtain a third audio duration of the current sub-voice, and obtain a fourth audio duration of the expected sub-voice; and based on a duration ratio between the fourth audio duration and the third audio duration and a first playing progress of the current sub-voice at the completion time of the expected sub-voice synthesis, obtain a second playing progress of the expected sub-voice that matches the first playing progress.

[0060] In some disclosed embodiments, the synthesized voice switching device 20 further comprises an interaction module (not shown) configured to display an interaction interface; wherein the interaction interface displays the to-be-synthesized text, and highlights a sub-text in the to-be-synthesized text that matches the first synthesized voice playing progress with a preset marker, and the interaction interface is further provided with a parameter configuration control for triggering the control instruction.

[0061] In some disclosed embodiments, the synthesized voice switching device 20 further comprises a prompt module (not shown) configured to output a prompt message in response to the control instruction; wherein the prompt message is used to prompt that the second synthesized voice is being synthesized, and the prompt message stops outputting when the second synthesized voice synthesis is completed.

[0062] Please refer to Figure 3 , Figure 3 is a frame schematic diagram of an embodiment of the electronic device 30. The electronic device 30 comprises a memory 31 and a processor 32, the memory 31 stores program instructions, and the processor 32 is configured to execute the program instructions to implement the steps in any of the above synthesized voice switching method embodiments. For details, please refer to the foregoing disclosed embodiments, which will not be repeated here. The electronic device 30 can specifically include but is not limited to: a server, a smart phone, a notebook computer, a tablet computer, a self-service machine, etc., which are not limited here.

[0063] Specifically, the processor 32 is configured to control itself and the memory 31 to implement the steps in any of the above-described embodiments of the method for switching synthesized speech. The processor 32 can also be referred to as a CPU (Central Processing Unit). The processor 32 can be an integrated circuit chip having a processing capability of signals. The processor 32 can also be a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field-Programmable Gate Array) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor. In addition, the processor 32 can be implemented by an integrated circuit chip together.

[0064] The above scheme is that the electronic device 30 plays the first synthesized speech synthesized based on the first configuration parameter for the text to be synthesized, continues to maintain playing the first synthesized speech when detecting the control instruction representing re-synthesizing speech based on the second configuration parameter, synthesizes speech for the text to be synthesized based on the second configuration parameter to obtain the second synthesized speech, and determines the second playing progress in the second synthesized speech that matches the first playing progress based on the first playing progress of the first synthesized speech at the time when the second synthesized speech is synthesized, so as to switch to playing the second synthesized speech from the second playing progress of the second synthesized speech. Therefore, during the process of synthesizing the second synthesized speech, the playing of the first synthesized speech is maintained, and based on the time when the second synthesized speech is synthesized, the playing is switched to the second playing progress in the second synthesized speech that matches the first playing progress, so as to realize smooth switching of synthesized speech. Therefore, the fluency of switching of synthesized speech can be improved, and seamless connection of playing of switching of synthesized speech is realized as much as possible.

[0065] Please refer to Figure 4 , Figure 4 is a schematic diagram of an embodiment of the computer readable storage medium 40 of the present application. The computer readable storage medium 40 stores program instructions 41 capable of being executed by the processor, and the program instructions 41 are used to implement the steps in any of the above-described embodiments of the method for switching synthesized speech.

[0066] In some embodiments, the apparatus provided by the embodiments of the present disclosure has functions or includes modules that can be used to execute the methods described in the above method embodiments, and the specific implementation can refer to the description of the above method embodiments. For brevity, it will not be described here.

[0067] The above description of the various embodiments tends to emphasize differences between the various embodiments, and the same or similar elements can be referred to each other, and for brevity, will not be repeated here.

[0068] In several embodiments provided in the present application, it should be understood that the disclosed methods and devices can be implemented in other ways. For example, the above-described device implementation is only schematic, and for example, the division of the modules or units is only a logical function division, and actual implementation can have another division manner, for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the displayed or discussed elements can be indirect coupling or communication connection through some interfaces, devices or units, and can be electrical, mechanical or other forms.

[0069] The units described as separate components can or can not be physically separate, and the components shown as units can or can not be physical units, that is, can be located in one place, or can be distributed on a plurality of network units. Part or all of the units can be selected according to actual needs to achieve the purpose of the present embodiment scheme.

[0070] In addition, the functional units in each embodiment of the present application can be integrated in one processing unit, or each unit can be physically present separately, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of a software functional unit.

[0071] If the integrated unit is realized in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer readable storage medium. Based on this understanding, the technical scheme of the present application or the essential part or all or part of the technical scheme that contributes to the prior art can be embodied in the form of a software product, which is stored in a storage medium and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the various embodiment methods of the present application. The foregoing storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various program code storage media.

[0072] If the technical solution of the present application involves personal information, the product applying the technical solution of the present application has clearly informed the personal information processing rules before processing the personal information and obtained the personal independent consent. If the technical solution of the present application involves sensitive personal information, the product applying the technical solution of the present application has obtained the personal independent consent before processing the sensitive personal information and at the same time meets the requirement of "explicit consent". For example, at the personal information collection device such as camera, a clear and prominent mark is set to inform that it has entered the personal information collection range and will collect personal information. If the individual voluntarily enters the collection range, it is considered to agree to collect personal information. Or on the device for processing personal information, through the pop-up information or by asking the individual to upload his / her personal information, the individual's authorization is obtained under the condition of using obvious mark / information to inform the personal information processing rules. The personal information processing rules can include personal information processor, personal information processing purpose, processing method and personal information type, etc.

Claims

1. A synthetic speech switching method, characterized in that: include: Playing a first synthesized speech synthesized based on the first configuration parameter to be synthesized text; In response to detecting a control instruction indicating re-speech synthesis based on second configuration parameters, performing speech synthesis on the text to be synthesized based on the second configuration parameters to obtain a second synthesized speech; determining, based on a first playback progress of the first synthesized speech at a time when synthesis of the second synthesized speech is completed, a second playback progress of the second synthesized speech that matches the first playback progress; Starting from the second playback progress of the second synthesized speech, the playback of the second synthesized speech is switched.

2. The method according to claim 1, characterized in that The determining, based on the first playback progress of the first synthesized speech at the moment when synthesis of the second synthesized speech is completed, a second playback progress in the second synthesized speech that matches the first playback progress includes: Obtaining a first audio duration of the first synthesized speech, and obtaining a second audio duration of the second synthesized speech; Based on the ratio of the second audio duration to the first audio duration and the first playback progress, the second playback progress matching the first playback progress is obtained.

3. The method according to claim 1, characterized in that The text to be synthesized includes a plurality of subtexts, the first synthesized speech includes a first sub-speech of the plurality of subtexts, and the speech synthesis of the text to be synthesized based on the second configuration parameter to obtain the second synthesized speech includes: Perform speech synthesis on each of the sub-texts in sequence based on the second configuration parameters to obtain a second sub-speech; wherein the second synthesized speech includes the second sub-speech; The determining, based on the first playback progress of the first synthesized speech at the moment when synthesis of the second synthesized speech is completed, a second playback progress in the second synthesized speech that matches the first playback progress includes: Selecting the first sub-speech currently being played as the current sub-speech, and detecting at least the synthesis progress of a desired sub-speech during the playing of the current sub-speech; wherein the desired sub-speech is a second sub-speech corresponding to the same subtext as the current sub-speech; In response to at least the expected sub-speech being synthesized, a second playback progress in the expected sub-speech that matches the first playback progress is determined based on the first playback progress of the current sub-speech at the time when the expected sub-speech is synthesized.

4. The method according to claim 3, characterized in that The step of sequentially performing speech synthesis on each of the sub-texts based on the second configuration parameter to obtain a second sub-speech includes: Selecting a subtext corresponding to a target sub-speech and its subsequent sub-texts as target sub-texts; wherein the target sub-speech is the first sub-speech played at the triggering moment of the control instruction; Based on the second configuration parameters, speech synthesis is performed on each of the target subtexts in sequence to obtain a second sub-speech corresponding to the target subtext; wherein the second synthesized speech at least includes the second sub-speech of each of the target subtexts.

5. The method according to claim 3, characterized in that The step of sequentially performing speech synthesis on each of the sub-texts based on the second configuration parameter to obtain a second sub-speech includes: Perform speech synthesis on each historical subtext in sequence based on the second configuration parameter to obtain a second subtext corresponding to the historical subtext; wherein the historical subtext is the subtext corresponding to the first subtext that has been played at the time the control instruction is triggered; The second sub-speech of the historical sub-text is sequentially added before the second sub-speech of the target sub-text to obtain the second synthesized speech; wherein, the target sub-text is the sub-text corresponding to the first sub-speech that is being played and has not been played at the triggering moment of the control instruction, and the historical sub-text is speech synthesized after the target sub-text.

6. The method according to claim 3, characterized in that The determining, based on the first playback progress of the current sub-speech at the time when the synthesis of the expected sub-speech is completed, a second playback progress of the expected sub-speech that matches the first playback progress, includes: Obtaining a third audio duration of the current sub-speech, and obtaining a fourth audio duration of the expected sub-speech; Based on the ratio of the fourth audio duration to the third audio duration and the first playback progress of the current sub-speech at the time when the synthesis of the expected sub-speech is completed, the second playback progress of the expected sub-speech that matches the first playback progress is obtained.

7. The method according to claim 1, characterized in that The method further comprises: Display the interactive interface; Among them, the interactive interface displays the text to be synthesized, and highlights the sub-text in the text to be synthesized that matches the playback progress of the first synthesized voice with a preset mark. The interactive interface is also provided with a parameter configuration control for triggering the control instruction.

8. The method according to claim 1, characterized in that The method further comprises: In response to the control instruction, a prompt message is output; wherein the prompt message is used to prompt that the second synthesized speech is being synthesized, and the prompt message stops being output when the synthesis of the second synthesized speech is completed.

9. A synthetic speech switching device, characterized in that: include: A playback module, configured to play a first synthesized speech synthesized from the text to be synthesized based on the first configuration parameters; A synthesis module is configured to, in response to detecting a control instruction indicating re-speech synthesis based on second configuration parameters, perform speech synthesis on the text to be synthesized based on the second configuration parameters to obtain a second synthesized speech; A determination module, configured to determine a second playback progress of the second synthesized speech that matches the first playback progress based on the first playback progress of the first synthesized speech at the moment when synthesis of the second synthesized speech is completed; The switching module is used to switch to playing the second synthesized speech starting from the second playing progress of the second synthesized speech.

10. An electronic device, characterized in that: The method comprises at least a memory and a processor coupled to each other, wherein the memory stores program instructions, and the processor is used to execute the program instructions to implement the synthetic speech switching method according to any one of claims 1 to 8.

11. A computer-readable storage medium, characterized in that Program instructions that can be executed by a processor are stored, and the program instructions are used to implement the synthetic speech switching method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Reading mode switching method and apparatus, and storage medium

    CN107704437A

  • Multimedia interaction method, related device, equipment and storage medium

    CN111741370A