Real-time speech synthesis method and system
Through the method of automatically synthesizing audio in user voice in real time, the pause problem of real-time voice synthesis in the prior art is solved, and a wider application and better user experience is achieved.
Patent Information
- Application Number
- CN202211184894.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-27
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2042-09-27
AI Technical Summary
The existing real-time voice synthesis technology requires all possible speech audio to be configured in advance, resulting in pause problems during playback, which cannot meet all speech scenarios, especially user names and other scenarios, and the user's sensory experience is poor.
By configuring the real-time tts speech required for the current human-computer call task, obtain real-time tts resource information, establish index information, and recognize the user's voice in real time during the call, automatically synthesize audio, and play after synthesis, and report the usage status after completion to adjust resource scheduling.
减少了配置时间,扩大了应用场景,满足所有需要机器人实时播放的场景,提升了用户体验。
Smart Images

Figure CN115633126B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of intelligent conversational robots, and in particular to a real-time speech synthesis method and system. Background Art
[0002] Current real-time speech synthesis usually uses pre-synthesis TTS technology. Before making a call, the user needs to upload user-related information, use this information to synthesize audio in advance, and play the audio during the call. Variable information that can only be clear during the call is indirectly met by generating all possible words that the user may say in advance.
[0003] For example, the robot asks the user what the last four digits of his mobile phone number are. After the user says the last four digits, the robot will play the last four digits of the user's mobile phone number again to confirm with the user. At this time, you need to configure the audio corresponding to each digit in advance. During actual playback, split the digits said by the user into four digits, and play the audio corresponding to each digit in sequence.
[0004] For example, the robot asks the user when he will repay the loan. After the user says the date, the robot will play the date the user said again for confirmation. At this time, all possible corresponding date audios need to be configured in advance (for example, all audio information from 01-01 to 12-31).
[0005] However, existing technologies usually require that the audio corresponding to all possible words that the user may say be configured in advance, and then generated when playing. From the user's perspective, there will be pauses when the robot plays the audio, causing a poor user experience. In addition, it can only meet the scenario of being able to traverse all words, such as user names. Summary of the invention
[0006] In view of the above-mentioned shortcomings of the prior art, an object of the present invention is to provide a real-time speech synthesis method and system for solving the above technical problems in the prior art.
[0007] To achieve the above-mentioned purpose and other related purposes, the present invention provides a real-time speech synthesis method, which includes: configuring one or more real-time TTS scripts required for the current human-computer call task; obtaining the real-time TTS resource information corresponding to each real-time TTS script required to be scheduled for the current human-computer call task; establishing corresponding index information based on each configured real-time TTS script; based on the index information and real-time TTS resource information corresponding to each real-time TTS script, recognizing the user's voice respectively when performing the current human-computer call task, and automatically synthesizing the corresponding real-time TTS audio when the identification information corresponding to each index information is obtained, and playing the real-time TTS audio after the synthesis is completed; reporting the usage of each real-time TTS script of the current human-computer call task and the actual generation of each corresponding real-time TTS audio after the call task is completed, so as to adjust the scheduling of the real-time TTS resource information corresponding to each real-time TTS script.
[0008] In one embodiment of the present invention, the real-time TTS resource information includes: the name and number of one or more scheduling resources corresponding to the current real-time TTS session.
[0009] In one embodiment of the present invention, based on each index information, based on the index information corresponding to each real-time TTS speech and the real-time TTS resource information, the user voice is recognized respectively when the current human-computer call task is performed, and the corresponding real-time TTS audio is automatically synthesized when the identification information corresponding to each index information is obtained. The following includes: based on each index information, the user voice is recognized in real time when the current human-computer call task is performed, and the identification information corresponding to each index information is obtained; when the identification information is obtained, the corresponding scheduling resources are scheduled based on the corresponding real-time TTS resource information, and the corresponding real-time TTS audio is automatically synthesized according to the identification information and the corresponding real-time TTS speech.
[0010] In one embodiment of the present invention, when the identification information is obtained, the corresponding scheduling resources are scheduled based on the corresponding real-time TTS resource information to automatically synthesize the corresponding real-time TTS audio according to the identification information and the corresponding real-time TTS speech, including: whenever the identification information is obtained, the corresponding one or more scheduling resources are scheduled based on the corresponding real-time TTS resource information to obtain the corresponding one or more audio data according to the identification information, and based on the real-time TTS speech corresponding to the index information corresponding to the identification information, the audio data are synthesized to obtain the real-time TTS audio corresponding to the real-time TTS speech.
[0011] In one embodiment of the present invention, based on the constructed audio data synthesis model, one or more corresponding audio data are obtained according to the identification information, and based on the real-time TTS speech corresponding to the index information corresponding to the identification information, each audio data level is synthesized to obtain real-time TTS audio corresponding to the real-time TTS speech.
[0012] In one embodiment of the present invention, the reporting of the usage of each real-time TTS speech of the current human-computer call task and the actual generation of each corresponding real-time TTS audio after the call task is completed, so as to adjust the scheduling of real-time TTS resource information corresponding to each real-time TTS speech, includes: reporting the usage of each real-time TTS speech of the current human-computer call task and the actual generation of each corresponding real-time TTS audio after the call task is completed, so as to statistically analyze whether the scheduling resources corresponding to each real-time TTS speech are sufficient, so as to adjust the scheduling of real-time TTS resource information corresponding to each real-time TTS speech.
[0013] In one embodiment of the present invention, the adjustment of the scheduling of real-time TTS resource information corresponding to each real-time TTS speech includes: when the corresponding scheduling resources are insufficient, increasing the scheduling resources of the corresponding real-time TTS speech and adjusting the corresponding real-time TTS resource information.
[0014] In one embodiment of the present invention, the usage of each real-time TTS speech includes: the number of times each real-time TTS speech is used and the scheduling time of the scheduling resources corresponding to each real-time TTS speech; the actual generation of each real-time TTS audio includes: the audio length of each corresponding real-time TTS audio and the generation time of each real-time TTS audio.
[0015] In an embodiment of the present invention, each index information includes: one or more index target information and an index order.
[0016] To achieve the above-mentioned purpose and other related purposes, the present invention provides a real-time speech synthesis system, the system comprising: a TTS speech configuration module, used to configure one or more real-time TTS speech required for the current human-machine call task; a resource information acquisition module, connected to the TTS speech configuration module, to obtain the real-time TTS resource information corresponding to each real-time TTS speech required to be scheduled for the current human-machine call task; an index establishment module, connected to the resource information acquisition module, used to establish corresponding index information based on each configured real-time TTS speech; an audio synthesis playback module, connected to the index establishment module, used to The index information and real-time TTS resource information corresponding to the TTS speech are used to recognize the user's voice respectively when performing the current human-computer call task, and automatically synthesize the corresponding real-time TTS audio when the identification information corresponding to each index information is obtained, and play it after the real-time TTS audio is synthesized; a scheduling adjustment module is connected to the audio synthesis and playback module, and is used to report the usage of each real-time TTS speech of the current human-computer call task and the actual generation of each corresponding real-time TTS audio after the call task is completed, so as to adjust the scheduling of the real-time TTS resource information corresponding to each real-time TTS speech.
[0017] As described above, the present invention is a real-time speech synthesis method and system, which has the following beneficial effects: the present invention obtains the real-time TTS resource information corresponding to each real-time TTS speech required for the current human-machine call task by configuring one or more real-time TTS speech required for the current human-machine call task, and establishes corresponding index information based on the configured real-time TTS speech; based on the index information and real-time TTS resource information corresponding to each real-time TTS speech, the user voice is recognized respectively when the current human-machine call task is performed, and the corresponding real-time TTS audio is automatically synthesized when the identification information corresponding to each index information is obtained, and the real-time TTS audio is played after the synthesis of the real-time TTS audio; after the call task is completed, the usage of each real-time TTS speech of the current human-machine call task and the actual generation of each corresponding real-time TTS audio are reported, so as to adjust the scheduling of the real-time TTS resource information corresponding to each real-time TTS speech. The present invention not only reduces the configuration time, does not need to configure all audio data in advance, but also expands the application scenarios, and all scenarios requiring real-time playback of robots can be met. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 Shown is a flowchart of a real-time speech synthesis method in an embodiment of the present invention.
[0019] Figure 2 Shown is a schematic diagram of a real-time speech synthesis method in an embodiment of the present invention.
[0020] Figure 3Shown is a structural schematic diagram of a real-time speech synthesis system in an embodiment of the present invention. DETAILED DESCRIPTION
[0021] The following describes the embodiments of the present invention by specific examples, and those skilled in the art can easily understand other advantages and effects of the present invention from the contents disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the following embodiments and features in the embodiments can be combined with each other without conflict.
[0022] It should be noted that in the following description, reference is made to the accompanying drawings, which describe several embodiments of the present invention. It should be understood that other embodiments may also be used, and that mechanical composition, structural, electrical and operational changes may be made without departing from the spirit and scope of the present invention. The following detailed description should not be considered restrictive, and the scope of the embodiments of the present invention is limited only by the claims of the published patents. The terms used herein are only for describing specific embodiments and are not intended to limit the present invention. Spatially related terms, such as "upper", "lower", "left", "right", "below", "below", "lower", "above", "upper", etc., may be used in the text to facilitate the description of the relationship between an element or feature shown in the figure and another element or feature.
[0023] Throughout the specification, when a part is said to be "connected" to another part, this includes not only the case of "direct connection" but also the case of "indirect connection" by placing other elements therebetween. In addition, when a part is said to "include" a certain constituent element, unless otherwise stated, it does not exclude other constituent elements, but means that other constituent elements may be included.
[0024] The terms first, second and third mentioned herein are used to describe various parts, components, regions, layers and / or segments, but are not limited thereto. These terms are only used to distinguish a certain part, component, region, layer or segment from other parts, components, regions, layers or segments. Therefore, the first part, component, region, layer or segment described below may refer to the second part, component, region, layer or segment within the scope of the present invention.
[0025] Furthermore, as used herein, the singular forms "a", "an" and "the" are intended to include the plural forms as well, unless there is an indication to the contrary in the context. It should be further understood that the terms "comprise", "include" indicate the presence of the described features, operations, elements, components, items, kinds, and / or groups, but do not exclude the presence, occurrence or addition of one or more other features, operations, elements, components, items, kinds, and / or groups. The terms "or" and "and / or" used herein are interpreted as inclusive, or mean any one or any combination. Therefore, "A, B or C" or "A, B and / or C" means "any of the following: A; B; C; A and B; A and C; B and C; A, B and C". Exceptions to this definition will only occur when the combination of elements, functions or operations is inherently mutually exclusive in some way.
[0026] Therefore, the present invention provides a real-time speech synthesis method and system, by configuring one or more real-time TTS speech required for the current human-machine call task, obtaining the real-time TTS resource information corresponding to each real-time TTS speech required to be scheduled for the current human-machine call task, and establishing corresponding index information based on the configured real-time TTS speech; based on the index information corresponding to each real-time TTS speech and the real-time TTS resource information, the user voice is recognized when the current human-machine call task is performed, and the corresponding real-time TTS audio is automatically synthesized when the identification information corresponding to each index information is obtained, and the real-time TTS audio is played after the synthesis of the real-time TTS audio; after the call task is completed, the usage of each real-time TTS speech of the current human-machine call task and the actual generation of each corresponding real-time TTS audio are reported, so as to adjust the scheduling of the real-time TTS resource information corresponding to each real-time TTS speech. The present invention not only reduces the configuration time, does not need to configure all audio data in advance, but also expands the application scenarios, and all scenarios that require real-time playback of robots can be met.
[0027] The following is a detailed description of the embodiments of the present invention with reference to the accompanying drawings so that those skilled in the art can easily implement the present invention. The present invention can be embodied in many different forms and is not limited to the embodiments described herein.
[0028] like Figure 1 A flowchart showing a real-time speech synthesis method in an embodiment of the present invention is shown.
[0029] The method comprises:
[0030] Step S11: configure one or more real-time TTS scripts required for the current human-machine communication task.
[0031] Optionally, each real-time TTS speech is equivalent to an audio speech rule. When the robot needs to play the speech, it automatically synthesizes the audio according to the rule.
[0032] Step S12: Obtain the real-time TTS resource information corresponding to each real-time TTS speech required to be scheduled for the current human-machine call task.
[0033] Optionally, the real-time TTS resource information includes: the name and number of one or more scheduling resources corresponding to the current real-time TTS script, such as the name and number of the called server.
[0034] Preferably, before the robot makes a call, the synthetic real-time TTS resource information is arranged in advance according to the intelligent scheduling system to prevent insufficient resources when actually synthesizing the audio, causing the synthetic audio timeout problem.
[0035] Step S13: Establish corresponding index information based on each configured real-time TTS script.
[0036] Optionally, each index information includes: one or more index target information and an index order corresponding to each index target.
[0037] It should be noted that the index target information may be audio feature information of the index target, which is not limited to feature information such as location and time, and depends on the needs.
[0038] Before the call starts, the index information is calculated in advance based on the real-time TTS script. For example, if a real-time TTS script requires the last four digits of the ID card to be generated, then it is only necessary to obtain the last four digits of the user's ID card.
[0039] Step S14: Based on the index information corresponding to each real-time TTS speech and the real-time TTS resource information, the user voice is recognized respectively when the current human-computer call task is performed, and the corresponding real-time TTS audio is automatically synthesized when the identification information corresponding to each index information is obtained, and the real-time TTS audio is played after the synthesis is completed.
[0040] If the audio is generated directly during playback, it will take about 100 to 200 milliseconds to synthesize the TTS audio. If it is generated only during playback, from the user's perspective, there will be pauses when the robot plays the audio, causing a poor user experience. Therefore, the audio is not synthesized during playback, but played after the real-time TTS audio is synthesized.
[0041] Optionally, step S14 includes:
[0042] Based on each index information, the user's voice is recognized in real time when the current human-machine communication task is being performed, and recognition information corresponding to each index information is obtained; specifically, each index target is recognized in real time when the current human-machine communication task is being performed, and recognition information of the corresponding text type is recognized;
[0043] When the identification information is obtained, the corresponding scheduling resources are scheduled based on the corresponding real-time TTS resource information to automatically synthesize the corresponding real-time TTS audio according to the identification information and the corresponding real-time TTS speech.
[0044] Optionally, when the identification information is obtained, scheduling the corresponding scheduling resources based on the corresponding real-time TTS resource information to automatically synthesize the corresponding real-time TTS audio according to the identification information and the corresponding real-time TTS speech includes:
[0045] Whenever identification information is obtained, one or more corresponding scheduling resources are scheduled based on the corresponding real-time TTS resource information to obtain the corresponding one or more audio data according to the identification information, and each audio data is synthesized based on the real-time TTS speech corresponding to the index information corresponding to the identification information to obtain real-time TTS audio corresponding to the real-time TTS speech.
[0046] It should be noted that based on the real-time TTS speech, we can adjust the order of each audio data or generate a real-time TTS audio in a certain format according to the template, so that we can get the audio playback method we want according to our needs.
[0047] Optionally, based on a trained audio data synthesis model, one or more corresponding audio data are obtained according to the recognition information, and based on the real-time TTS speech corresponding to the index information corresponding to the recognition information, each audio data level is synthesized to obtain real-time TTS audio corresponding to the real-time TTS speech; therefore, there is no need to configure all the audio data in advance.
[0048] Step S15: After the call task is completed, the usage of each real-time TTS speech of the current human-machine call task and the actual generation of each corresponding real-time TTS audio are reported to adjust the scheduling of the real-time TTS resource information corresponding to each real-time TTS speech.
[0049] Optionally, step S15 includes: reporting the usage of each real-time TTS speech of the current human-machine call task and the actual generation of each corresponding real-time TTS audio after the call task is completed, so as to statistically analyze whether the scheduling resources corresponding to each real-time TTS speech are sufficient, so as to adjust the scheduling of the real-time TTS resource information corresponding to each real-time TTS speech. Provide intelligent scheduling system with more reasonable planning of TTS resources
[0050] Optionally, the adjustment of the scheduling of real-time TTS resource information corresponding to each real-time TTS speech includes: when the corresponding scheduling resources are insufficient, increasing the scheduling resources of the corresponding real-time TTS speech and adjusting the corresponding real-time TTS resource information.
[0051] Preferably, when TTS resources are insufficient, the synthesis time of TTS audio may be greatly delayed to more than 1 second. In this case, the scheme of early generation can accept a delay of more than 1 second, because the time from generation to actual playback may be much longer than 1 second.
[0052] Therefore, the usage of each real-time TTS speech includes: the number of uses of each real-time TTS speech and the scheduling time of the scheduling resources corresponding to each real-time TTS speech; the actual generation of each real-time TTS audio includes: the audio length of each corresponding real-time TTS audio and the generation time of each real-time TTS audio. It provides a more reasonable planning of TTS resources for the intelligent scheduling system.
[0053] In order to better describe the real-time speech synthesis method, the following specific embodiments are provided for illustration;
[0054] Embodiment 1: A real-time speech synthesis method; Figure 2 Schematic diagram of real-time speech synthesis method.
[0055] The method comprises:
[0056] 1. The robot's speech supports the configuration of real-time TTS speech. When the robot needs to play the speech, it automatically synthesizes the audio according to the speech.
[0057] 2. Before the robot calls, the real-time TTS resource information is arranged in advance according to the intelligent scheduling system to prevent insufficient resources during the actual audio synthesis, which causes the audio synthesis time to time out.
[0058] 3. Before the call starts, calculate the index information in advance based on the real-time TTS script. For example, if a real-time TTS script requires the last four digits of the ID card to generate, then as long as the last four digits of the user's ID card are obtained, the real-time TTS audio will be generated immediately, and the generated TTS audio will be played directly when it is needed.
[0059] in,
[0060] 3.1. If the audio is generated directly during playback, it will take about 100 to 200 milliseconds to synthesize the TTS audio. If it is generated only during playback, then from the user's perspective, there will be pauses when the robot plays the audio, causing a poor user experience.
[0061] 3.2. When TTS resources are insufficient, the time for synthesizing TTS audio may be greatly delayed, reaching more than 1 second. At this time, the scheme of early generation can accept a delay of more than 1 second, because the time from generation to actual playback may be much longer than 1 second.
[0062] 4. After the call is completed, the robot will report how many real-time TTS scripts it used and the time each script was generated. This will be used for statistical analysis and for the intelligent scheduling system to plan TTS resources more reasonably.
[0063] Similar to the principle of the above embodiment, the present invention provides a real-time speech synthesis system.
[0064] The following provides specific embodiments in conjunction with the accompanying drawings:
[0065] like Figure 3 A structural schematic diagram of a real-time speech synthesis system in an embodiment of the present invention is shown.
[0066] The system comprises:
[0067] A tts speech configuration module 31 is used to configure one or more real-time tts speech required for the current human-machine communication task;
[0068] The resource information acquisition module 32 is connected to the TTS speech configuration module 31 to obtain the real-time TTS resource information corresponding to each real-time TTS speech required to be scheduled for the current human-machine conversation task;
[0069] An index establishment module 33, connected to the resource information acquisition module 32, is used to establish corresponding index information based on each configured real-time TTS speech;
[0070] The audio synthesis and playback module 34 is connected to the index establishment module 33, and is used to recognize the user's voice respectively when performing the current human-machine call task based on the index information corresponding to each real-time TTS speech and the real-time TTS resource information, and automatically synthesize the corresponding real-time TTS audio when the identification information corresponding to each index information is obtained, and play the real-time TTS audio after the synthesis is completed;
[0071] The scheduling adjustment module 35 is connected to the audio synthesis and playback module 34, and is used to report the usage of each real-time TTS speech of the current human-computer call task and the actual generation of each corresponding real-time TTS audio after the call task is completed, so as to adjust the scheduling of the real-time TTS resource information corresponding to each real-time TTS speech.
[0072] It should be understood that Figure 3The division of each module in the system embodiment is only a division of logical functions. In actual implementation, they can be fully or partially integrated into one physical entity, or physically separated. And these units can all be implemented in the form of software calling through processing elements; they can also be all implemented in the form of hardware; some units can also be implemented in the form of processing elements calling software, and some units can be implemented in the form of hardware.
[0073] Since the implementation principle of the real-time speech synthesis system has been described in the above embodiments, it will not be repeated here.
[0074] Optionally, the real-time TTS resource information includes: the name and number of one or more scheduling resources corresponding to the current real-time TTS script.
[0075] Optionally, based on each index information, based on the index information corresponding to each real-time TTS speech and the real-time TTS resource information, the user voice is recognized respectively when the current human-computer call task is performed, and the corresponding real-time TTS audio is automatically synthesized when the identification information corresponding to each index information is obtained, and the real-time TTS audio is played after the synthesis is completed. It includes: based on each index information, the user voice is recognized in real time when the current human-computer call task is performed, and the identification information corresponding to each index information is obtained; when the identification information is obtained, the corresponding scheduling resources are scheduled based on the corresponding real-time TTS resource information, and the corresponding real-time TTS audio is automatically synthesized according to the identification information and the corresponding real-time TTS speech.
[0076] Optionally, when the identification information is obtained, the corresponding scheduling resources are scheduled based on the corresponding real-time TTS resource information to automatically synthesize the corresponding real-time TTS audio according to the identification information and the corresponding real-time TTS speech, including: whenever the identification information is obtained, the corresponding one or more scheduling resources are scheduled based on the corresponding real-time TTS resource information to obtain the corresponding one or more audio data according to the identification information, and based on the real-time TTS speech corresponding to the index information corresponding to the identification information, each audio data is synthesized to obtain the real-time TTS audio corresponding to the real-time TTS speech.
[0077] Optionally, based on the constructed audio data synthesis model, one or more corresponding audio data are obtained according to the identification information, and based on the real-time TTS speech corresponding to the index information corresponding to the identification information, each audio data level is synthesized to obtain real-time TTS audio corresponding to the real-time TTS speech.
[0078] Optionally, the reporting of the usage of each real-time TTS speech of the current human-computer call task and the actual generation of each corresponding real-time TTS audio after the call task is completed, so as to adjust the scheduling of real-time TTS resource information corresponding to each real-time TTS speech, includes: reporting the usage of each real-time TTS speech of the current human-computer call task and the actual generation of each corresponding real-time TTS audio after the call task is completed, so as to statistically analyze whether the scheduling resources corresponding to each real-time TTS speech are sufficient, so as to adjust the scheduling of real-time TTS resource information corresponding to each real-time TTS speech.
[0079] Optionally, the adjustment of the scheduling of real-time TTS resource information corresponding to each real-time TTS speech includes: when the corresponding scheduling resources are insufficient, increasing the scheduling resources of the corresponding real-time TTS speech and adjusting the corresponding real-time TTS resource information.
[0080] Optionally, the usage of each real-time TTS script includes: the number of times each real-time TTS script is used and the scheduling time of the scheduling resources corresponding to each real-time TTS script; the actual generation of each real-time TTS audio includes: the audio length of each corresponding real-time TTS audio and the generation time of each real-time TTS audio.
[0081] Optionally, each index information includes: one or more index target information and an index order.
[0082] In summary, the real-time speech synthesis system of the present invention obtains the real-time TTS resource information corresponding to each real-time TTS speech required for the current human-machine call task by configuring one or more real-time TTS speech techniques required for the current human-machine call task, and establishes corresponding index information based on the configured real-time TTS speech techniques; based on the index information corresponding to each real-time TTS speech technique and the real-time TTS resource information, the user voice is recognized respectively when the current human-machine call task is performed, and the corresponding real-time TTS audio is automatically synthesized when the identification information corresponding to each index information is obtained, and the real-time TTS audio is played after the synthesis of the real-time TTS audio; after the call task is completed, the usage of each real-time TTS speech technique of the current human-machine call task and the actual generation of each corresponding real-time TTS audio are reported, so as to adjust the scheduling of the real-time TTS resource information corresponding to each real-time TTS speech technique. The present invention not only reduces the configuration time, does not need to configure all the audio data in advance, but also expands the application scenarios, and all scenarios requiring real-time playback of robots can be met. Therefore, the present invention effectively overcomes the various shortcomings in the prior art and has a high industrial utilization value.
[0083] The above embodiments are merely illustrative of the principles and effects of the present invention, and are not intended to limit the present invention. Anyone familiar with the technology may modify or change the above embodiments without violating the spirit and scope of the present invention. Therefore, all equivalent modifications or changes made by a person of ordinary skill in the art without departing from the spirit and technical ideas disclosed by the present invention shall still be covered by the claims of the present invention.
Claims
1. A real-time speech synthesis method, It is characterized in that The method comprises: Before the robot calls, configure one or more real-time TTS scripts required for the current human-machine call task; Get the real-time TTS resource information corresponding to each real-time TTS speech required to be scheduled for the current human-machine call task; Establish corresponding index information based on each configured real-time TTS script; Based on the index information and real-time TTS resource information corresponding to each real-time TTS speech, the user voice is recognized respectively when the current human-computer call task is performed, and the corresponding real-time TTS audio is automatically synthesized when the identification information corresponding to each index information is obtained, and the real-time TTS audio is played after the real-time TTS audio is synthesized; the automatic synthesis includes the following steps: based on each index information, the user voice is recognized in real time when the current human-computer call task is performed, and the identification information corresponding to each index information is obtained; when the identification information is obtained, the corresponding scheduling resources are scheduled based on the corresponding real-time TTS resource information; according to the identification information and the corresponding real-time TTS speech, the corresponding real-time TTS audio is automatically synthesized; wherein, whenever the identification information is obtained, the corresponding one or more scheduling resources are scheduled based on the corresponding real-time TTS resource information; according to the identification information, the corresponding one or more audio data are obtained, and based on the real-time TTS speech corresponding to the index information corresponding to the identification information, each audio data is synthesized to obtain the real-time TTS audio corresponding to the real-time TTS speech; After the call task is completed, the usage of each real-time TTS script of the current human-computer call task and the actual generation of the corresponding real-time TTS audio are reported to adjust the scheduling of the real-time TTS resource information corresponding to each real-time TTS script.
2. According to the real-time speech synthesis method described in claim 1, It is characterized in that The real-time TTS resource information includes: the name and number of one or more scheduling resources corresponding to the current real-time TTS script.
3. According to the real-time speech synthesis method described in claim 2, It is characterized in that Based on the constructed audio data synthesis model, one or more corresponding audio data are obtained according to the recognition information, and based on the real-time TTS speech corresponding to the index information corresponding to the recognition information, each audio data is synthesized to obtain real-time TTS audio corresponding to the real-time TTS speech.
4. According to the real-time speech synthesis method described in claim 1, It is characterized in that The method of reporting the usage of each real-time TTS speech of the current human-machine conversation task and the actual generation of each corresponding real-time TTS audio after the conversation task is completed, so as to adjust the scheduling of the real-time TTS resource information corresponding to each real-time TTS speech, includes: After the call task is completed, the usage of each real-time TTS speech of the current human-computer call task and the actual generation of each corresponding real-time TTS audio are reported to statistically analyze whether the scheduling resources corresponding to each real-time TTS speech are sufficient, so as to adjust the scheduling of the real-time TTS resource information corresponding to each real-time TTS speech.
5. According to the real-time speech synthesis method described in claim 4, It is characterized in that The adjustment of the scheduling of the real-time TTS resource information corresponding to each real-time TTS speech includes: When the corresponding scheduling resources are insufficient, increase the scheduling resources for the corresponding real-time TTS speech and adjust the corresponding real-time TTS resource information.
6. The real-time speech synthesis method according to claim 4, It is characterized in that The usage of each real-time TTS script includes: the number of times each real-time TTS script is used and the scheduling time of the scheduling resources corresponding to each real-time TTS script; the actual generation of each real-time TTS audio includes: the audio length of each corresponding real-time TTS audio and the generation time of each real-time TTS audio.
7. According to the real-time speech synthesis method described in claim 1, It is characterized in that Each index information includes: one or more index target information and an index order.
8. A real-time speech synthesis system, It is characterized in that The system comprises: The tts speech configuration module is used to configure one or more real-time tts speech required for the current human-machine conversation task before the robot calls; A resource information acquisition module is connected to the TTS speech configuration module to obtain the real-time TTS resource information corresponding to each real-time TTS speech required to be scheduled for the current human-machine conversation task; An index establishment module, connected to the resource information acquisition module, for establishing corresponding index information based on each configured real-time TTS speech; The audio synthesis and playback module is connected to the index establishment module and is used to recognize the user's voice based on the index information corresponding to each real-time TTS speech and the real-time TTS resource information when performing the current human-computer call task, and automatically synthesize the corresponding real-time TTS audio when the identification information corresponding to each index information is obtained, and after synthesizing the real-time TTS The automatic synthesis includes the following steps: based on each index information, the user's voice is recognized in real time when the current human-computer call task is performed, and the identification information corresponding to each index information is obtained; when the identification information is obtained, the corresponding scheduling resources are scheduled based on the corresponding real-time TTS resource information; according to the identification information and the corresponding real-time TTS speech, the corresponding real-time TTS audio is automatically synthesized; wherein, whenever the identification information is obtained, the corresponding one or more scheduling resources are scheduled based on the corresponding real-time TTS resource information; according to the identification information, the corresponding one or more audio data are obtained, and based on the real-time TTS speech corresponding to the index information corresponding to the identification information, the audio data are synthesized to obtain the real-time TTS audio corresponding to the real-time TTS speech; The scheduling adjustment module is connected to the audio synthesis and playback module and is used to report the usage of each real-time TTS speech of the current human-computer call task and the actual generation of each corresponding real-time TTS audio after the call task is completed, so as to adjust the scheduling of the real-time TTS resource information corresponding to each real-time TTS speech.
Citation Information
Patent Citations
Multilingual intelligent voice conversation method and system
CN111128126A
Outbound voice output method, device and equipment
CN112735372A
Speech synthesis method and device, computer equipment and storage medium
CN113421549A