Voice acquisition method and device, server, client and storage medium
By sending streaming URL address information of speech synthesis in advance in the dialogue generation scenario and accumulating streaming text in parallel, the problems of inactivity, incoherence and high resource consumption of audio streaming in the prior art are solved, and more efficient voice acquisition and real-time experience are achieved.
Patent Information
- Application Number
- CN202311755625.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-19
- Publication Date
- 2025-06-20
AI Technical Summary
In the prior art, in the dialogue generation scenario, when the audio stream after speech synthesis is sent to the client, there are problems such as long time to obtain the first packet of audio data, incoherence of audio and excessive consumption of server resources.
When responding to the client's voice interaction request message on the server, the streaming URL address information corresponding to the target voice to be synthesized is sent in advance, so that the client can obtain continuous audio streams in real time. At the same time, the streaming text output from the large model can be accumulated and distributed in parallel, avoiding waiting for the entire text accumulation and audio synthesis to complete.
It saves time to obtain the first packet of audio data, improves the real-timeness of the client to obtain the target voice, avoids the problem of audio incoherence, and reduces the consumption of server resources.
Smart Images

Figure CN120183409A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of speech technology, and in particular, to a method, apparatus, server, client, and storage medium for obtaining speech. Background Art
[0002] With the development of deep learning and artificial intelligence technologies, large model technology has gradually become one of the research hotspots. Using large model technology (such as models like GPT-4, Meena, and DialoGPT, etc.) for dialogue generation has become the mainstream technology in the field of dialogue generation. By learning and training a large model with a large-scale corpus, the trained large model has powerful language understanding and generation capabilities, and can generate text responses in natural language with a certain degree of coherence and rationality according to the dialogue context and user input.
[0003] In the dialogue generation scenario, speech synthesis is an essential link. When using a large model for dialogue generation, the response text is usually very long and is in a streaming format. To improve the user experience during the dialogue generation process, it is necessary to smoothly and naturally send the audio converted from these streaming texts to the client. Summary of the Invention
[0004] To overcome the problems existing in the related art, the present disclosure provides a method, apparatus, server, client, and storage medium for obtaining speech.
[0005] According to the first aspect of the embodiments of the present disclosure, a method for obtaining speech is provided, which is applied to a server and includes:
[0006] In response to receiving a speech interaction request message sent by a client, sending address information corresponding to a target speech to be synthesized to the client, so that the client obtains the target speech according to the address information, and the target speech is used to respond to the speech interaction request message;
[0007] Obtaining the streaming text corresponding to the target speech;
[0008] Performing speech synthesis on the streaming text to obtain the target speech.
[0009] Optionally, the performing speech synthesis on the streaming text to obtain the target speech includes:
[0010] In response to receiving a speech download request message sent by the client according to the address information, performing speech synthesis on the streaming text to obtain the target speech.
[0011] Optionally, the streaming text includes a text sequence composed of multiple first text units, and the target speech includes an audio stream;
[0012] Performing speech synthesis on the streaming text to obtain the target speech includes:
[0013] Repeatedly execute the speech synthesis step until the text sequence is fully traversed;
[0014] The speech synthesis step includes:
[0015] Traverse the text sequence in turn, and according to the preset text division rule, sequentially merge one or more first text units obtained by traversal to obtain a second text unit;
[0016] Perform speech synthesis on the second text unit to obtain the audio stream;
[0017] When the first text units that have not been traversed in the text sequence are not empty, traverse the first text units that have not been traversed and obtain an updated second text unit according to the preset text division rule.
[0018] Optionally, performing speech synthesis on the second text unit to obtain the audio stream includes:
[0019] Perform streaming audio synthesis on the second text unit through a preset text-to-speech (TTS) speech synthesis model to obtain the audio stream.
[0020] Optionally, obtaining the streaming text corresponding to the target speech includes:
[0021] Obtain the streaming text corresponding to the target speech according to the speech interaction request message.
[0022] Optionally, obtaining the streaming text corresponding to the target speech according to the speech interaction request message includes:
[0023] Obtain the request text corresponding to the speech interaction request message;
[0024] Output the streaming text according to the request text through a preset dialogue generation model.
[0025] Optionally, the method further includes:
[0026] Send the target speech to the client.
[0027] According to the second aspect of the embodiments of the present disclosure, there is provided a speech acquisition method, which is applied to a client, and the method includes:
[0028] Send a speech interaction request message to the server;
[0029] Receive the address information sent by the server, where the address information is the address information of the target voice to be synthesized sent by the server in response to receiving the voice interaction request message, and the target voice is used to respond to the voice interaction request message;
[0030] Obtain the target voice from the server according to the address information.
[0031] Optionally, the obtaining the target voice from the server according to the address information includes:
[0032] Send a voice download request message to the server according to the address information, so that the server synthesizes the voice of the streaming text corresponding to the target voice after receiving the voice download request message, and obtains the target voice;
[0033] Receive the target voice sent by the server.
[0034] According to the third aspect of the embodiments of the present disclosure, there is provided a voice acquisition device, which is applied to a server and includes:
[0035] A first sending module, configured to send the address information corresponding to the target voice to be synthesized to the client in response to receiving the voice interaction request message sent by the client, so that the client obtains the target voice according to the address information, and the target voice is used to respond to the voice interaction request message;
[0036] An acquisition module, configured to acquire the streaming text corresponding to the target voice;
[0037] A voice synthesis module, configured to perform voice synthesis on the streaming text to obtain the target voice.
[0038] According to the fourth aspect of the embodiments of the present disclosure, there is provided a voice acquisition device, which is applied to a client and includes:
[0039] A second sending module, configured to send a voice interaction request message to the server;
[0040] A first receiving module, configured to receive the address information sent by the server, where the address information is the address information of the target voice to be synthesized sent by the server in response to receiving the voice interaction request message, and the target voice is used to respond to the voice interaction request message;
[0041] A voice acquisition module, configured to obtain the target voice from the server according to the address information.
[0042] According to the fifth aspect of the embodiments of the present disclosure, there is provided a server, including:
[0043] A processor;
[0044] A memory for storing instructions executable by the processor;
[0045] Wherein, the processor is configured to: execute the steps of the method described in the first aspect of the present disclosure.
[0046] According to a sixth aspect of the embodiments of the present disclosure, a client is provided, including:
[0047] A processor;
[0048] A memory for storing instructions executable by the processor;
[0049] Wherein, the processor is configured to: execute the steps of the method described in the second aspect of the present disclosure.
[0050] According to a seventh aspect of the embodiments of the present disclosure, a computer-readable storage medium is provided, on which computer program instructions are stored, and when the program instructions are executed by a processor, the steps of the voice acquisition method provided in the first aspect or the second aspect of the present disclosure are implemented.
[0051] The technical solutions provided by the embodiments of the present disclosure may include the following beneficial effects: In response to receiving the voice interaction request message sent by the client, the server sends the address information corresponding to the target voice to be synthesized to the client, so that the address information can be sent in advance without relying on the accumulation of the streaming text output by the large model and the completion of audio synthesis. Since the large model may be outputting the streaming text corresponding to the target voice while the server sends the address information to the client, this realizes the parallelization of the two processes of sending the address information and accumulating the streaming text, thereby saving the time for obtaining the first packet of audio data and improving the real-time performance of the client to obtain the target voice.
[0052] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] The accompanying drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present disclosure and used together with the specification to explain the principles of the present disclosure.
[0054] Figure 1 FIG. is a schematic diagram of a process for implementing streaming text audio synthesis based on a non-streaming URL provided in the related art.
[0055] Figure 2 FIG. is a flowchart of a voice acquisition method shown according to an exemplary embodiment.
[0056] Figure 3 is a flowchart of a voice acquisition method shown according to Figure 1 the embodiment shown.
[0057] Figure 4 is a flowchart of a voice acquisition method shown according to Figure 2 the embodiment shown.
[0058] Figure 5 is a flowchart of a voice acquisition method shown according to Figure 2 the embodiment shown.
[0059] Figure 6 is a schematic diagram of a process for implementing streaming text audio synthesis based on a streaming URL shown according to an exemplary embodiment.
[0060] Figure 7 is a flowchart of a voice acquisition method shown according to an exemplary embodiment.
[0061] Figure 8 is a block diagram of a voice acquisition device shown according to an exemplary new embodiment.
[0062] Figure 9 is a block diagram of a voice acquisition device shown according to Figure 8 the embodiment shown.
[0063] Figure 10 is a block diagram of a voice acquisition device shown according to an exemplary embodiment.
[0064] Figure 11 is a block diagram of a device for voice acquisition shown according to an exemplary embodiment.
[0065] Figure 12 is a block diagram of a device for voice acquisition shown according to an exemplary embodiment. Detailed implementation manners
[0066] Here, the exemplary embodiments will be described in detail, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present disclosure. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.
[0067] It should be noted that all actions of obtaining signals, information, or data in the present disclosure are carried out on the premise of complying with the corresponding data protection regulations and policies of the country where it is located and obtaining the authorization given by the owner of the corresponding device.
[0068] The present disclosure is mainly applied to speech synthesis in dialogue generation scenarios. The audio stream after speech synthesis is generally sent to the client in two ways: binary audio stream and URL (Uniform Resource Locator).
[0069] Among them, binary audio stream refers to the transmission of speech synthesized audio data to the client in binary form, and the client can directly play this audio stream. The advantages of this method are fast transmission speed, high real-time and stability of audio data, and it is suitable for scenarios with high real-time requirements, such as real-time voice calls, real-time voice interactions, etc. However, there are two disadvantages: First, the storage space occupancy is high: the binary audio stream needs to be generated in real time during the transmission process, and needs to occupy the corresponding storage space, which will cause a certain burden for devices with limited storage space (such as speakers, etc.); second, it cannot be easily interrupted: the binary audio stream is transmitted once, and once the transmission starts, it needs to be transmitted until the end, otherwise the audio playback will be incoherent or unable to play. The URL method refers to storing the speech synthesized audio data on the server, and returning the URL link to access the audio data to the client, and the client can download the audio data and play it through the URL link. The advantage of this method is that the audio data can be downloaded and played at any time, and the user can interrupt or switch the audio at any time, which has better flexibility. Therefore, in some application scenarios, the URL method may be more suitable, such as cockpit voice navigation, voice books and other applications that need to switch audio at any time in the middle. In scenarios such as large-model floor-standing speakers and cockpits, due to device storage space limitations and the possibility of interrupting the transmission of the audio stream at any time, in dialogue generation scenarios in the fields of voice assistants and smart cockpits, the URL method will be preferred to implement the delivery of the audio stream for speech synthesis.
[0070] In the process of sending audio streams based on URLs, since the large model generates text in a streaming manner, it needs to be segmented and transmitted continuously. If a non-streaming URL is used for sending, it is necessary to wait for the server to process the entire text before returning the audio file, so the delay is large. In the related art, the streaming text generated by the large model can be divided into multiple parts, each of which generates a non-streaming URL, thereby solving the problem of long time to generate audio streams.
[0071] Figure 1 It is a schematic diagram of a process of realizing streaming text and audio synthesis based on a non-streaming URL provided in the related art, such as Figure 1As shown below, it includes the following steps: Step 1: The large model performs continuous streaming text generation (for example, the streaming text output by the large model is: text1 = "Daytime", text2 = "mountains lean", text3 = "Yellow River", text4 = "flows into the sea", text5 = "flows"). After the voice control module obtains the streaming text output by the large model, it accumulates and cuts the text based on a preset text accumulation rule. For example, when accumulating to 10 characters, these 10 characters are used as an accumulated text (such as the accumulated text 1 or accumulated text 2 in Figure 1 ), and it is transmitted to the voice synthesis module for audio synthesis, or when encountering a preset punctuation mark, the accumulated text is requested for audio synthesis once. The voice synthesis module returns the synthesized audio corresponding to the input accumulated text (such as the synthesized audio 1 or synthesized audio 2 in Figure 1 ) to the voice control module. After the voice synthesis module obtains the complete synthesized audio corresponding to the accumulated text, it stores it in the cloud audio storage medium and generates a URL corresponding to the currently stored synthesized audio. Here, the URL corresponds one-to-one with the accumulated text. The voice synthesis module can send multiple non-streaming URLs obtained to the client, and each time the client receives a non-streaming URL, it will throw it to the player in sequence. When the player plays the URL, it will first establish a connection with the cloud through the URL, and then obtain the audio stored in the cloud. The client will piece together the audio corresponding to these URLs into a complete audio file for playback.
[0072] However Figure 1 For the method of realizing streaming text audio synthesis based on non-streaming URLs shown below, the streaming text generated by the large model is cut into multiple paragraphs or sentences, and a corresponding URL is generated for each paragraph or sentence. When the client pieces together the audio corresponding to these URLs into a complete audio file for playback, the following problems will exist:
[0073] 1. The time for the user to obtain the first package of audio from the large model is relatively long: Although a single accumulated large model text is shorter than the complete streaming text, it still needs to wait until the voice synthesis module finishes synthesizing all the accumulated single texts and stores them in the audio storage medium before generating the corresponding URL. This process is executed serially, and the time-consuming is positively correlated with the text length.
[0074] 2. The audio splicing may be incoherent: Since the audio for each paragraph or sentence is generated independently, there may be a problem of incoherent audio when the client plays and splices them, affecting the user experience.
[0075] 3. The number of URLs may be large, which increases the resource consumption and latency of establishing links: A network connection needs to be established for each URL corresponding to the accumulated text, and the process of establishing a network connection usually consumes a certain amount of time and resources. If the number of generated URLs is too large, the server will need to handle a large number of network connection requests simultaneously, easily resulting in connection timeouts or connection failures. This will lead to a decrease in the availability of the system and also affect the user experience.
[0076] To solve the above problems, the present disclosure provides a voice acquisition method, device, server, client, and storage medium. The following will describe the specific embodiments of the present disclosure in detail with reference to the accompanying drawings.
[0077] Figure 2 is a flowchart of a voice acquisition method shown according to an exemplary embodiment, and this method can be applied to a server. As Figure 2 shown, it includes the following steps.
[0078] In step S201, in response to receiving a voice interaction request message sent by the client, send the address information corresponding to the target voice to be synthesized to the client, so that the client can obtain the target voice according to the address information, and the target voice is used to respond to the voice interaction request message.
[0079] Among them, the voice interaction request message can be a voice request message sent by the user through the client. For example, based on the activation of the voice assistant function on the client, the client collects a voice interaction request message "Please play a poem" sent by the user, and the client uploads the voice interaction request message to the server so that the server can respond to the voice interaction request message.
[0080] The address information may include the URL address corresponding to the target voice, and the URL address is a streaming URL. Through this streaming URL, the client can obtain the audio stream of the target voice that is continuous and transmitted in chronological order in real time. A non-streaming URL usually refers to a traditional URL link. Usually, the voice data obtained based on a non-streaming URL is the voice data that has been pre-audio synthesized by the server. Based on a non-streaming URL, it is not possible to obtain the audio stream data in real time, and it is necessary to wait until all the audio is obtained and then send it to the client at one time.
[0081] In this step, in response to receiving the voice interaction request message sent by the client, the server immediately sends the address information corresponding to the target voice to be synthesized to the client, avoiding the lag operation of only being able to send the URL after the streaming text output by the large model is accumulated and the audio is synthesized, and saving the time for the client to obtain the synthesized audio.
[0082] In step S202, obtain the streaming text corresponding to the target voice.
[0083] Among them, the streaming text can be understood as text continuously output in chronological order, and this streaming text can include a text sequence composed of multiple first text units. For example, the streaming text output by the large model can be: text1 = "Daytime", text2 = "Mountains bend", text3 = "Yellow River", text4 = "Flows into the sea", text5 = "Flows". text1, text2, text3, etc. are all the first text units.
[0084] In addition, this streaming text is the response text of the voice interaction request message sent by the user.
[0085] In this step, the streaming text corresponding to the target voice can be obtained according to the voice interaction request message.
[0086] In one implementation, the request text corresponding to the voice interaction request message can be obtained; the streaming text is output according to the request text through a preset dialogue generation model. Among them, this preset dialogue generation model is a pre-set large model, and this large model can be model-trained through a large-scale corpus. After training, the large model has powerful language understanding and generation capabilities, and can generate a streaming text with a certain degree of coherence and rationality according to the dialogue context and user input.
[0087] Exemplarily, assume that the voice interaction request message sent by the user is "Please play a poem". After the server receives this voice interaction request message, based on speech recognition technology, the voice information can be converted into text data (i.e., the request text "Please play a poem"), and then this request text is input into the preset dialogue generation model obtained by pre-training. After that, the response text of "Please play a poem" output by the preset dialogue generation model is "text1 = 'Daytime', text2 = 'Mountains bend', text3 = 'Yellow River', text4 = 'Flows into the sea', text5 = 'Flows'". The above example is only for illustration, and the present disclosure is not limited thereto.
[0088] In step S203, perform speech synthesis on the streaming text to obtain the target voice.
[0089] Among them, the target voice can include an audio stream continuously output in chronological order.
[0090] In a possible implementation of this step, the target voice can be obtained after performing speech synthesis on the streaming text through a preset TTS (Text To Speech) speech synthesis model.
[0091] Using the above method, upon receiving the voice interaction request message sent by the client, the server sends the address information corresponding to the target voice to be synthesized to the client, enabling the address information to be sent in advance without relying on the accumulation of the streaming text output by the large model and the completion of audio synthesis. Since the large model may be outputting the streaming text corresponding to the target voice while the server sends the address information to the client, this realizes the parallelization of the two processes of sending the address information and accumulating the streaming text, thereby saving the time to obtain the first packet of audio data and improving the real-time performance of the client to obtain the target voice.
[0092] In addition, the client can obtain the target voice based on this address information (i.e., the streaming URL), thus avoiding the problem of audio incoherence caused by splicing the audio of multiple URLs on the client. Moreover, the client uses only one address information to obtain the target voice, which reduces the unnecessary consumption of link resources and latency compared with using multiple URLs, and also avoids the problem of reduced availability of the server due to simultaneously processing a large number of network connection requests, improving the user experience.
[0093] Figure 3 is based on Figure 1 The flowchart of a voice acquisition method shown in the illustrated embodiment is as Figure 3 shown, and step S203 includes the following sub-steps:
[0094] In step S2031, in response to receiving the voice download request message sent by the client according to the address information, perform voice synthesis on the streaming text to obtain the target voice.
[0095] Wherein, the voice download request message includes the address information.
[0096] After the client obtains the address information sent by the server, it will be handed over to the player for playback. The player can establish a connection with the server by accessing the address information and wait for the server to send the streaming target voice for playback. Among them, during the process of the client establishing a connection with the server through the address information, it can generate and send the voice download request message to the server.
[0097] After the server receives the voice download request message, it can start performing voice synthesis on the streaming text.
[0098] Figure 4 is based on Figure 2 The flowchart of a voice acquisition method shown in the illustrated embodiment is as Figure 4 shown, and step S203 includes the following sub-steps:
[0099] In step S2032, the speech synthesis step is repeatedly executed until the text sequence is completely traversed. The speech synthesis step includes: sequentially traversing the text sequence, and sequentially combining one or more first text units obtained by traversal according to a preset text division rule to obtain a second text unit; performing speech synthesis on the second text unit to obtain the audio stream; in the case where the first text units that have not been traversed in the text sequence are not empty, traversing the first text units that have not been traversed, and obtaining an updated second text unit according to the preset text division rule.
[0100] As described above, while the server sends a streaming URL to the client, the large model may be returning streaming text to the server. The server can store the obtained streaming text in a preset message queue. It can be understood that the streaming text is stored in the preset message queue in the form of a text sequence.
[0101] After receiving the voice download request message, the server can start performing speech synthesis on the streaming text. In an implementation manner of the present disclosure, the text sequence of the streaming text stored in the preset message queue can be sequentially traversed, and one or more first text units obtained by traversal are sequentially combined according to a preset text division rule to obtain a second text unit. For example, the preset text division rule can be to perform text merging once every 10 characters accumulated to obtain the second text unit. Alternatively, the preset text division rule can be to sequentially combine one or more first text units obtained by traversal according to a preset punctuation mark to obtain a second text unit.
[0102] After obtaining a second text unit through one text accumulation, speech synthesis can be performed on the second text unit to obtain an audio stream; here, the second text unit can be subjected to streaming audio synthesis through a preset TTS speech synthesis model to obtain the audio stream.
[0103] It should be noted that during the process of performing streaming audio synthesis on the second text unit through the preset TTS speech synthesis model, the preset TTS speech synthesis model can continuously output streaming audio data (i.e., the audio stream), so that the audio stream after speech synthesis can be transmitted to the client in real time, without waiting until the preset TTS speech synthesis model has completely synthesized the current second text unit and then transmitting the synthesized audio to the client. Thus, the first-pack waiting time for the user to obtain the streaming TTS audio of the large model can be greatly reduced, and the real-time performance of the client to obtain the target voice can be further improved.
[0104] Figure 5 is based on Figure 2 The flowchart of a voice acquisition method shown in the illustrated embodiment, as Figure 5As shown, the method further comprises the following steps:
[0105] In step S204, the target voice is sent to the client.
[0106] In response to receiving the voice download request message sent by the client according to the address information, the server performs voice synthesis on the streaming text to obtain the target voice, and then transmits the target voice to the client in real time in the form of an audio stream.
[0107] As mentioned above, the address information in the present disclosure includes a streaming URL.
[0108] For example, Figure 6 FIG. 1 is a schematic diagram showing a process of implementing streaming text and audio synthesis based on a streaming URL according to an exemplary embodiment. Figure 6 As shown, the streaming text audio synthesis based on the streaming URL includes the following steps: S1: The voice control module (deployed to the server) sends the received voice interaction request message uploaded by the client to the big model, so as to request the big model to generate the response text (i.e., streaming text) of the voice interaction request message. S2: The big model continuously generates the streaming text, and the voice control module can store the streaming text obtained by the server to the message queue of the streaming text. At the same time, in response to receiving the voice interaction request message uploaded by the client, the voice control module on the server can send the streaming URL to the client (the streaming URL is used for the client to download the synthesized target voice from the server), so that the step of accumulating the streaming text and sending the URL to the client can be asynchronous. S3: After the client receives the streaming URL, it initiates a voice download request message through the streaming URL. S4: In response to receiving the voice download request message sent by the client, the voice control module reads the streaming text from the message queue of the streaming text in sequence (such as cutting the streaming text according to punctuation marks), and the cut text (i.e., the second text unit) is sent to the speech synthesis module for speech synthesis. S5: The speech synthesis module returns the synthesized audio stream, which can be continuously transmitted to the client, so that the speech synthesized audio stream can be transmitted to the client in real time, without waiting until the speech synthesis module completes the audio synthesis of the current second text unit before transmitting the synthesized audio to the client, thereby greatly reducing the waiting time for the user to obtain the first packet of the large model streaming audio, and further improving the real-time performance of the client to obtain the target speech. The above examples are only for illustration, and the present disclosure does not limit this.
[0109] Figure 7 is a flow chart of a method for acquiring voice according to an exemplary embodiment. The method can be applied to a client, such as Figure 7 As shown, the method comprises the following steps:
[0110] In step S701, a voice interaction request message is sent to the server.
[0111] Among them, the voice interaction request message can be a voice request message sent by the user through the client. For example, based on the activation of the voice assistant function on the client, the client collects a voice interaction request message "Please play a poem" sent by the user, and the client uploads the voice interaction request message to the server so that the server can respond to the voice interaction request message.
[0112] In step S702, address information sent by the server is received. The address information is the address information of the target voice to be synthesized sent by the server in response to receiving the voice interaction request message. The target voice is used to respond to the voice interaction request message.
[0113] Among them, the address information can include the URL address corresponding to the target voice, and the URL address is a streaming URL. Through this streaming URL, the client can obtain the audio stream of the target voice that is continuous and transmitted in chronological order in real time. A non-streaming URL usually refers to a traditional URL link. Usually, the voice data obtained based on a non-streaming URL is the voice data after the server pre-performs audio synthesis. Based on a non-streaming URL, it is not possible to obtain the audio stream data in real time, and it is necessary to wait until all the audio is obtained and then send it to the client all at once.
[0114] In response to receiving the voice interaction request message sent by the client, the server immediately sends the address information of the target voice to be synthesized to the client, avoiding the lag operation of only being able to send the URL after the streaming text output by the large model has been accumulated and the audio has been synthesized, saving the time for the client to obtain the synthesized audio.
[0115] In step S703, the target voice is obtained from the server according to the address information.
[0116] In this step, a voice download request message can be sent to the server according to the address information, so that in response to receiving the voice download request message, the server performs voice synthesis on the streaming text corresponding to the target voice to obtain the target voice; and receives the target voice sent by the server.
[0117] Among them, the streaming text can be understood as text continuously output in chronological order, and the streaming text can include a text sequence composed of multiple first text units. In addition, the streaming text is the response text of the voice interaction request message sent by the user. The target voice can include an audio stream continuously output in chronological order.
[0118] Using the above method, after the client sends a voice interaction request message to the server, it can receive the address information of the target voice to be synthesized sent by the server in response to receiving the voice interaction request message sent by the client, so that the address information can be sent in advance without relying on the accumulation of the streaming text output by the large model and the completion of audio synthesis. Since the large model may be outputting the streaming text corresponding to the target voice while the server sends the address information to the client, this realizes the parallelization of the two processes of sending the address information and accumulating the streaming text, thereby saving the time for obtaining the first packet of audio data and improving the real-time performance of the client to obtain the target voice.
[0119] In addition, the client can obtain the target voice according to this address information (i.e., the streaming URL), thus avoiding the problem of audio incoherence caused by splicing the audio of multiple URLs on the client. Moreover, the client only uses one address information to obtain the target voice. Compared with using multiple URLs, it reduces unnecessary link resource consumption and latency, and also avoids the problem of reduced availability of the server due to simultaneously processing a large number of network connection requests, improving the user experience.
[0120] Figure 8 It is a block diagram of a voice acquisition device shown according to an exemplary new embodiment, applied to a server, as Figure 8 shown, the device includes:
[0121] A first sending module 801, configured to respond to receiving a voice interaction request message sent by a client, and send address information corresponding to a target voice to be synthesized to the client, so that the client obtains the target voice according to the address information, and the target voice is used to respond to the voice interaction request message;
[0122] An acquisition module 802, configured to acquire the streaming text corresponding to the target voice;
[0123] A voice synthesis module 803, configured to perform voice synthesis on the streaming text to obtain the target voice.
[0124] Optionally, the voice synthesis module 803 is configured to respond to receiving a voice download request message sent by the client according to the address information, and perform voice synthesis on the streaming text to obtain the target voice.
[0125] Optionally, the streaming text includes a text sequence composed of multiple first text units, and the target voice includes an audio stream;
[0126] The voice synthesis module 803 is configured to repeatedly execute the voice synthesis step until the text sequence is traversed;
[0127] The voice synthesis steps include:
[0128] Traverse the text sequence in turn, and according to the preset text division rule, sequentially merge one or more first text units traversed to obtain a second text unit;
[0129] Perform voice synthesis on the second text unit to obtain the audio stream;
[0130] In the case that the first text units not traversed in the text sequence are not empty, traverse the first text units not traversed, and obtain an updated second text unit according to the preset text division rule.
[0131] Optionally, the voice synthesis module 803 is configured to perform streaming audio synthesis on the second text unit through a preset text-to-speech (TTS) voice synthesis model to obtain the audio stream.
[0132] Optionally, the acquisition module 802 is configured to obtain the streaming text corresponding to the target voice according to the voice interaction request message.
[0133] Optionally, the acquisition module 802 is configured to obtain the request text corresponding to the voice interaction request message; and output the streaming text according to the request text through a preset dialogue generation model.
[0134] Optionally, Figure 9 is a block diagram of a voice acquisition device shown according to the embodiments shown, as Figure 8 shown, the device further includes: Figure 9 shown, the device further includes:
[0135] A third sending module 804, configured to send the target voice to the client.
[0136] Figure 10 is a block diagram of a voice acquisition device shown according to an exemplary embodiment, applied to a client, as Figure 10 shown, the device includes:
[0137] A second sending module 1001, configured to send a voice interaction request message to the server;
[0138] A first receiving module 1002, configured to receive address information sent by the server, where the address information is the address information corresponding to the target voice to be synthesized sent by the server in response to receiving the voice interaction request message, and the target voice is used to respond to the voice interaction request message;
[0139] The voice acquisition module 1003 is configured to acquire the target voice from the server according to the address information.
[0140] Optionally, the voice acquisition module 1003 is configured to send a voice download request message to the server according to the address information, so that the server synthesizes the voice of the streaming text corresponding to the target voice after receiving the voice download request message, and then obtains the target voice; and receive the target voice sent by the server.
[0141] Regarding the device in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated herein.
[0142] The present disclosure also provides a computer-readable storage medium, on which computer program instructions are stored, and when the program instructions are executed by a processor, the steps of the voice acquisition method provided by the present disclosure are implemented.
[0143] Figure 11 is a block diagram of a device for voice acquisition shown according to an exemplary embodiment. For example, the device 1100 may be provided as a server. Refer to Figure 11 , the device 1100 includes a processing component 1122, which further includes one or more processors, and memory resources represented by a memory 1132 for storing instructions executable by the processing component 1122, such as application programs. The application programs stored in the memory 1132 may include one or more modules each corresponding to a set of instructions. In addition, the processing component 1122 is configured to execute instructions to perform the above-mentioned voice acquisition method.
[0144] The device 1100 may further include a power supply component 1126 configured to perform power management of the device 1100, a wired or wireless network interface 1150 configured to connect the device 1100 to a network, and an input / output interface 1158. The device 1100 may operate based on an operating system stored in the memory 1132, such as Windows Server TM , Mac OS X TM , Unix TM , Linux TM , FreeBSD TM or the like.
[0145] Figure 12FIG. is a block diagram of an apparatus for voice acquisition according to an exemplary embodiment. For example, apparatus 1200 may be a client (such as a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.).
[0146] Referring to Figure 12 , apparatus 1200 may include one or more of the following components: processing component 1202, memory 1204, power component 1206, multimedia component 1208, audio component 1210, input / output interface 1212, sensor component 1214, and communication component 1216.
[0147] Processing component 1202 generally controls the overall operation of apparatus 1200, such as operations associated with display, telephone calls, data communications, camera operations, and recording operations. Processing component 1202 may include one or more processors 1220 to execute instructions to complete all or part of the steps of the above-described voice acquisition method. In addition, processing component 1202 may include one or more modules to facilitate the interaction between processing component 1202 and other components. For example, processing component 1202 may include a multimedia module to facilitate the interaction between multimedia component 1208 and processing component 1202.
[0148] Memory 1204 is configured to store various types of data to support the operation of apparatus 1200. Examples of such data include instructions for any application or method operating on apparatus 1200, contact data, phone book data, messages, pictures, videos, etc. Memory 1204 may be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disk.
[0149] Power component 1206 provides power to the various components of apparatus 1200. Power component 1206 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for apparatus 1200.
[0150] The multimedia component 1208 includes a screen that provides an output interface between the device 1200 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors can sense not only the boundaries of a touch or swipe action but also detect the duration and pressure associated with the touch or swipe operation. In some embodiments, the multimedia component 1208 includes a front camera and / or a rear camera. When the device 1200 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera can receive external multimedia data. Each of the front camera and the rear camera can be a fixed optical lens system or have a focal length and optical zoom capabilities.
[0151] The audio component 1210 is configured to output and / or input audio signals. For example, the audio component 1210 includes a microphone (MIC) that is configured to receive external audio signals when the device 1200 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signals can be further stored in the memory 1204 or transmitted via the communication component 1216. In some embodiments, the audio component 1210 further includes a speaker for outputting audio signals.
[0152] The input / output interface 1212 provides an interface between the processing component 1202 and a peripheral interface module, which can be a keyboard, a click wheel, buttons, etc. These buttons can include but are not limited to: a home button, a volume button, a power button, and a lock button.
[0153] The sensor component 1214 includes one or more sensors for providing an assessment of various aspects of the state of the device 1200. For example, the sensor component 1214 can detect the on / off state of the device 1200, the relative positioning of components, such as the display and keypad of the device 1200. The sensor component 1214 can also detect a change in the position of the device 1200 or a component of the device 1200, the presence or absence of user contact with the device 1200, the orientation or acceleration / deceleration of the device 1200, and a change in the temperature of the device 1200. The sensor component 1214 can include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor component 1214 can also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor component 1214 can further include an acceleration sensor, a gyro sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.
[0154] The communication component 1216 is configured to facilitate communication between the device 1200 and other devices in a wired or wireless manner. The device 1200 can access a communication standard-based wireless network, such as WiFi, 2G, or 3G, or a combination thereof. In an exemplary embodiment, the communication component 1216 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 1216 further includes a Near Field Communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on Radio Frequency Identification (RFID) technology, Infrared Data Association (IrDA) technology, Ultra Wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.
[0155] In an exemplary embodiment, the device 1200 can be implemented by one or more Application Specific Integrated Circuits (ASICs), Digital Signal Processors (DSPs), Digital Signal Processing Devices (DSPDs), Programmable Logic Devices (PLDs), Field Programmable Gate Arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components for performing the above-described voice acquisition method.
[0156] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 1204 including instructions, and the above instructions can be executed by a processor 1220 of the device 1200 to complete the above-described voice acquisition method. For example, the non-transitory computer-readable storage medium can be a ROM, Random Access Memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc.
[0157] In another exemplary embodiment, a computer program product is also provided, and the computer program product includes a computer program capable of being executed by a programmable device, and the computer program has a code portion for performing the above-described voice acquisition method when executed by the programmable device.
[0158] Those skilled in the art will readily conceive of other embodiments of the present disclosure after considering the specification and practicing the present disclosure. The present disclosure is intended to cover any variations, uses, or adaptations of the present disclosure, which follow the general principles of the present disclosure and include known common knowledge or conventional technical means in the technical field not disclosed in the present disclosure. The specification and embodiments are only regarded as exemplary, and the true scope and spirit of the present disclosure are pointed out by the following claims.
[0159] It should be understood that the present disclosure is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.
Claims
1. A method for obtaining voice, characterized in that, Applied to a server, including: In response to receiving a voice interaction request message sent by a client, send address information corresponding to a target voice to be synthesized to the client, so that the client can obtain the target voice according to the address information, and the target voice is used to respond to the voice interaction request message; Obtain the streaming text corresponding to the target voice; Perform voice synthesis on the streaming text to obtain the target voice.
2. The method according to claim 1, characterized in that, The performing voice synthesis on the streaming text to obtain the target voice includes: In response to receiving a voice download request message sent by the client according to the address information, perform voice synthesis on the streaming text to obtain the target voice.
3. The method according to claim 1, characterized in that, The streaming text includes a text sequence composed of multiple first text units, and the target voice includes an audio stream; The performing voice synthesis on the streaming text to obtain the target voice includes: Loop and execute the voice synthesis step until the text sequence is traversed; The voice synthesis step includes: Traverse the text sequence in sequence, and according to a preset text division rule, sequentially merge one or more traversed first text units to obtain a second text unit; Perform voice synthesis on the second text unit to obtain the audio stream; In the case that the first text units not traversed in the text sequence are not empty, traverse the first text units not traversed and obtain an updated second text unit according to the preset text division rule.
4. The method according to claim 3, characterized in that, The performing voice synthesis on the second text unit to obtain the audio stream includes: Perform streaming audio synthesis on the second text unit through a preset text-to-speech (TTS) voice synthesis model to obtain the audio stream.
5. The method according to claim 1, characterized in that, The obtaining the streaming text corresponding to the target voice includes: Obtain the streaming text corresponding to the target voice according to the voice interaction request message.
6. The method according to claim 5, characterized in that, The obtaining the streaming text corresponding to the target voice according to the voice interaction request message includes: Obtain a request text corresponding to the voice interaction request message; Output the streaming text according to the request text through a preset dialogue generation model.
7. The method according to any one of claims 1-6, characterized in that, The method further includes: Send the target voice to the client.
8. A method for obtaining voice, characterized in that, Applied to a client, the method includes: Send a voice interaction request message to a server; Receive address information sent by the server, where the address information is the address information corresponding to a target voice to be synthesized sent by the server in response to receiving the voice interaction request message, and the target voice is used to respond to the voice interaction request message; Obtain the target voice from the server according to the address information.
9. The method according to claim 8, characterized in that, The obtaining the target voice from the server according to the address information includes: Send a voice download request message to the server according to the address information, so that the server performs voice synthesis on the streaming text corresponding to the target voice and then obtains the target voice; Receive the target voice sent by the server.
10. A voice acquisition device, characterized in that, Applied to a server, including: A first sending module, configured to, in response to receiving a voice interaction request message sent by a client, send address information corresponding to a target voice to be synthesized to the client, so that the client obtains the target voice according to the address information, and the target voice is used to respond to the voice interaction request message; An obtaining module, configured to obtain streaming text corresponding to the target voice; A voice synthesis module, configured to perform voice synthesis on the streaming text to obtain the target voice.
11. A voice acquisition device, characterized in that, Applied to a client, including: A second sending module, configured to send a voice interaction request message to a server; A first receiving module, configured to receive address information sent by the server, where the address information is address information corresponding to a target voice to be synthesized sent by the server to the client in response to receiving the voice interaction request message, and the target voice is used to respond to the voice interaction request message; A voice obtaining module, configured to obtain the target voice from the server according to the address information.
12. A server, characterized in that, Including: A processor; A memory for storing processor-executable instructions; Wherein, the processor is configured to: execute the steps of the method according to any one of claims 1-7.
13. A client, characterized in that, Including: A processor; A memory for storing processor-executable instructions; Wherein, the processor is configured to: execute the steps of the method according to claim 8 or 9.
14. A computer-readable storage medium, on which computer program instructions are stored, characterized in that, When the program instructions are executed by the processor, the steps of the method according to any one of claims 1 to 7 or 8 to 9 are implemented.
Citation Information
Cited By
Streaming audio synthesis method and device, storage medium and electronic device
CN121600905A