Audio generation method, system and device and storage medium
By receiving and converting the second streaming text in the audio generation method, the delay problem when the language model output text is solved, and timely replaying of audio playback is realized.
Patent Information
- Application Number
- CN202311541040.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-17
- Publication Date
- 2025-05-20
AI Technical Summary
In some application scenarios, delay problems often occur when the text output by the language model is converted into audio, resulting in the failure of the next audio to be broadcast in time after the previous audio broadcast ends.
After receiving the first streaming text and converting it to the first audio, receiving the second streaming text and determining the target time point based on its number of characters or character interval duration, obtaining the unplayed time of the first audio, and converting the second streaming text into the second audio within the time range defined by the unplayed time and the playback interval duration.
It effectively solves the problem of audio playback delay, ensuring that the second audio can be played no later than the playback interval after the first audio playback is finished.
Smart Images

Figure CN120020942A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical field of audio processing, and particularly to an audio generation method, system, device, and storage medium. Background Art
[0002] Currently, in some application scenarios, it is necessary to convert the text output by a language model into audio for playback. Since the language model outputs text character by character or phrase by phrase, when converting the text output by the language model into audio, it is usually necessary to first accumulate the text. When the accumulated text quantity meets the requirements, the accumulated text can be segmented, and the segmented sentences can be converted into audio.
[0003] In some technologies, on the one hand, due to the performance reasons of the language model itself, the speed of text generation is unstable; on the other hand, due to network transmission reasons, the text output by the language model may take a long time to be transmitted to the audio conversion server, which often causes a problem of audio broadcast delay. That is, the problem that the audio of the previous sentence has been broadcast for a long time and the audio of the next sentence has not been connected yet.
[0004] In view of this, there is an urgent need for a method that can solve the audio playback delay. Summary of the Invention
[0005] In view of this, the embodiments of the present disclosure provide an audio generation method, an audio generation system, an electronic device, and a computer-readable storage medium, which can solve the problem of audio playback delay.
[0006] On the one hand, the present disclosure provides an audio generation method, the method comprising:
[0007] Receiving a first streaming text and converting the first streaming text into a first audio;
[0008] Receiving a second streaming text located after the first streaming text, and determining a target time point during the reception of the second streaming text based on the number of characters of the received second streaming text or the time interval between adjacent characters received;
[0009] Obtaining the audio duration of the first audio after the target time point as the unplayed duration of the first audio;
[0010] Starting from the target time point, within the time range limited by the unplayed duration and the playback interval duration, converting the second streaming text into a second audio, where the playback interval duration represents the maximum time interval between the end time point of the first audio and the start time point of the second audio.
[0011] On the other hand, the present disclosure also provides an audio generation system, the system comprising:
[0012] A first receiving module, configured to receive a first streaming text and convert the first streaming text into a first audio;
[0013] A second receiving module, configured to receive a second streaming text located after the first streaming text, and determine a target time point during the receiving process of the second streaming text based on the number of characters of the received second streaming text or the interval duration between adjacent characters;
[0014] A duration obtaining module, configured to obtain the unplayed duration of the first audio after the target time point;
[0015] A conversion module, configured to start from the target time point and convert the second streaming text into a second audio within the duration range limited by the unplayed duration and the playback interval duration, where the playback interval duration represents the maximum time interval between the end time point of the first audio and the start time point of the second audio.
[0016] On the other hand, the present disclosure also provides a computer-readable storage medium, which is used to store a computer program. When the computer program is executed by a processor, the method described above is implemented.
[0017] On the other hand, the present disclosure also provides an electronic device, which includes a processor and a memory. The memory is used to store a computer program. When the computer program is executed by the processor, the method described above is implemented.
[0018] In the technical solutions of some embodiments of the present application, after converting the received first streaming text into a first audio, during the process of receiving the second streaming text, the unplayed duration of the first audio after the target time point is obtained, and the second streaming text is converted into a second audio within the duration range limited by the unplayed duration and the playback interval duration. In this way, after the first audio finishes playing, the second audio can be played at the latest no more than the playback interval duration, effectively solving the problem of audio playback delay. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] The features and advantages of the present disclosure will be more clearly understood by referring to the accompanying drawings. The drawings are schematic and should not be construed as imposing any limitation on the present disclosure. In the drawings:
[0020] Figure 1 The schematic diagram of an audio conversion system provided by an embodiment of the present application is shown;
[0021] Figure 2 The flowchart of an audio generation method provided by an embodiment of the present application is shown;
[0022] Figure 3 Shows the timing schematic diagram of the audio generation method provided by an embodiment of the present application;
[0023] Figure 4 Shows the timing schematic diagram of the audio generation method provided by another embodiment of the present application;
[0024] Figure 5 Shows the timing schematic diagram of the audio generation method provided by another embodiment of the present application;
[0025] Figure 6 Shows the module schematic diagram of the audio generation system provided by an embodiment of the present application;
[0026] Figure 7 Shows the schematic diagram of the electronic device provided by an embodiment of the present application. Detailed implementation manners
[0027] To make the objectives, technical solutions, and advantages of the embodiments of the present disclosure clearer, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present disclosure. Apparently, the described embodiments are some but not all of the embodiments of the present disclosure. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present disclosure without creative efforts shall fall within the scope of protection of the present disclosure.
[0028] Please refer to Figure 1 , which is the schematic diagram of the audio conversion system 100 provided by an embodiment of the present application. Figure 1 In, the audio conversion system 100 includes a client 11 and a server 12. The client 11 is communicatively connected to the server 12. The server 12 may be deployed with a language model 121 and an audio conversion device 122. The language model 121 can be used to implement functions such as text translation, question answering, and text classification. The working process of the audio conversion system 100 can be as follows:
[0029] 1) The client 11 receives the target content to be processed by the language model 121 and sends the target content to the language model 121. For example, when the language model 121 is used to implement the text translation function, the target content may be the text to be translated; when the language model 121 is used to implement the question answering function, the target content may be the question that needs to be answered by the language model 121.
[0030] 2) The language model 121 processes the target content and sends the processing result to the audio conversion device 122 and the client 11. For example, when the language model 121 is used to implement the text translation function, the processing result may include the result obtained after translating the target content; when the language model 121 is used to implement the question and answer function, the processing result may include the result obtained after answering the target content.
[0031] The processing result may specifically be a streaming text. Here, the streaming text refers to the characters serially output by the language model 121 in chronological order. Simply put, the language model 121 can output the processing result serially in units of one or more characters within a certain time range. For example, the language model 121 may output "Today" at 1.1 seconds, "is" at 2 seconds, "the" at 2.5 seconds, "weather" at 2.7 seconds, "relatively" at 3 seconds, "good" at 3.1 seconds, ",", "suitable" at 3.1 seconds, and "to go out" at 3.5 seconds, that is, sequentially output "Today>is>the>weather>relatively>good>,>suitable>to go out" as the processing result.
[0032] 3) The audio conversion device 122 receives and accumulates the processing results sent by the language model 121, and when the accumulated processing results can be segmented into one or more sentences, it segments the accumulated processing results and converts the segmented sentences into audio. For example, after the audio conversion device 122 accumulates "Today's weather is relatively good, suitable to go out", it can segment the accumulated processing result at the comma, obtaining two sentences "Today's weather is relatively good" and "Suitable to go out", and convert the two sentences into audio respectively; or, after the audio conversion device 122 accumulates "Today's weather is relatively good", it can perform segmentation and convert the sentence "Today's weather is relatively good" into audio.
[0033] 4) The audio conversion device 122 returns the converted audio to the client 11.
[0034] 5) The client 11 plays the received audio and displays the processing result returned by the language model 121 in text form.
[0035] In this way, the client 11 can support viewing the processing result of the target content through both visual and auditory means.
[0036] In Figure 1In the illustrated audio conversion system 100, on the one hand, due to the performance reasons of the language model 121 itself or network transmission reasons, there may be a problem of playback delay. Among them, delay means that after the previous audio is played, a long time interval is required to continue playing the next audio. In this case, the playback of the next audio belongs to playback delay. For example, after the audio conversion device 122 sends the audio of "The weather is nice today" to the client 11, it then continues to receive "suitable", but due to network and other reasons, it has not received "to go out" for a long time. Since "suitable" is not sufficient as a sentence, the audio conversion device 122 can continue to wait to accumulate more characters. In this case, after the client 11 plays the audio of "The weather is nice today", the audio conversion device 122 may not have converted the next audio yet, resulting in a long time interval after the audio "The weather is nice today" is played before the next audio can be played, that is, the next audio has a playback delay. On the other hand, if a sentence is long and the previous audio has been played, the sentence may not have been completely received yet, which will also cause the audio of this sentence to have a playback delay.
[0037] In view of this, the present application provides an audio generation method, which can solve the problem of audio playback delay. The audio generation method can be applied to an audio conversion device. Referring jointly to Figure 2 and Figure 3 . Figure 2 It is a flowchart of the audio generation method provided by an embodiment of the present application. Figure 3 It is a timing diagram of the audio generation method provided by an embodiment of the present application. Figure 2 In it, the audio generation method includes the following steps:
[0038] Step S21, receive the first streaming text and convert the first streaming text into the first audio.
[0039] Referring jointly to Figure 3 . The first streaming text can be the streaming text that the audio conversion device has received and completed audio conversion. After converting the first streaming text into the first audio, the first audio can be sent to the client for playback.
[0040] Figure 3 In it, the audio conversion device receives the first streaming text between time periods AB, converts the first streaming text into the first audio between time periods BC, and returns the first audio to the client for playback at time point C.
[0041] Step S22, receive the second streaming text located after the first streaming text, and determine the target time point during the reception of the second streaming text based on the number of characters of the received second streaming text or the interval duration between adjacent characters received.
[0042] Refer to in combination Figure 3 The second streaming text can be the streaming text that the audio conversion device is receiving but has not completed audio conversion yet. During the process of playing the first audio on the client, the audio conversion device can receive the second streaming text simultaneously.
[0043] The target time point refers to the time point that may cause a playback delay in the audio of the second streaming text during the process of receiving the second streaming text. Specifically, in combination with Figure 1 the relevant description, on the one hand, during the process of receiving the second streaming text, if the interval duration between two adjacent characters received is relatively long, it will cause the audio conversion device to wait for a long time when accumulating characters, which may lead to a problem of playback delay in the audio of the second streaming text. On the other hand, if the sentence in the second streaming text is relatively long, it will cause the audio conversion device to only receive the second streaming text for a long time but not perform the audio conversion of the second streaming text, which may also lead to a problem of playback delay in the audio of the second streaming text. Based on the above two reasons, the determination of the target time point is described below.
[0044] In some embodiments, during the process of receiving the second streaming text, starting from the first time point when the last character was received, if no next character is received within a preset duration, the first time point is taken as the target time point. Here, the preset duration refers to the maximum duration that the audio conversion device is allowed to wait for the next character under normal circumstances. If the duration for which the audio conversion device waits for the next character exceeds this maximum duration, it means that the duration for which the audio conversion device waits for the next character exceeds the allowed normal duration range. In this case, it may cause the audio conversion device to wait for a long time for character accumulation, thereby resulting in a playback delay in the audio of the second streaming text. Therefore, the first time point can be used as the target time point.
[0045] In some embodiments, during the process of receiving the second streaming text, the number of received characters can be counted, and the second time point when the number of characters reaches a specified number is taken as the target time point. Here, the specified number refers to the maximum number of characters that the audio conversion device can accumulate under normal circumstances. If the number of characters accumulated by the audio conversion device exceeds the maximum number of characters, it means that the number of characters accumulated by the audio conversion exceeds the allowed normal range. In this case, it may cause the audio conversion device to only receive the second streaming text for a long time but not perform the audio conversion of the second streaming text, thereby resulting in a playback delay in the audio of the second streaming text. Therefore, the second time point can be used as the target time point.
[0046] Step S23, obtain the audio duration of the first audio after the target time point as the unplayed duration of the first audio.
[0047] For ease of understanding, with reference to Figure 3 , assuming that time point D is the target time point, then the audio duration of the first audio after the target time point is time period DF.
[0048] Step S24: Starting from the target time point, within the time range limited by the unplayed duration and the playback interval duration, convert the second streaming text into a second audio, where the playback interval duration represents the maximum time interval between the end time point of the first audio and the start time point of the second audio.
[0049] Specifically, the playback interval duration can be set according to the actual situation. If it is required that the first audio and the second audio can be played continuously (i.e., the playback delay of the second audio is small), the playback interval duration can be set to a small value; if it is allowed that there can be a delay in the second audio, the playback interval duration can be set to a large value. In Figure 3 , use time period FH to represent the playback interval duration. The playback interval duration is after the unplayed duration of the first audio.
[0050] With reference to Figure 3 , it can be understood that if it is necessary to continue playing the second audio within the playback interval duration after the first audio finishes playing, then it is necessary to complete the audio conversion of the second streaming text within the time range limited by the unplayed duration and the playback interval duration.
[0051] Specifically, converting the second streaming text into a second audio within the time range limited by the unplayed duration and the playback interval duration can include two cases:
[0052] One case is that within the time range limited by the unplayed duration and the playback interval duration, the received second streaming text has accumulated relatively much. In this case, the second streaming text can be segmented according to the first segmentation logic, and the segmented sentences can be converted into a second audio. The individual sentences segmented according to the first segmentation logic can be relatively long.
[0053] Another case is that within the time range limited by the unplayed duration and the playback interval duration, the received second streaming text has accumulated relatively little. In this case, the second streaming text can be segmented according to the second segmentation logic, and the segmented sentences can be converted into a second audio. The individual sentences segmented according to the second segmentation logic can be relatively short. For example, a phrase can be used as a sentence.
[0054] Normally, when segmenting the second streaming text, first judge whether the accumulated second streaming text can be segmented according to the first segmentation logic. If it can, segment it according to the first segmentation logic. If not, then segment it based on the second segmentation logic.
[0055] Refer to Figure 4 , which is a timing schematic diagram of the audio generation method provided for another embodiment of this application. Figure 4 In , assume that time point D is the target time point, and time point G is within the unplayed duration of the first audio. If at time point G, the second streaming text accumulated by the audio conversion device can already be segmented according to the first segmentation logic, then the second streaming text can be segmented at time point G, and the sentences obtained by segmentation can be converted into the second audio. In this case, during the playback of the first audio, the audio conversion device can return the second audio. After the first audio is played, the second audio can be played continuously.
[0056] Refer to Figure 3 , and also assume that time point D is the target time point. Within the duration range defined by the unplayed duration DF and the playback interval duration FH of the first audio, time point H is the last time point. If at time point H, the second streaming text accumulated by the audio conversion device still cannot be segmented according to the first segmentation logic, then at time point H, the second streaming text can be segmented according to the second segmentation logic, and the sentences obtained by segmentation can be converted into the second audio. In this case, there is a playback delay for the second audio, and there is a maximum time interval between the end time point of the first audio and the start time point of the second audio.
[0057] In summary, in the technical solutions of some embodiments of this application, after converting the received first streaming text into the first audio, during the process of receiving the second streaming text, obtain the unplayed duration of the first audio after the target time point, and within the duration range defined by the unplayed duration and the playback interval duration, convert the second streaming text into the second audio. In this way, after the first audio is played, the second audio can be played at the latest not exceeding the playback interval duration, effectively solving the problem of audio playback delay.
[0058] In addition, this application determines the duration range for audio conversion of the second streaming text based on the unplayed duration of the first audio. When the unplayed duration of the first audio is relatively long, the corresponding duration range for audio conversion of the second streaming text is also relatively long, which is beneficial for accumulating more second streaming text to facilitate segmenting longer sentences according to the first segmentation logic. When the unplayed duration of the first audio is relatively short, the corresponding duration range for audio conversion of the second streaming text is also relatively short, which is beneficial for timely audio conversion of the second streaming text to avoid a large delay in the second audio.
[0059] The following further describes the solution of this application.
[0060] In some embodiments, starting from the target time point, within the time range defined by the unplayed duration and the playback interval duration, converting the second streaming text into second audio may specifically include:
[0061] Taking the second streaming text received before the target time point as the text to be converted, and determining the conversion duration required for audio-converting the text to be converted;
[0062] Taking the sum of the unplayed duration and the playback interval duration as the total duration;
[0063] Taking the difference between the total duration and the conversion duration as the maximum waiting duration;
[0064] Starting from the target time point, within the maximum waiting duration, converting the text to be converted into second audio.
[0065] In these embodiments, within the time range defined by the unplayed duration and the playback interval duration (i.e., the total duration), after subtracting the conversion duration for audio-converting the text to be converted, the audio conversion of the text to be converted is completed within the obtained maximum waiting duration. In this way, time can be reserved for the audio conversion of the text, ensuring that after the audio conversion is completed, the end time point of the first audio and the start time point of the second audio are within the maximum time interval.
[0066] For ease of understanding, taking Figure 3 as an example, the time period EH may represent the conversion duration of the text to be converted. The above determination of the maximum waiting duration may specifically include:
[0067] The unplayed duration DF of the first audio + the maximum time interval FH = the total duration DH
[0068] The total duration DH - the conversion duration EF = the initial waiting duration DE
[0069] When taking the initial waiting duration DE as the maximum waiting duration, the latest time point for audio-converting the text to be converted cannot exceed time point E.
[0070] Further, in this embodiment, in order to ensure that the determined maximum waiting duration is within a reasonable time range, after obtaining the maximum waiting duration, the audio generation method of the present application may further include:
[0071] Comparing the maximum waiting duration with a preset maximum duration. If the maximum waiting duration is greater than the maximum duration, taking the maximum duration as the maximum waiting duration; and / or
[0072] Comparing the maximum waiting duration with a preset minimum duration. If the maximum waiting duration is less than the minimum duration, taking the minimum duration as the maximum waiting duration.
[0073] Specifically, the maximum duration can represent the maximum value allowed for the maximum waiting duration, and the minimum duration can represent the minimum value allowed for the maximum waiting duration. If the maximum waiting duration is not within the range defined by the maximum duration and the minimum duration, the maximum duration or the minimum duration can be used as the maximum waiting duration. In this way, the rationality of the maximum waiting duration can be ensured, and the control accuracy during audio conversion of the second streaming text can be improved.
[0074] The following describes how to determine the unplayed duration of the first audio.
[0075] With reference to Figure 5 , a timing schematic diagram of the audio generation method provided by another embodiment of this application is shown. Figure 5 In, the first streaming text includes k sub-texts. Wherein, k is an integer greater than or equal to 1. The audio conversion device can sequentially convert each sub-text into a sub-audio. That is, the first audio includes k sub-audios. There may be a pause duration between two adjacent sub-audios, or two adjacent sub-audios may be continuous. In the case where there is a pause duration between two adjacent sub-audios, the pause duration can include the playback interval duration and the delay caused by network transmission. For example Figure 5 in, between the first sub-audio and the second sub-audio, the time period BC can represent the playback interval duration, and the time period CD can represent the delay caused by network transmission.
[0076] Based on Figure 5 , in some embodiments, the unplayed duration of the first audio can be determined based on the following method:
[0077] 11) Based on the total audio duration of the k sub-audios and the total pause duration between the k sub-audios, obtain the playback duration of the first audio.
[0078] Specifically, the total audio duration can be the sum of the audio durations of the k sub-audios, and the total pause duration can be the sum of the pause durations between the k sub-audios. This process can be shown as in expression (1):
[0079] total_time = accumulate_audio[k] + accumulate_waiting[k] (1)
[0080] Wherein, total_time represents the playback duration of the first audio, accumulate_audio[k] represents the total audio duration of the k sub-audios, and accumulate_waiting[k] represents the total pause duration between the k sub-audios.
[0081] 12) Use the difference between the first time point when the last character was received and the third time point when the first sub - audio was converted as the played duration of the first audio.
[0082] This process can be shown as in Expression (2):
[0083] past_time = last_recv_time - package_time[1] (2)
[0084] Where past_time represents the played duration of the first audio, last_recv_time represents the first time point when the last character was received, and package_time[1] represents the third time point when the first sub - audio was converted.
[0085] 13) Use the difference between the playing duration of the first audio and the played duration as the unplayed duration of the first audio.
[0086] In this way, the unplayed duration of the first audio can be obtained. In the above method for obtaining the unplayed duration, the audio duration, the first time point, the third time point, etc. can be directly obtained on the audio conversion device side, which is convenient for calculation.
[0087] The following describes the method for determining the total pause duration accumulate_waiting[k] between the above k sub - audios.
[0088] Refer to Figure 5 , since there are no other sub - audios before the first sub - audio and no other sub - audios after the k - th sub - audio, the total pause duration accumulate_waiting[k] between sub - audios should actually be the sum of the pause durations between the n - th sub - audio and the (n - 1) - th sub - audio, where n ranges from 2 to k. In view of this, in some embodiments, the total pause duration accumulate_waiting[k] between the above k sub - audios can be determined based on the following method:
[0089] 21) Use the difference between the time point when the n - th sub - audio was converted and the time point when the first sub - audio was converted as the playing duration of the first n - 1 sub - audios.
[0090] This process can be shown as in Expression (3):
[0091] total_time[n - 1] = package_time[n] - package_time[1] (3)
[0092] Among them, total_time[n - 1] represents the playing duration of the first n - 1 sub - audio clips, package_time[n] represents the time point when the nth sub - audio clip is converted, and package_time[1] represents the time point when the first sub - audio clip is converted.
[0093] 22) Take the difference between the playing duration of the first n - 1 sub - audio clips and the total audio duration of the first n - 1 sub - audio clips as the pause duration between the nth sub - audio clip and the (n - 1)th sub - audio clip.
[0094] This process can be shown as in expression (4):
[0095] waiting[n]=total_time[n - 1]-accumulate_audio[n - 1](4)
[0096] Among them, waiting[n] represents the pause duration between the nth sub - audio clip and the (n - 1)th sub - audio clip, and accumulate_audio[n - 1] represents the total audio duration of the first n - 1 sub - audio clips.
[0097] Furthermore, the total audio duration accumulate_audio[n - 1] of the first n - 1 sub - audio clips can be shown as in expression (5):
[0098] accumulate_audio[n - 1]=audio[1]+audio[2]+……+audio[n - 1](5)
[0099] Among them, audio[1], audio[2], ……, audio[n - 1] represent the audio durations of each sub - audio clip.
[0100] Furthermore, considering that the difference between the playing duration of the first n - 1 sub - audio clips and the total audio duration of the first n - 1 sub - audio clips (i.e., total_time[n - 1]-accumulate_audio[n - 1]) may be negative. In practice, the minimum pause duration between the nth sub - audio clip and the (n - 1)th sub - audio clip can only be 0 (i.e., there is no pause duration between the nth sub - audio clip and the (n - 1)th sub - audio clip, and they will play continuously). In view of this, taking the difference between the playing duration of the first n - 1 sub - audio clips and the total audio duration of the first n - 1 sub - audio clips as the pause duration between the nth sub - audio clip and the (n - 1)th sub - audio clip can include:
[0101] If the difference between the playing duration of the first n - 1 sub - audio clips and the total audio duration of the first n - 1 sub - audio clips is greater than or equal to 0, take this difference as the pause duration between the kth sub - audio clip and the (k - 1)th sub - audio clip;
[0102] If the difference between the playback duration of the first n-1 sub-audios and the total audio duration of the first n-1 sub-audios is less than 0, set the pause duration between the nth sub-audio and the n-1th sub-audio to 0.
[0103] This is to avoid the situation where the difference between the playback duration of the first n-1 sub-audios and the total audio duration of the first n-1 sub-audios is a negative value.
[0104] 23) Based on the pause duration between the nth sub-audio and the n-1th sub-audio, the total pause duration between the above k sub-audios is obtained accumulate_waiting[k].
[0105] In this embodiment, accumulate_waiting[k] can be expressed as expression (6):
[0106] accumulate_waiting[k]=accumulate_waiting[n-1]+waiting[n](6)
[0107] Accumulate_waiting[n-1] indicates the total duration of pauses between the first n-1 sub-audios, and waiting[n] indicates the duration of pauses between the nth sub-audio and the n-1th sub-audio.
[0108] This completes the description of the audio generation method of this application.
[0109] Corresponding to the audio generation method, the present application also provides an audio generation system. Please refer to Figure 6 , is a module diagram of an audio generation system provided by an embodiment of the present application. The audio generation system includes:
[0110] A first receiving module, configured to receive a first streaming text and convert the first streaming text into a first audio;
[0111] A second receiving module is used to receive a second stream text located after the first stream text, and determine a target time point in the second stream text receiving process based on the number of characters of the received second stream text or the length of interval between adjacent characters received;
[0112] Duration acquisition module, used to acquire the unplayed duration of the first audio after the target time point;
[0113] A conversion module is used to convert the second streaming text into the second audio starting from the target time point within the time range defined by the unplayed time and the play interval time, wherein the play interval time represents the maximum time interval between the end time point of the first audio and the start time point of the second audio.
[0114] Please refer to Figure 7 , which is a schematic diagram of an electronic device provided for an embodiment of the present application. The electronic device includes a processor and a memory. The memory is used to store a computer program. When the computer program is executed by the processor, the above-mentioned method is implemented.
[0115] Among them, the processor may be a central processing unit (CPU). The processor may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. chips, or combinations of the above types of chips.
[0116] As a non-transitory computer-readable storage medium, the memory can be used to store non-transitory software programs, non-transitory computer-executable programs, and modules, such as program instructions / modules corresponding to the methods in the embodiments of the present invention. By running the non-transitory software programs, instructions, and modules stored in the memory, the processor can execute various functional applications and data processing of the processor, that is, implement the methods in the above method embodiments.
[0117] The memory may include a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store data created by the processor, etc. In addition, the memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory may optionally include a memory remotely provided with respect to the processor, and these remote memories can be connected to the processor through a network. Examples of the above networks include, but are not limited to, the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0118] An embodiment of the present application also provides a computer-readable storage medium. The computer-readable storage medium is used to store a computer program. When the computer program is executed by the processor, the above-mentioned method is implemented.
[0119] Although the embodiments of the present disclosure have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the present disclosure, and such modifications and variations all fall within the scope defined by the appended claims.
Claims
1. An audio generation method, characterized in that: The method comprises: Receiving a first streaming text, and converting the first streaming text into a first audio; receiving a second stream text following the first stream text, and determining a target time point in a receiving process of the second stream text based on the number of characters of the second stream text received or the length of intervals between adjacent characters received; Obtaining the audio duration of the first audio after the target time point as the unplayed duration of the first audio; Starting from the target time point, the second streaming text is converted into a second audio within the time range defined by the unplayed time and the play interval time, wherein the play interval time represents the maximum time interval between the end time point of the first audio and the start time point of the second audio.
2. The method according to claim 1, characterized in that Determining the target time point based on the interval duration between receiving adjacent characters includes: During the receiving process of the second stream text, starting from the first time point when the last character is received, if the next character is not received within a preset time length, the first time point is used as the target time point.
3. The method according to claim 1, characterized in that Determining the target time point based on the number of characters of the second stream text received includes: During the receiving process of the second streaming text, the number of characters received is counted, and a second time point when the number of characters reaches a specified number is used as the target time point.
4. The method according to claim 1, characterized in that The step of converting the second streaming text into a second audio from the target time point within a time range defined by the unplayed time and the play interval time includes: Taking the second streaming text received before the target time point as the text to be converted, and determining the conversion time required for performing audio conversion on the text to be converted; The sum of the non-playing time and the playing interval time is taken as the total time; The difference between the total duration and the conversion duration is used as the maximum waiting duration; Starting from the target time point, within the maximum waiting time, the text to be converted is converted into the second audio.
5. The method according to claim 4, characterized in that After obtaining the maximum waiting time, the method further includes: Compare the maximum waiting time with a preset maximum waiting time, and if the maximum waiting time is greater than the maximum waiting time, use the maximum waiting time as the maximum waiting time; and / or The maximum waiting time is compared with a preset minimum waiting time. If the maximum waiting time is less than the minimum waiting time, the minimum waiting time is used as the maximum waiting time.
6. The method according to claim 1 or 4, characterized in that: The first audio includes k sub-audios; and the unplayed time is determined based on the following method: Obtaining a playback duration of the first audio based on a total audio duration of the k sub-audios and a total pause duration between the k sub-audios; The difference between the first time point when the character is received for the last time and the third time point when the first sub-audio is converted is used as the playing time of the first audio; The difference between the playing time of the first audio and the played time is used as the unplayed time of the first audio; The value of k is an integer greater than or equal to 1.
7. The method according to claim 6, characterized in that The total freeze duration between the k sub-audios is determined based on the following method: The difference between the time point at which the nth sub-audio is converted and the time point at which the first sub-audio is converted is used as the playback duration of the first n-1 sub-audios; The difference between the playing duration of the first n-1 sub-audios and the total audio duration of the first n-1 sub-audios is used as the pause duration between the kth sub-audio and the k-1th sub-audio; Based on the freeze duration between the nth sub-audio and the n-1th sub-audio, obtain the total freeze duration between the k sub-audios; The value of n is between 2 and k.
8. The method according to claim 7, characterized in that The step of using the difference between the playing duration of the first n-1 sub-audios and the total audio duration of the first n-1 sub-audios as the freeze duration between the nth sub-audio and the n-1th sub-audio includes: If the difference between the playback duration of the first n-1 sub-audios and the total audio duration of the first n-1 sub-audios is greater than or equal to 0, the difference is used as the freeze duration between the kth sub-audio and the k-1th sub-audio; If the difference between the playback duration of the first n-1 sub-audios and the total audio duration of the first n-1 sub-audios is less than 0, the pause duration between the nth sub-audio and the n-1th sub-audio is set to 0.
9. An audio generation system, characterized in that: The system comprises: A first receiving module, configured to receive a first streaming text and convert the first streaming text into a first audio; A second receiving module is used to receive a second streaming text located after the first streaming text, and determine a target time point in the receiving process of the second streaming text based on the number of characters of the received second streaming text or the interval length between adjacent characters; A duration acquisition module, used to acquire the unplayed duration of the first audio after the target time point; A conversion module is used to convert the second streaming text into a second audio starting from the target time point within the time range specified by the unplayed time and the play interval time, wherein the play interval time represents the maximum time interval between the end time point of the first audio and the start time point of the second audio.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium is used to store a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 8 is implemented.
11. An electronic device, characterized in that: The electronic device comprises a processor and a memory, wherein the memory is used to store a computer program, and when the computer program is executed by the processor, the method according to any one of claims 1 to 8 is implemented.