Audio broadcast method, device, system, electronic device and storage medium
By refreshing service push information in real time and dynamically adjusting the synthesis speed, the problem of audio broadcast strategy pauses when there is network delay or uneven data frames is solved, the smoothness and real-time nature of audio broadcast is achieved, and the user experience is improved.
Patent Information
- Application Number
- CN202411754114.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-02
- Publication Date
- 2025-09-19
- Estimated Expiration
- 2044-12-02
AI Technical Summary
The existing technology cannot dynamically adjust the audio broadcast strategy when there is network delay or uneven data frame transmission, resulting in pauses or discontinuities in audio playback, affecting the user experience.
After receiving the audio frames sent by the synthesis engine, the service push information is refreshed in real time according to the push time node and length, the synthesis speed is dynamically adjusted, and silent segments are inserted into the phoneme chain to optimize the audio broadcast strategy.
It achieves global smoothness and real-time performance of audio broadcasting, improves user experience, and avoids pauses or discontinuities caused by network delays or uneven data frame transmission.
Smart Images

Figure CN119728658B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of audio processing technology, and in particular to an audio broadcasting method, device, system, electronic device and storage medium. Background Art
[0002] In real-time audio broadcast scenarios, the audio data transmission rate is affected by a variety of factors. These factors can cause data frames to arrive at the front end at uneven times, leading to pauses, freezes, backlogs, and other issues during the audio broadcast process, resulting in a poor user experience. Therefore, how to provide high-quality audio broadcasts to enhance the user experience is an important topic that needs urgent research.
[0003] In related technologies, a fixed playback rate or a larger playback buffer is usually used to store the transmitted data frames to optimize the audio broadcast strategy. However, when there is network delay or uneven data frame transmission, it is impossible to dynamically adjust the audio broadcast strategy, resulting in the problem of pauses or discontinuities in audio playback. Summary of the Invention
[0004] The present invention provides an audio broadcast method, device, system, electronic device and storage medium, which are used to solve the defect in the prior art that the playback strategy cannot be adjusted when the network is delayed or the data frame transmission is uneven, resulting in pauses or discontinuities in the audio playback process, thereby improving the global fluency and real-time performance of the audio broadcast.
[0005] The present invention provides an audio broadcasting method, comprising:
[0006] Upon receiving the last synthesized audio frame sent by the synthesis engine, refreshing the service push information according to the push time node and push time length of the last synthesized audio frame and the current push time node to obtain the current service push information;
[0007] Refreshing the synthesis speed of the synthesis engine according to the phoneme chain and the current service push information to obtain the current synthesis speed; the phoneme chain is constructed based on the phoneme text information synthesized and returned by the synthesis engine;
[0008] According to the phoneme chain and the current service push information, a silent segment is configured for the current synthesized audio frame sent by the synthesis engine to obtain an audio frame to be broadcast; the current synthesized audio frame is synthesized by the synthesis engine according to the current synthesis speed;
[0009] The audio frame to be broadcast is pushed to an audio processing end, and the audio processing end is used to broadcast the audio frame to be broadcast.
[0010] According to an audio broadcast method provided by the present invention, refreshing service push information based on the push time node and push time length of the previous synthesized audio frame and the current push time node to obtain current service push information includes:
[0011] Summing the push time node and the push time length of the previous synthesized audio frame to obtain a first target time node;
[0012] According to the comparison result between the current push time node and the first target time node, the service push information is refreshed to obtain the current service push information.
[0013] According to an audio broadcast method provided by the present invention, refreshing the service push information according to a comparison result between the current push time node and the first target time node to obtain the current service push information includes:
[0014] When the current push time node is less than or equal to the first target time node, refreshing the service push information according to the first refresh strategy to obtain the current service push information;
[0015] When the current push time node is greater than the first target time node, refreshing the service push information according to the second refresh strategy to obtain the current service push information;
[0016] Among them, the first refresh strategy includes a strategy of keeping the push time node of the previous synthesized audio frame unchanged, and a strategy of updating the push time length of the previous synthesized audio frame to the sum of the push time length of the previous synthesized audio frame and the current push time length; the second refresh strategy includes a strategy of updating the push time node of the previous synthesized audio frame to the current push time node, and a strategy of updating the push time length of the previous synthesized audio frame to the current push time length.
[0017] According to an audio broadcast method provided by the present invention, the method further includes:
[0018] Summing the first target time node and the broadcast freeze time threshold to obtain a second target time node;
[0019] Summing the first target time node and the broadcast redundancy time threshold to obtain a third target time node;
[0020] In a case where the current push time node is greater than the first target time node, if it is determined that the current push time node is less than or equal to the second target time node, determining that the second refresh strategy also includes a strategy of updating the broadcast jam identifier to be true;
[0021] In the case that the current push time node is greater than the first target time node, if it is determined that the current push time node is greater than the third target time node, it is determined that the second refresh strategy also includes a strategy of updating the broadcast redundant identifier to be true.
[0022] According to an audio broadcasting method provided by the present invention, refreshing the synthesis speed of the synthesis engine according to the phoneme chain and the current service push information to obtain the current synthesis speed includes:
[0023] According to the phoneme chain, the number of characters of the currently received synthesized audio frame is updated;
[0024] Determining the number of characters in the currently sent unsynthesized audio frame based on the number of characters in the currently received synthesized audio frame;
[0025] When it is determined that the number of characters of the currently sent unsynthesized audio frame is greater than or equal to a preset number, and when it is determined based on the current service push information that the current broadcast state is a broadcast redundant state, the synthesis speed of the synthesis engine is refreshed according to the preset speed to obtain the current synthesis speed.
[0026] According to an audio broadcast method provided by the present invention, configuring a silent segment for the current synthesized audio frame sent by the synthesis engine according to the phoneme chain and the current service push information to obtain an audio frame to be broadcast, including:
[0027] If it is determined that the number of characters in the currently sent unsynthesized audio frame is less than the preset number, and if it is determined according to the current service push information that the current broadcast state is a broadcast blocked state, matching and obtaining a target phoneme frame having a rhythmic pause node in the currently synthesized audio frame based on the rhythmic pause identifiers of each phoneme frame in the phoneme chain;
[0028] According to the rhythmic pause configuration position corresponding to the target phoneme frame, a silent segment is inserted into the current synthesized audio frame to obtain the audio frame to be broadcasted.
[0029] According to an audio broadcasting method provided by the present invention, the phoneme chain is constructed based on the following steps:
[0030] Parsing the phoneme text information to obtain phoneme information of each word;
[0031] splicing the phoneme information of each word in the phoneme text information in a manner of splicing different syllables of the same word to obtain a phoneme frame corresponding to each word;
[0032] According to the sorting order of each word in the phoneme text information, links are established between the phoneme frames corresponding to each word to obtain an initial phoneme chain;
[0033] The pronunciation duration, synthesis time-consuming information and rhythmic pause identifier of the phoneme frame corresponding to each word are configured in the initial phoneme chain to obtain the phoneme chain.
[0034] According to an audio broadcast method provided by the present invention, the method further includes:
[0035] In the case that the last synthesized audio frame is not received, obtaining the number of characters of the currently sent unsynthesized audio frame and the length of the currently pushed but unbroadcasted audio frame;
[0036] Obtaining a current broadcast state according to the number of characters of the currently sent unsynthesized audio frame and the length of the currently pushed but unbroadcasted audio frame;
[0037] When it is determined that the current broadcast state is a broadcast redundant state, the synthesis speed of the synthesis engine is accelerated to obtain the current synthesis speed.
[0038] The present invention also provides an audio broadcasting device, comprising:
[0039] A first refreshing unit is configured to, upon receiving a last synthesized audio frame sent by the synthesis engine, refresh the service push information according to the push time node and push time length of the last synthesized audio frame and the current push time node, thereby obtaining the current service push information;
[0040] A second refreshing unit is configured to refresh the synthesis speed of the synthesis engine according to the phoneme chain and the current service push information to obtain a current synthesis speed; the phoneme chain is constructed based on the phoneme text information synthesized and returned by the synthesis engine;
[0041] an audio configuration unit, configured to configure a silent segment for a current synthesized audio frame sent by the synthesis engine according to the phoneme chain and the current service push information, to obtain an audio frame to be broadcast; the current synthesized audio frame is synthesized by the synthesis engine according to the current synthesis speed;
[0042] The audio broadcast unit is used to push the audio frame to be broadcast to the audio processing end, and the audio processing end is used to broadcast the audio frame to be broadcast.
[0043] The present invention also provides an audio broadcast system, comprising an audio processing terminal, a broadcast monitor and a synthesis engine;
[0044] The broadcast monitor is connected to the audio processing end and the synthesis engine respectively, and is used to execute any of the above-mentioned audio broadcast methods.
[0045] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements any of the above-described audio broadcast methods when executing the program.
[0046] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which implements any of the above-mentioned audio broadcasting methods when executed by a processor.
[0047] The present invention also provides a computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the computer program implements any of the above-mentioned audio broadcasting methods.
[0048] The audio broadcast method, device, system, electronic device and storage medium provided by the present invention, after receiving the last synthesized audio frame sent by the synthesis engine, refresh the service push information in real time according to the push time node, push time length and current time node of the last synthesized audio frame, and based on the current service push information refreshed in real time and the phoneme chain constructed by the phoneme text information synthesized and returned by the synthesis engine, finely and intelligently perform dynamic adjustment of the synthesis speed at the phoneme level and dynamic configuration of the silent segments, thereby achieving more finely and intelligent dynamic optimization of the audio broadcast strategy, so as to more effectively cope with the adverse effects caused by network delay or uneven data frame transmission, thereby improving the overall fluency and real-time performance of the audio broadcast and enhancing the user experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] In order to more clearly illustrate the technical solutions in the present invention or the prior art, a brief introduction is given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0050] Figure 1 It is a flow chart of the audio broadcasting method provided by the present invention.
[0051] Figure 2 It is a schematic diagram of the interaction of various components in the audio broadcasting system provided by the present invention.
[0052] Figure 3 It is a flowchart of the phoneme chain construction provided by the present invention.
[0053] Figure 4This is one of the distribution diagrams of the audio broadcast time provided by the present invention.
[0054] Figure 5 This is the second distribution diagram of the audio broadcast time provided by the present invention.
[0055] Figure 6 This is the third distribution diagram of the audio broadcast time provided by the present invention.
[0056] Figure 7 This is the fourth distribution diagram of the audio broadcast time provided by the present invention.
[0057] Figure 8 It is a structural diagram of the audio broadcasting device provided by the present invention.
[0058] Figure 9 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION
[0059] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are only some of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.
[0060] In real-time audio broadcast scenarios, the audio data transmission rate is affected by many factors, such as network status, bandwidth fluctuations, data transmission delays, and the real-time processing speed of the synthesis engine. These factors may cause the data frames to arrive at the front end at uneven times, thereby causing pauses, freezes, backlogs, and other problems during the audio broadcast process. This situation is particularly evident in high-demand real-time application scenarios, such as simultaneous voice interpretation and virtual voice assistants, which require continuous and stable audio output to ensure that users can receive information without obstacles. Therefore, there is an urgent need to provide an intelligent real-time audio broadcast method to cope with the adverse effects of data frame fluctuations, ensure smooth audio playback without freezes and backlogs, improve the quality of audio broadcasting, and enhance user experience.
[0061] Common audio broadcast systems mostly use a fixed playback rate. Regardless of how the data transmission situation changes, the audio output speed remains basically unchanged. Although this method is simple, when there is network delay or uneven data frame transmission, pauses or discontinuities in audio playback may occur. In addition, to alleviate the problem of playback interruption, some systems set up a larger playback buffer to store the transmitted data frames, and read the content from the buffer when insufficient data arrives, thereby avoiding playback pauses. However, when there is network delay or uneven data frame transmission, this method cannot dynamically adjust the audio broadcast strategy in real time, and may still cause playback delays. Once the buffer is exhausted, pauses are still difficult to avoid. Therefore, it still has the problem of pauses or discontinuities in audio playback.
[0062] Although, some technical solutions attempt to dynamically adjust the audio playback rate. For example, the speech speed adjustment algorithm is used to accelerate the part with less transmission data to balance the problem of insufficient data; or when insufficient data arrives, silent segments are inserted to maintain the continuity of the audio. Although such solutions have improved the coherence of the broadcast to a certain extent, they are usually unable to take into account the problems of unsmooth broadcast caused by data backlog and speech speed changes. In addition, the existing audio speed adjustment algorithms generally lack support for fine-grained adjustment at the phoneme level in practical applications, and cannot accurately insert silence or acceleration adjustments on each phoneme chain, making it difficult to ensure global fluency.
[0063] In this regard, this embodiment provides an audio broadcast method, which dynamically handles the backlog problem of content to be broadcast by predicting the relationship between the content to be broadcast and the current backlog; and based on the refined and intelligent dynamic adjustment of the synthesis speed at the phoneme level and the dynamic configuration of silent segments, so as to achieve the continuity and stability of audio content push, thereby improving the overall fluency and real-time performance of the audio broadcast.
[0064] Figure 1 The figure is a flow chart of the audio broadcast method provided by the present invention. The method can be applied to an audio broadcast system, which at least includes an audio processing terminal, a broadcast monitor and a synthesis engine.
[0065] Figure 2 Schematic diagram of the interaction of various components in the audio broadcast system provided by the present invention; Figure 2As shown, the audio processing end is used to receive voice data transmitted by the user, convert or translate the voice data into text data, and transmit the text data as data to be synthesized to the broadcast monitor; the broadcast monitor is used to receive the data to be synthesized transmitted by the upstream audio processing end, and by maintaining the phoneme chain and throughput content, determine in real time whether synthesis speed adjustment is currently required to generate the corresponding synthesis speed, and send the synthesis speed and data to be synthesized to the downstream synthesis engine so that the downstream synthesis engine performs audio frame synthesis according to the synthesis speed and data to be synthesized; in addition, the broadcast monitor is also used to receive the audio frames synthesized by the downstream synthesis engine, and by maintaining the phoneme chain and throughput content, perform backlog judgment in real time to determine whether rhythm point silence insertion is currently required to control the push of audio frames, and then push the synthesized audio frames to the audio processing end to achieve smooth broadcast of audio frames. Therefore, through the mutual cooperation of various components in the audio broadcast system, dynamic optimization of audio broadcast is achieved to achieve the consistency and stability of audio content push, thereby improving the overall fluency and real-time performance of audio broadcast.
[0066] It should be noted that this method can be widely applied to various audio broadcast scenarios, including but not limited to simultaneous voice interpretation scenarios and virtual voice assistant scenarios, and this embodiment does not specifically limit these. The following describes the method provided by this embodiment using the simultaneous voice interpretation scenario as an example. For other scenarios, refer to this scenario to adaptively implement the audio broadcast method.
[0067] like Figure 1 As shown, the method includes: step 110, step 120, step 130 and step 140.
[0068] Step 110: upon receiving the last synthesized audio frame sent by the synthesis engine, refresh the service push information according to the push time node and push time length of the last synthesized audio frame and the current push time node to obtain the current service push information.
[0069] Optionally, during the current audio broadcast process, the broadcast monitor may first determine whether the previous synthesized audio frame sent by the synthesis engine has been received. When the previous synthesized audio frame is received, the push time node P and push time length Dp of the previous synthesized audio frame are obtained according to the push record of the previous synthesized audio frame, and the current push time node N is loaded to dynamically determine the current broadcast status by combining the push time node P, push time length Dp and the current push time node N, and determine the corresponding service push information refresh strategy based on the current broadcast status, and dynamically refresh the service push information according to the corresponding service push information refresh strategy to obtain the current service push information, so as to dynamically adjust the synthesis speed based on the current service push information, thereby ensuring the continuity and stability of the audio content push.
[0070] The different broadcast states here correspond to different service push information refresh strategies, such as the first refresh strategy corresponding to the non-stuck broadcast state, the second refresh strategy corresponding to the stuck broadcast state, etc., wherein the first refresh strategy and the second refresh strategy contain different refresh strategies for the push time node P and the push time length Dp.
[0071] The method for obtaining the current broadcast status here includes performing multiple comparisons and judgments on the push time node and push time length of the previous synthetic audio frame, as well as the current push time node, to obtain the current broadcast status; or, inputting the push time node and push time length of the previous synthetic audio frame, as well as the current push time node into a pre-trained recognition model, and the recognition model applies the push time node and push time length of the previous synthetic audio frame, as well as the current push time node, to perform nonlinear calculations to identify the current broadcast status. The specific method can be determined according to actual needs, and this embodiment does not specifically limit this.
[0072] Step 120: Refresh the synthesis speed of the synthesis engine according to the phoneme chain and the current service push information to obtain the current synthesis speed; the phoneme chain is constructed according to the phoneme text information synthesized and returned by the synthesis engine.
[0073] The phoneme chain here is a chain data structure of multiple phoneme frames linked by parsing the phoneme text information returned by the synthesis engine. Each node can represent the pronunciation information of the synthesized word, including but not limited to pronunciation duration, synthesis time information and prosodic pause identifiers. The prosodic pause identifiers include those used to represent different prosodic pause types, including but not limited to single-word pauses, word-level pauses (L1), short pauses (Sp in the sentence), complete sentence pauses (Sil) and other prosodic pause types. Among them, phonemes refer to the index information of the synthesized audio returned by the synthesis engine, including the text corresponding to the synthesized audio and the pronunciation duration of each synthesized monosyllable.
[0074] For example, in some embodiments, the phoneme chain is constructed based on the following steps:
[0075] Parsing the phoneme text information to obtain phoneme information of each word;
[0076] splicing the phoneme information of each word in the phoneme text information in a manner of splicing different syllables of the same word to obtain a phoneme frame corresponding to each word;
[0077] According to the sorting order of each word in the phoneme text information, links are established between the phoneme frames corresponding to each word to obtain an initial phoneme chain;
[0078] The pronunciation duration, synthesis time-consuming information and rhythmic pause identifier of the phoneme frame corresponding to each word are configured in the initial phoneme chain to obtain the phoneme chain.
[0079] Figure 3 This is a flow chart of the phoneme chain construction provided by the present invention. Figure 3 As shown, the specific process of building a phoneme chain includes: initializing a queue; parsing the phoneme text information returned by the synthesis engine to obtain the phoneme information of each character; splicing the phoneme information of each character in the phoneme text information according to the splicing method of different syllables of the same character to obtain the phoneme frame corresponding to each character (also called synthesized phoneme frame); establishing links between the phoneme frames corresponding to each character according to the sorting order of each character in the phoneme text information to obtain the initial phoneme chain. The pronunciation duration of each character is then calculated to obtain the pronunciation duration of the phoneme frame corresponding to each character; the total synthesis time information starting from the first frame is accumulated to obtain the synthesis time information of the phoneme frame corresponding to each character; and according to the prosodic pause configuration information of the synthesis agreement, the identifiers allowed for pauses and silence insertion, namely prosodic pause identifiers, are parsed in the phoneme frame corresponding to each character. Finally, by configuring the pronunciation duration, synthesis time information and rhythmic pause identifier of the phoneme frames corresponding to each word in the initial phoneme chain, the final complete phoneme chain can be obtained, so that the synthesis speed can be dynamically adjusted and the silent segments can be dynamically configured based on the phoneme chain. This will significantly improve the fluency and naturalness of the audio broadcast, ensuring a smooth and uninterrupted broadcast process.
[0080] Optionally, after obtaining the current service push information, the phoneme chain constructed by combining the current service push information and the factor text information returned by the synthesis engine can be used to identify the current broadcast status and the number of characters in the unsynthesized audio frame. Based on the identified current broadcast status and the number of characters in the unsynthesized audio frame, the synthesis speed of the synthesis engine can be dynamically refreshed so that the refreshed current synthesis speed matches the current broadcast speed, thereby improving the naturalness and fluency of the audio broadcast, and ensuring the real-time nature of the audio broadcast, thereby enhancing the user experience.
[0081] Step 130 : Based on the phoneme chain and the current service push information, a silent segment is configured for the current synthesized audio frame sent by the synthesis engine to obtain an audio frame to be broadcast; the current synthesized audio frame is synthesized by the synthesis engine according to the current synthesis speed.
[0082] Optionally, after obtaining the current synthesis speed, the current data to be synthesized and the current synthesis speed received from the upstream audio processor can be transmitted to the synthesis engine, so that the synthesis engine can synthesize the audio frames of the current data to be synthesized according to the current synthesis speed to obtain the current synthesized audio frame.
[0083] After receiving the current synthesized audio frame transmitted by the synthesis engine, the current broadcast status and the number of characters in the unsynthesized audio frame can be identified in combination with the current service push information and the phoneme chain. Based on the identified current broadcast status and the number of characters in the unsynthesized audio frame, a silent segment configuration strategy can be used to determine whether a silent segment needs to be inserted into the current synthesized audio frame. The silent segment configuration strategy can be used to configure the current synthesized audio frame sent by the synthesis engine according to the silent segment configuration strategy, thereby achieving fine control of the playback rhythm and making the pauses of the required audio frames to be broadcast consistent with human speech habits, avoiding discordant playback due to missing content, improving the naturalness and smoothness of the audio broadcast, ensuring the real-time nature of the audio broadcast, and thus enhancing the user experience. The silent segment configuration strategy here includes a strategy for adding silent segments or a strategy for not adding silent segments. Accordingly, the silent segment configuration includes adding silent segments or not adding silent segments.
[0084] Step 140: Push the audio frame to be broadcast to an audio processing end, and the audio processing end is used to broadcast the audio frame to be broadcast.
[0085] Optionally, after the audio frame to be broadcast is acquired, the audio frame to be broadcast may be pushed to the audio processing end so that the audio processing end broadcasts the audio frame to be broadcast using a broadcast voice.
[0086] The method provided in this embodiment, after receiving the last synthesized audio frame sent by the synthesis engine, refreshes the service push information in real time according to the push time node, push time length and current time node of the last synthesized audio frame, and based on the current service push information refreshed in real time and the phoneme chain constructed by the phoneme text information synthesized and returned by the synthesis engine, performs dynamic adjustment of the synthesis speed at the phoneme level and dynamic configuration of the silent segments in a refined and intelligent manner, thereby achieving more refined and intelligent dynamic optimization of the audio broadcast strategy, so as to more effectively deal with the adverse effects brought about by network delays or uneven data frame transmission, thereby improving the overall fluency and real-time performance of the audio broadcast and enhancing the user experience.
[0087] In some embodiments, step 110 specifically includes:
[0088] Summing the push time node and the push time length of the previous synthesized audio frame to obtain a first target time node;
[0089] According to the comparison result between the current push time node and the first target time node, the service push information is refreshed to obtain the current service push information.
[0090] Optionally, during the service push information refresh process, the broadcast monitor can record two key pieces of information during the audio push process: the push time node P and the push duration Dp of the previous synthesized audio frame. When the audio frame is pushed at the current push time node N, the service push information can be refreshed using the following refresh steps to obtain the current service push information:
[0091] First, sum the push time node P and the push time length Dp to get the first target time node ,in, , this time node Represents the expected time point when the previous synthesized audio frame has been fully played.
[0092] In addition, the current push time node N is obtained. And by comparing the current push time node N with the first target time node , to determine whether the current push time node is within the time window of the previous synthesized audio frame playback, or whether it has exceeded this time window.
[0093] Based on the comparison and judgment results, the current broadcast status is obtained, and based on the current broadcast status, it is determined whether the current audio broadcast has not been stuck. Based on whether the current audio broadcast has not been stuck, different service push information refresh strategies are determined, and then the service push information is refreshed according to the corresponding service push information refresh strategy to obtain the current service push information. If the current audio broadcast has not been stuck, the service push information refresh strategy is to keep part of the content in the previous service push information unchanged and refresh the other part; if the current audio broadcast has been stuck, the service push information refresh strategy is to refresh all the content in the previous service push information to ensure that the subsequent audio broadcast can be closely connected to reduce stuck or delays.
[0094] The method provided in this embodiment sums the push time node and push time length of the previous synthesized audio frame and compares it with the current push time node, so as to be able to judge the current audio playback status in real time and accurately, and refresh the service push information accordingly, thereby ensuring that the audio broadcast can be closely connected and smooth, and effectively avoiding audio playback pauses or discontinuities caused by network delays or uneven data frame transmission.
[0095] In some embodiments, refreshing the service push information according to the comparison result between the current push time node and the first target time node to obtain the current service push information includes:
[0096] When the current push time node is less than or equal to the first target time node, refreshing the service push information according to the first refresh strategy to obtain the current service push information;
[0097] When the current push time node is greater than the first target time node, refreshing the service push information according to the second refresh strategy to obtain the current service push information;
[0098] Among them, the first refresh strategy includes a strategy of keeping the push time node of the previous synthesized audio frame unchanged, and a strategy of updating the push time length of the previous synthesized audio frame to the sum of the push time length of the previous synthesized audio frame and the current push time length; the second refresh strategy includes a strategy of updating the push time node of the previous synthesized audio frame to the current push time node, and a strategy of updating the push time length of the previous synthesized audio frame to the current push time length.
[0099] Figure 4 This is one of the distribution diagrams of the audio broadcast time provided by the present invention; Figure 5 This is the second schematic diagram of the distribution of audio broadcast time provided by the present invention; Figure 6 This is the third distribution diagram of the audio broadcast time provided by the present invention; Figure 7 This is the fourth distribution diagram of the audio broadcast time provided by the present invention.
[0100] Optionally, the current service push information refresh strategy further includes:
[0101] like Figure 4 As shown, if the current push time node N is earlier than the first target time node , that is, N P+Dp, the current audio frame broadcast is not stuck, the current network status or data frame transmission is stable, and there is no need to adjust the time node. Therefore, at this time, the service push information is refreshed according to the first refresh strategy, specifically keeping the push time node of the last synthesized audio frame. The push time length of the previous synthetic audio frame remains unchanged, and the push time length of the previous synthetic audio frame is updated to the sum of the push time length of the previous synthetic audio frame and the current push time length, that is, Dp=Dp+Dn. The distribution of the audio broadcast time in the current service push information updated accordingly is as follows: Figure 5 shown.
[0102] like Figure 6 As shown, if the current push time node N is later than the first target time node , that is, N P+Dp, the current audio frame broadcast has been stuck, and the playback of the previous synthesized audio frame has ended. It is necessary to calculate the new push time node and push time length from the current time, that is, the time node and push time length need to be adjusted synchronously. Therefore, at this time, the service push information is refreshed according to the second refresh strategy. Specifically, the push time node of the previous synthesized audio frame is updated to the current push time node, that is, P=N, and the push time length of the previous synthesized audio frame is updated to the current push time length, that is, Dp=Dn. The distribution of audio broadcast time in the current service push information updated accordingly is as follows: Figure 7 The current push time node and current push duration here are predicted and set based on factors such as the current network conditions, data frame size, and playback speed.
[0103] The method provided in this embodiment dynamically refreshes the push time node and push time length in the service push information by adopting a corresponding refresh strategy based on the comparison result between the current push time node and the first target time node, thereby ensuring that the system accurately tracks the broadcast status in real time, so that in the subsequent broadcast process, the status of the audio broadcast can be more accurately judged, and the broadcast strategy can be adjusted accordingly, thereby ensuring that the audio broadcast can be closely connected and smooth, effectively avoiding problems such as audio playback pauses, delays or jumps.
[0104] In some embodiments, the method further comprises:
[0105] Summing the first target time node and the broadcast freeze time threshold to obtain a second target time node;
[0106] Summing the first target time node and the broadcast redundancy time threshold to obtain a third target time node;
[0107] When the current push time node is greater than the first target time node, if it is determined that the current push time node is less than or equal to the second target time node, it is determined that the second refresh policy further includes a policy of updating the broadcast stutter identifier to true;
[0108] When the current push time node is greater than the first target time node, if it is determined that the current push time node is greater than the third target time node, it is determined that the second refresh policy further includes a policy of updating the broadcast redundancy identifier to true.
[0109] The broadcast stutter time threshold here is used to measure whether there is a stutter situation during the broadcast; the broadcast redundancy time threshold is used to measure whether there is a broadcast redundancy situation during the broadcast. For the convenience of description, hereinafter, the broadcast stutter time threshold is defined as Ts (unit: millisecond ms), and the broadcast redundancy time threshold is defined as Tr (unit: millisecond ms).
[0110] Optionally, when it is determined that the current push time node is greater than the first target time node, the sum of the first target time node and the broadcast stutter time threshold can be further calculated to obtain the second target time node P + Dp + Ts, so as to judge whether there is a broadcast stutter risk during the broadcast; and the sum of the first target time node and the broadcast redundancy time threshold is calculated to obtain the third target time node P + Dp + Tr, so as to judge whether there is a broadcast redundancy risk during the broadcast.
[0111] Furthermore, in order to update the broadcast status in real time, when it is determined that the current push time node is greater than the first target time node, in addition to updating the push time node of the previous synthesized audio frame to the current push time node and updating the push time length of the previous synthesized audio frame to the current push time length, it is also necessary to further judge whether the current push time node is less than or equal to the second target time node, or greater than the third target time node.
[0112] If it is further determined that the current push time node is less than or equal to the second target time node, that is, P + Dp < N <= P + Dp + Ts, it is determined that there is a broadcast stutter risk during the broadcast. At this time, it is necessary to synchronously update the broadcast stutter identifier maintained by the broadcast monitor to true (also called True).
[0113] If it is further determined that the current push time node is greater than the third target time node, that is, N > P + Dp + Tr, it is determined that there is a broadcast redundancy risk during the broadcast. At this time, it is necessary to synchronously update the broadcast redundancy identifier maintained by the broadcast monitor to true.
[0114] The method provided in this embodiment monitors the front-end playback status when pushing audio frames by setting a broadcast freeze time threshold and a broadcast redundancy time threshold, thereby realizing real-time identification of freeze and redundancy situations, so as to facilitate subsequent adjustment of the playback strategy and dynamic update of the playback status to ensure the consistency and stability of audio content push.
[0115] In some embodiments, step 130 specifically includes:
[0116] According to the phoneme chain, the number of characters of the currently received synthesized audio frame is updated;
[0117] Determining the number of characters in the currently sent unsynthesized audio frame based on the number of characters in the currently received synthesized audio frame;
[0118] When it is determined that the number of characters of the currently sent unsynthesized audio frame is greater than or equal to a preset number, and when it is determined based on the current service push information that the current broadcast state is a broadcast redundant state, the synthesis speed of the synthesis engine is refreshed according to the preset speed to obtain the current synthesis speed.
[0119] Optionally, after obtaining the phoneme chain, the currently received synthesized content can be refreshed based on the phoneme chain, thereby obtaining the number of characters in the currently received synthesized audio frame. The specific implementation logic includes: performing content matching based on the phoneme chain to determine the end of the corresponding character returned by the previous synthesized audio frame, thereby locating the position of the characters of the currently synthesized content, and then determining the currently received synthesized content based on the position of the characters of the currently synthesized content, and determining the number of characters in the currently received synthesized audio frame based on the number of characters contained in the currently received synthesized content. It should be noted that since there is no punctuation information in the phoneme information, a certain amount of redundancy can be designed when performing content matching. If the currently matched character is a correct match, backward retrieval is allowed; if the redundancy is reached and there is still no correct match, the character is ignored and content matching is performed after the next audio frame is received.
[0120] After obtaining the number of characters of the currently received synthesized audio frame, the difference between the number of characters of the currently sent content to be synthesized and the number of characters of the currently received synthesized audio frame can be calculated to obtain the number of characters of the currently sent unsynthesized audio frame.
[0121] After obtaining the number of characters in the currently transmitted unsynthesized audio frame, the number of characters in the currently transmitted unsynthesized audio frame is compared with a preset number, and a determination is made based on the current service push information whether the current broadcast status is a broadcast redundant state. Determining the current broadcast status based on the current service push information may involve determining whether the current broadcast status is a broadcast redundant state based on updated information about the push time node and push duration in the current service push information, or determining whether the current broadcast status is a broadcast redundant state based on a broadcast redundant identifier, which is specifically defined in this embodiment.
[0122] If it is determined that the number of characters in the currently sent unsynthesized audio frames is greater than or equal to the preset number, and the current broadcast status is a broadcast redundant state, it indicates that there is sufficient playback redundancy on the current front end and there is a lot of unsynthesized content. At this time, the synthesis parameters of the current queue of the synthesis engine can be refreshed through the specified synthesis speed, that is, the preset speed (accelerated speed) to obtain the current synthesis speed.
[0123] The method provided in this embodiment dynamically obtains the number of characters of the currently sent unsynthesized audio frame through the phoneme chain, and dynamically obtains the current broadcast status in real time through the current service push information, and dynamically refreshes the synthesis speed in combination with the number of characters of the currently sent unsynthesized audio frame and the current broadcast status, so as to achieve more refined and intelligent dynamic optimization of the audio broadcast strategy, thereby improving the global smoothness and real-time performance of the audio broadcast and enhancing the user experience.
[0124] In some embodiments, step 140 specifically includes:
[0125] If it is determined that the number of characters in the currently sent unsynthesized audio frame is less than the preset number, and if it is determined according to the current service push information that the current broadcast state is a broadcast blocked state, matching and obtaining a target phoneme frame having a rhythmic pause node in the currently synthesized audio frame based on the rhythmic pause identifiers of each phoneme frame in the phoneme chain;
[0126] According to the rhythmic pause configuration position corresponding to the target phoneme frame, a silent segment is inserted into the current synthesized audio frame to obtain the audio frame to be broadcasted.
[0127] Optionally, after obtaining the number of characters in the currently sent unsynthesized audio frame based on the phoneme chain, the number of characters in the currently sent unsynthesized audio frame is compared with a preset number, and a determination is made based on the current service push information whether the current broadcast state is a broadcast-blocked state. The method for determining the current broadcast state based on the current service push information herein may be to determine whether the current broadcast state is a broadcast-blocked state based on updated information regarding the push time node and push time length in the current service push information, or to determine whether the current broadcast state is a broadcast-blocked state based on a broadcast redundancy identifier and a broadcast jam identifier, which are specifically defined in this embodiment.
[0128] If it is determined that the number of characters in the currently sent unsynthesized audio frame is less than the preset number and the current broadcast status is a broadcast congestion state, it indicates that the current front-end is about to or has been blocked in playback and there is insufficient unsynthesized content. In order to ensure user experience, a silent segment needs to be inserted at an appropriate position to avoid broadcast interruption. The specific implementation steps are as follows:
[0129] According to the prosodic pause identifier of each phoneme frame in the phoneme chain, a target phoneme frame having a prosodic pause node is matched and obtained in the current synthesized audio frame.
[0130] According to the rhythmic pause configuration position corresponding to the target phoneme frame, a silent segment of a preset length (the specific length can be adjusted according to user preferences or system settings) is inserted at the corresponding position of the current synthesized audio frame to generate a new audio frame containing the silent segment, thereby obtaining the audio frame to be broadcast.
[0131] The method provided in this embodiment dynamically determines, based on the phoneme chain and the current service push information, that when the current front-end playback detects that playback congestion is about to or has occurred and there is insufficient synthesized content, it dynamically determines the rhythmic pause node through the phoneme chain to rationally insert silence, so that when playback is blocked, the playback content can be rationally segmented at the rhythmic pause nodes, thereby achieving fine control of the playback rhythm, ensuring the playback rhythm and naturalness of the synthesized audio, and ensuring that the pauses and playback rhythm of the synthesized audio are consistent with human voice habits, avoiding disharmonious playback and abrupt playback caused by missing content, thereby improving the user experience.
[0132] In some embodiments, the method further comprises:
[0133] In the case that the last synthesized audio frame is not received, obtaining the number of characters of the currently sent unsynthesized audio frame and the length of the currently pushed but unbroadcasted audio frame;
[0134] Obtaining a current broadcast state according to the number of characters of the currently sent unsynthesized audio frame and the length of the currently pushed but unbroadcasted audio frame;
[0135] When it is determined that the current broadcast state is a broadcast redundant state, the synthesis speed of the synthesis engine is accelerated to obtain the current synthesis speed.
[0136] Alternatively, in the broadcast process, if the last synthesized audio frame is not received, that is, before the synthesized frame is delivered, then according to the interactive information between the broadcast monitor and the synthesis engine, determine the number of characters of the currently sent unsynthesized audio frame, and according to the interactive information between the broadcast monitor and the audio processor, determine the length of the currently unreported audio frame that has been pushed. Thus, by combining the number of characters of the currently sent unsynthesized audio frame and the length of the currently unreported audio frame that has been pushed, it is judged whether the current broadcast state is the broadcast redundant state. As, the number of characters of the currently sent unsynthesized audio frame is compared with a quantity threshold value, the length of the currently unreported audio frame that has been pushed is compared with a length threshold value, and when the number of characters of the currently sent unsynthesized audio frame exceeds the quantity threshold value, and the length of the currently unreported audio frame that has been pushed exceeds the length threshold value, then determine that the current broadcast state is the broadcast redundant state.
[0137] If the current broadcast state is determined to be redundant based on the above determination, synthesis acceleration is triggered to speed up the synthesis engine, thereby effectively reducing playback wait time and ensuring the continuity and smoothness of the broadcast process. Specifically, acceleration can be achieved by adjusting the internal parameters of the synthesis engine, such as increasing the number of parallel processing threads or optimizing algorithms, to improve synthesis efficiency.
[0138] The following describes the audio broadcast method provided by this embodiment using a simultaneous voice interpretation scenario as an example.
[0139] like Figure 2 As shown, in the scenario of simultaneous voice interpretation, the specific steps of the audio broadcast method include:
[0140] The broadcast monitor receives the data to be synthesized sent by the upstream audio processing end, such as the deterministic translation result; the broadcast monitor controls the synthesis speed and sends the synthesis speed and the data to be synthesized to the downstream synthesis engine; the broadcast monitor receives the audio frames synthesized and returned by the downstream synthesis engine; the broadcast monitor performs backlog judgment to control the push of audio frames, that is, sends the audio frames to be broadcast to the audio processing end for audio broadcast.
[0141] Among them, before receiving the audio frame synthesized and returned by the synthesis engine, the broadcast monitor can determine the number of characters in the currently sent unsynthesized audio frame and the length of the currently pushed unbroadcasted audio frame to determine whether the current broadcast state is a broadcast redundant state. If so, synthesis acceleration is triggered to speed up the synthesis speed.
[0142] When receiving the audio frames synthesized by the synthesis engine, the main logic executed by the broadcast monitor includes: refreshing the current service push information, constructing a phoneme chain, and refreshing the currently received synthesized content.
[0143] For the step of refreshing the current service push information, its specific implementation logic includes: recording the push time node P and the push time length Dp of the previous synthesized audio frame, obtaining the current push time node N, and the push time length Dn; if the current push time node N is earlier than P + Dp, that is, N <= P + Dp, it indicates that there is no lag in the current front-end audio, and the refreshed push information is: P remains unchanged, Dp = Dp + Dn; if the current time node N is later than P + Dp, that is, N > P + Dp, the current front-end audio playback has experienced lag, and the refreshed push information is: P = N, Dp = Dn; if P + Dp < N <= P + Dp + Ts, the refreshed push information is: set the broadcast lag identifier maintained by the broadcast monitor to True; if N > P + Dp + Tr, set the broadcast redundancy identifier maintained by the broadcast monitor to True.
[0144] For the step of constructing a phoneme chain, its specific implementation logic includes: initializing a queue; parsing the phoneme text information synthesized and returned by the synthesis engine to obtain the phoneme information of each character; splicing the phoneme information of each character in the phoneme text information in the way of splicing different syllables of the same character to obtain the phoneme frames corresponding to each character; establishing links between the phoneme frames corresponding to each character according to the sorting order of each character in the phoneme text information to obtain an initial phoneme chain. And counting the pronunciation duration of each character to obtain the pronunciation duration of the phoneme frames corresponding to each character; accumulating the total synthesis time-consuming information starting from the first frame to obtain the synthesis time-consuming information of the phoneme frames corresponding to each character; according to the prosodic pause configuration information agreed upon in the synthesis, parsing and obtaining the identifiers that are allowed to be inserted as pauses and silences in the phoneme frames corresponding to each character, that is, the prosodic pause identifiers. Finally, configuring the pronunciation duration, synthesis time-consuming information, and prosodic pause identifiers of the phoneme frames corresponding to each character in the initial phoneme chain can obtain the final complete phoneme chain.
[0145] For the step of refreshing the currently received synthesized content, the specific implementation logic includes: performing content matching based on the phoneme chain to determine the character ending corresponding to the previous synthesized audio frame return, so as to locate the position of the character where the currently synthesized content is located, and then determining the currently received synthesized content based on the position of the character where the currently synthesized content is located. It should be noted that since there is no punctuation information in the phoneme information, when performing content matching, a certain amount of redundancy can be designed. If the currently matched character is a correct match, backward retrieval is allowed; if the correct match is not achieved after reaching the redundancy amount, the character is ignored, and content matching is performed again after receiving the next audio frame.
[0146] Among them, when receiving the audio frames synthesized and returned by the synthesis engine and the synthesis content and synthesis speed need to be sent, if the broadcast monitor identifies and determines that there is sufficient playback redundancy at the current front end and there is a lot of unsynthesized content based on the refreshed currently received synthesis content and the refreshed current service push information, it will refresh the synthesis parameters of the current queue of the synthesis engine by specifying the synthesis speed.
[0147] Among them, when receiving the audio frame synthesized and returned by the synthesis engine and when the audio frame needs to be pushed, the broadcast monitor identifies and determines that the current front end is about to or has already encountered playback congestion and insufficient unsynthesized content based on the refreshed currently received synthesis content and the refreshed current service push information. If it is further identified based on the phoneme chain that there is a target phoneme frame with a rhythmic pause node in the current synthesized audio frame, the silent segment difference is actively performed at the rhythmic pause configuration position corresponding to the target phoneme frame, so that the synthetic audio congestion played by the front end is optimized through the silent segment at a reasonable rhythmic pause point, thereby ensuring the global smoothness of the audio playback.
[0148] In summary, the method provided in this embodiment, through the dynamic control of the broadcast monitor and the fine management of the phoneme chain, can perform dynamic adjustment of the synthesis speed at the phoneme level and dynamic configuration of silent segments in a refined and intelligent manner, thereby achieving more refined and intelligent dynamic optimization of the audio broadcast strategy, so as to more effectively deal with the adverse effects of network delays or uneven data frame transmission, thereby improving the overall smoothness and real-time performance of the audio broadcast and enhancing the user experience.
[0149] The audio broadcast device provided by the present invention is described below. The audio broadcast device described below and the audio broadcast method described above can be referenced to each other.
[0150] Figure 8 Schematic diagram of the structure of the audio broadcast device provided by the present invention; Figure 8 As shown, the device includes:
[0151] The first refreshing unit 810 is configured to, upon receiving a previous synthesized audio frame sent by the synthesis engine, refresh the service push information according to the push time node and push time length of the previous synthesized audio frame and the current push time node, thereby obtaining the current service push information;
[0152] The second refreshing unit 820 is used to refresh the synthesis speed of the synthesis engine according to the phoneme chain and the current service push information to obtain the current synthesis speed; the phoneme chain is constructed based on the phoneme text information synthesized and returned by the synthesis engine;
[0153] The audio configuration unit 830 is configured to configure a silent segment for the current synthesized audio frame sent by the synthesis engine according to the phoneme chain and the current service push information, thereby obtaining an audio frame to be broadcast; the current synthesized audio frame is synthesized by the synthesis engine according to the current synthesis speed;
[0154] The audio broadcast unit 840 is used to push the audio frame to be broadcast to the audio processing end, and the audio processing end is used to broadcast the audio frame to be broadcast.
[0155] The device provided in this embodiment, after receiving the last synthesized audio frame sent by the synthesis engine, refreshes the service push information in real time according to the push time node, push time length and current time node of the last synthesized audio frame, and based on the current service push information refreshed in real time and the phoneme chain constructed by the phoneme text information synthesized and returned by the synthesis engine, performs dynamic adjustment of the synthesis speed at the phoneme level and dynamic configuration of the silent segments in a refined and intelligent manner, thereby achieving more refined and intelligent dynamic optimization of the audio broadcast strategy, so as to more effectively deal with the adverse effects of network delays or uneven data frame transmission, thereby improving the overall fluency and real-time performance of the audio broadcast and enhancing the user experience.
[0156] In some embodiments, the first refresh unit is specifically configured to:
[0157] Summing the push time node and the push time length of the previous synthesized audio frame to obtain a first target time node;
[0158] According to the comparison result between the current push time node and the first target time node, the service push information is refreshed to obtain the current service push information.
[0159] In some embodiments, the first refresh unit is further configured to:
[0160] When the current push time node is less than or equal to the first target time node, refreshing the service push information according to the first refresh strategy to obtain the current service push information;
[0161] When the current push time node is greater than the first target time node, refreshing the service push information according to the second refresh strategy to obtain the current service push information;
[0162] Among them, the first refresh strategy includes a strategy of keeping the push time node of the previous synthesized audio frame unchanged, and a strategy of updating the push time length of the previous synthesized audio frame to the sum of the push time length of the previous synthesized audio frame and the current push time length; the second refresh strategy includes a strategy of updating the push time node of the previous synthesized audio frame to the current push time node, and a strategy of updating the push time length of the previous synthesized audio frame to the current push time length.
[0163] In some embodiments, the first refresh unit is further configured to:
[0164] Summing the first target time node and the broadcast freeze time threshold to obtain a second target time node;
[0165] Summing the first target time node and the broadcast redundancy time threshold to obtain a third target time node;
[0166] In a case where the current push time node is greater than the first target time node, if it is determined that the current push time node is less than or equal to the second target time node, determining that the second refresh strategy also includes a strategy of updating the broadcast jam identifier to be true;
[0167] In the case that the current push time node is greater than the first target time node, if it is determined that the current push time node is greater than the third target time node, it is determined that the second refresh strategy also includes a strategy of updating the broadcast redundant identifier to be true.
[0168] In some embodiments, the second refresh unit is specifically configured to:
[0169] According to the phoneme chain, the number of characters of the currently received synthesized audio frame is updated;
[0170] Determining the number of characters in the currently sent unsynthesized audio frame based on the number of characters in the currently received synthesized audio frame;
[0171] When it is determined that the number of characters of the currently sent unsynthesized audio frame is greater than or equal to a preset number, and when it is determined based on the current service push information that the current broadcast state is a broadcast redundant state, the synthesis speed of the synthesis engine is refreshed according to the preset speed to obtain the current synthesis speed.
[0172] In some embodiments, the audio configuration unit is specifically configured to:
[0173] If it is determined that the number of characters in the currently sent unsynthesized audio frame is less than the preset number, and if it is determined according to the current service push information that the current broadcast state is a broadcast blocked state, matching and obtaining a target phoneme frame having a rhythmic pause node in the currently synthesized audio frame based on the rhythmic pause identifiers of each phoneme frame in the phoneme chain;
[0174] According to the rhythmic pause configuration position corresponding to the target phoneme frame, a silent segment is inserted into the current synthesized audio frame to obtain the audio frame to be broadcasted.
[0175] In some embodiments, the device further includes a phoneme chain construction unit, specifically configured to:
[0176] Parsing the phoneme text information to obtain phoneme information of each word;
[0177] splicing the phoneme information of each word in the phoneme text information in a manner of splicing different syllables of the same word to obtain a phoneme frame corresponding to each word;
[0178] According to the sorting order of each word in the phoneme text information, links are established between the phoneme frames corresponding to each word to obtain an initial phoneme chain;
[0179] The pronunciation duration, synthesis time-consuming information and rhythmic pause identifier of the phoneme frame corresponding to each word are configured in the initial phoneme chain to obtain the phoneme chain.
[0180] In some embodiments, the second refresh unit is further configured to:
[0181] In the case that the last synthesized audio frame is not received, obtaining the number of characters of the currently sent unsynthesized audio frame and the length of the currently pushed but unbroadcasted audio frame;
[0182] Obtaining a current broadcast state according to the number of characters of the currently sent unsynthesized audio frame and the length of the currently pushed but unbroadcasted audio frame;
[0183] When it is determined that the current broadcast state is a broadcast redundant state, the synthesis speed of the synthesis engine is accelerated to obtain the current synthesis speed.
[0184] The device provided by the present invention is used to execute the above-mentioned method embodiments. Please refer to the above-mentioned embodiments for the specific processes and detailed contents, which will not be repeated here.
[0185] In some embodiments, this embodiment further provides an audio broadcast system, which includes an audio processing terminal, a broadcast monitor, and a synthesis engine.
[0186] The audio processing end is used to convert the voice initiated by the user into data to be synthesized; and to broadcast the audio frames to be broadcasted pushed by the broadcast monitor in real time; the synthesis engine is used to synthesize the audio frames according to the synthesis speed and the data to be synthesized sent by the broadcast monitor, and return the synthesized audio frames to the broadcast monitor; the broadcast monitor is used to execute the audio broadcast method provided by the above embodiments, which can be found in detail. Figure 1 The execution process shown is not repeated here.
[0187] The system provided in this embodiment can perform fine-grained and intelligent dynamic adjustment of the synthesis speed at the phoneme level and dynamic configuration of silent segments through real-time interaction between the broadcast monitor and the audio processing end and the synthesis engine, so as to achieve more fine-grained and intelligent dynamic optimization of the audio broadcast strategy, so as to more effectively deal with the adverse effects of network delays or uneven data frame transmission, thereby improving the overall smoothness and real-time performance of the audio broadcast and enhancing the user experience.
[0188] Figure 9 An example of a physical structure diagram of an electronic device is shown below. Figure 9 As shown, the electronic device may include: a processor 910 , a communication interface 920 , a memory 930 and a communication bus 940 , wherein the processor 910 , the communication interface 920 and the memory 930 communicate with each other via the communication bus 940 . The processor 910 can call the logic instructions in the memory 930 to execute the audio broadcast method, which includes: when receiving the last synthesized audio frame sent by the synthesis engine, refreshing the service push information according to the push time node and push time length of the last synthesized audio frame, and the current push time node to obtain the current service push information; refreshing the synthesis speed of the synthesis engine according to the phoneme chain and the current service push information to obtain the current synthesis speed; the phoneme chain is constructed according to the phoneme text information synthesized and returned by the synthesis engine; according to the phoneme chain and the current service push information, configuring the current synthesized audio frame sent by the synthesis engine with a silent segment to obtain the audio frame to be broadcast; the current synthesized audio frame is synthesized by the synthesis engine according to the current synthesis speed; pushing the audio frame to be broadcast to the audio processing end, and the audio processing end is used to broadcast the audio frame to be broadcast.
[0189] Furthermore, the logic instructions in the aforementioned memory 930 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product, stored in a storage medium, includes instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0190] On the other hand, the present invention also provides a computer program product, which includes a computer program, which can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the audio broadcast method provided by the above-mentioned methods, which includes: upon receiving the last synthesized audio frame sent by the synthesis engine, refreshing the service push information according to the push time node and push time length of the last synthesized audio frame, as well as the current push time node, to obtain the current service push information; refreshing the synthesis speed of the synthesis engine according to the phoneme chain and the current service push information to obtain the current synthesis speed; the phoneme chain is constructed according to the phoneme text information synthesized and returned by the synthesis engine; according to the phoneme chain and the current service push information, configuring a silent segment for the current synthesized audio frame sent by the synthesis engine to obtain an audio frame to be broadcast; the current synthesized audio frame is synthesized by the synthesis engine according to the current synthesis speed; pushing the audio frame to be broadcast to the audio processing end, and the audio processing end is used to broadcast the audio frame to be broadcast.
[0191] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to execute the audio broadcast method provided by the above-mentioned methods, the method comprising: upon receiving the last synthesized audio frame sent by the synthesis engine, refreshing the service push information according to the push time node and push time length of the last synthesized audio frame, as well as the current push time node, to obtain the current service push information; refreshing the synthesis speed of the synthesis engine according to the phoneme chain and the current service push information, to obtain the current synthesis speed; the phoneme chain is constructed according to the phoneme text information synthesized and returned by the synthesis engine; configuring a silent segment for the current synthesized audio frame sent by the synthesis engine according to the phoneme chain and the current service push information, to obtain an audio frame to be broadcast; the current synthesized audio frame is synthesized by the synthesis engine according to the current synthesis speed; pushing the audio frame to be broadcast to the audio processing end, which is used to broadcast the audio frame to be broadcast.
[0192] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.
[0193] Through the above description of the embodiments, those skilled in the art will clearly understand that each embodiment can be implemented using software plus a necessary general-purpose hardware platform, or of course, hardware. Based on this understanding, the essence of the above technical solution, or the portion that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, or an optical disk, and includes a number of instructions for causing a computer device (such as a personal computer, server, or network device) to execute the methods described in each embodiment or certain portions of the embodiments.
[0194] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. An audio broadcast method, characterized in that: include: Upon receiving the last synthesized audio frame sent by the synthesis engine, refreshing the service push information according to the push time node and push time length of the last synthesized audio frame and the current push time node to obtain the current service push information; Refreshing the synthesis speed of the synthesis engine according to the phoneme chain and the current service push information to obtain the current synthesis speed; the phoneme chain is constructed based on the phoneme text information synthesized and returned by the synthesis engine; According to the phoneme chain and the current service push information, a silent segment is configured for the current synthesized audio frame sent by the synthesis engine to obtain an audio frame to be broadcast; the current synthesized audio frame is synthesized by the synthesis engine according to the current synthesis speed; The audio frame to be broadcast is pushed to an audio processing end, and the audio processing end is used to broadcast the audio frame to be broadcast.
2. The audio broadcast method according to claim 1, wherein: The refreshing of the service push information according to the push time node and push time length of the previous synthesized audio frame and the current push time node to obtain the current service push information includes: Summing the push time node and the push time length of the previous synthesized audio frame to obtain a first target time node; According to the comparison result between the current push time node and the first target time node, the service push information is refreshed to obtain the current service push information.
3. The audio broadcast method according to claim 2, characterized in that: The refreshing of the service push information according to the comparison result between the current push time node and the first target time node to obtain the current service push information includes: When the current push time node is less than or equal to the first target time node, refreshing the service push information according to the first refresh strategy to obtain the current service push information; When the current push time node is greater than the first target time node, refreshing the service push information according to the second refresh strategy to obtain the current service push information; Among them, the first refresh strategy includes a strategy of keeping the push time node of the previous synthesized audio frame unchanged, and a strategy of updating the push time length of the previous synthesized audio frame to the sum of the push time length of the previous synthesized audio frame and the current push time length; the second refresh strategy includes a strategy of updating the push time node of the previous synthesized audio frame to the current push time node, and a strategy of updating the push time length of the previous synthesized audio frame to the current push time length.
4. The audio broadcast method according to claim 3, characterized in that: The method further comprises: Summing the first target time node and the broadcast freeze time threshold to obtain a second target time node; Summing the first target time node and the broadcast redundancy time threshold to obtain a third target time node; In a case where the current push time node is greater than the first target time node, if it is determined that the current push time node is less than or equal to the second target time node, determining that the second refresh strategy also includes a strategy of updating the broadcast jam identifier to be true; In the case that the current push time node is greater than the first target time node, if it is determined that the current push time node is greater than the third target time node, it is determined that the second refresh strategy also includes a strategy of updating the broadcast redundant identifier to be true.
5. The audio broadcasting method according to any one of claims 1 to 4, characterized in that: The step of refreshing the synthesis speed of the synthesis engine according to the phoneme chain and the current service push information to obtain the current synthesis speed includes: According to the phoneme chain, the number of characters of the currently received synthesized audio frame is updated; Determining the number of characters in the currently sent unsynthesized audio frame based on the number of characters in the currently received synthesized audio frame; When it is determined that the number of characters of the currently sent unsynthesized audio frame is greater than or equal to a preset number, and when it is determined based on the current service push information that the current broadcast state is a broadcast redundant state, the synthesis speed of the synthesis engine is refreshed according to the preset speed to obtain the current synthesis speed.
6. The audio broadcasting method according to claim 5, characterized in that: The step of configuring a silent segment for the current synthesized audio frame sent by the synthesis engine according to the phoneme chain and the current service push information to obtain an audio frame to be broadcast includes: If it is determined that the number of characters in the currently sent unsynthesized audio frame is less than the preset number, and if it is determined according to the current service push information that the current broadcast state is a broadcast blocked state, matching and obtaining a target phoneme frame having a rhythmic pause node in the currently synthesized audio frame based on the rhythmic pause identifiers of each phoneme frame in the phoneme chain; According to the rhythmic pause configuration position corresponding to the target phoneme frame, a silent segment is inserted into the current synthesized audio frame to obtain the audio frame to be broadcasted.
7. The audio broadcasting method according to any one of claims 1 to 4, characterized in that: The phoneme chain is constructed based on the following steps: Parsing the phoneme text information to obtain phoneme information of each word; splicing the phoneme information of each word in the phoneme text information in a manner of splicing different syllables of the same word to obtain a phoneme frame corresponding to each word; According to the sorting order of each word in the phoneme text information, links are established between the phoneme frames corresponding to each word to obtain an initial phoneme chain; The pronunciation duration, synthesis time-consuming information and rhythmic pause identifier of the phoneme frame corresponding to each word are configured in the initial phoneme chain to obtain the phoneme chain.
8. The audio broadcasting method according to any one of claims 1 to 4, characterized in that: The method further comprises: In the case that the last synthesized audio frame is not received, obtaining the number of characters of the currently sent unsynthesized audio frame and the length of the currently pushed but unbroadcasted audio frame; Obtaining a current broadcast state according to the number of characters of the currently sent unsynthesized audio frame and the length of the currently pushed but unbroadcasted audio frame; When it is determined that the current broadcast state is a broadcast redundant state, the synthesis speed of the synthesis engine is accelerated to obtain the current synthesis speed.
9. An audio broadcasting device, characterized in that: include: A first refreshing unit is configured to, upon receiving a last synthesized audio frame sent by the synthesis engine, refresh the service push information according to the push time node and push time length of the last synthesized audio frame and the current push time node, thereby obtaining the current service push information; A second refreshing unit is configured to refresh the synthesis speed of the synthesis engine according to the phoneme chain and the current service push information to obtain a current synthesis speed; the phoneme chain is constructed based on the phoneme text information synthesized and returned by the synthesis engine; an audio configuration unit, configured to configure a silent segment for a current synthesized audio frame sent by the synthesis engine according to the phoneme chain and the current service push information, to obtain an audio frame to be broadcast; the current synthesized audio frame is synthesized by the synthesis engine according to the current synthesis speed; The audio broadcast unit is used to push the audio frame to be broadcast to the audio processing end, and the audio processing end is used to broadcast the audio frame to be broadcast.
10. An audio broadcast system, characterized in that: Includes audio processing end, broadcast monitor and synthesis engine; The broadcast monitor is connected to the audio processing end and the synthesis engine respectively, and is used to execute the audio broadcast method as described in any one of claims 1 to 8.
11. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the audio broadcast method according to any one of claims 1 to 8 is implemented.
12. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the audio broadcast method according to any one of claims 1 to 8 is implemented.
Citation Information
Patent Citations
Method, system and device for propelling multimedia data
CN101938606A
Incremental-type speech online synthesis method based on statistic parameter model
CN102592594A