Audio playing control method, related device, equipment and storage medium
By expanding the terminal cache space, the problem of starting audio playback from the beginning was solved, allowing audio to continue playing from the paused position, thus improving the user experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-09
- Publication Date
- 2026-04-10
AI Technical Summary
In existing technology, when a user pauses and then resumes audio playback, the audio starts from the beginning again, resulting in a poor user experience.
By expanding the terminal cache space to store audio data acquired after pausing, the system ensures that audio resumes playback from the paused position when the user resumes playback.
This effectively prevents audio data loss and improves the user experience.
Smart Images

Figure CN121832874A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular to a control method, related apparatus, device, and storage medium for audio playback. Background Technology
[0002] With the rise of natural language processing models, the capabilities of voice assistants have also been significantly improved. Users can retrieve information through voice assistants, and the retrieved information will be read aloud in voice format. If a user pauses listening midway, the audio data packet acceleration strategy is easily triggered when resuming listening, resulting in a poor listening experience.
[0003] Currently, in related technologies, manufacturers deploy audio pause capabilities on the server side, meaning the pause function is implemented on the server side. Based on this, when a user triggers a pause playback request through the terminal, the server stops pushing audio data. When the user triggers a resume playback request through the terminal, the server resumes pushing audio data to the terminal.
[0004] However, the inventors discovered that the current solution has at least the following problem: if audio playback resumes after the terminal has paused, it restarts from the beginning. This causes inconvenience for users, resulting in a poor user experience. Therefore, an effective method to solve this problem is urgently needed. Summary of the Invention
[0005] This application provides an audio playback control method, related apparatus, device, and storage medium. It not only effectively prevents audio data loss, but also ensures that when a user triggers a resume playback operation, the audio resumes playback from the paused position, eliminating the need to start playback from the beginning, thereby improving the user experience.
[0006] In view of this, this application provides a method for controlling audio playback, including:
[0007] The first audio data stored in the first cache space is played, wherein the first audio data comes from the target audio data and is the audio data that has already been obtained;
[0008] In response to a pause playback operation on the audio, second audio data is obtained, wherein the second audio data originates from the target audio data and is the audio data obtained after the audio playback is paused;
[0009] Based on the second audio data, the first cache space is expanded to obtain a second cache space, wherein the second cache space includes the first cache space and the target cache space, and the second cache space is used to store the first audio data and the second audio data;
[0010] In response to a continue playback operation for audio, audio data stored in a second buffer space is played, wherein the audio data includes at least the second audio data.
[0011] This application also provides an audio playback control device, comprising:
[0012] The playback module is used to play audio from the first audio data stored in the first cache space, wherein the first audio data comes from the target audio data and is the audio data that has already been obtained;
[0013] The acquisition module is used to acquire second audio data in response to a pause playback operation for audio. The second audio data is derived from the target audio data and is the audio data acquired after the audio playback is paused.
[0014] The processing module is used to expand the first cache space according to the second audio data to obtain a second cache space, wherein the second cache space includes the first cache space and the target cache space, and the second cache space is used to store the first audio data and the second audio data;
[0015] The playback module is also used to play audio data stored in the second buffer space in response to a continued playback operation for audio, wherein the audio data includes at least the second audio data.
[0016] In one possible design, in another implementation of another aspect of the embodiments of this application, the audio playback control device further includes a sending module and a storage module;
[0017] The acquisition module is also used to acquire target voice data through a voice acquisition device before playing the first audio data stored in the first cache space.
[0018] The sending module is used to send target speech data to the server so that the server can convert the target speech data into raw text data and generate target text data based on the raw text data. The server is used to convert the target text data into target audio data.
[0019] The storage module is used to store the received audio data in the first cache space during the process of receiving the target audio data sent by the receiving server.
[0020] In one possible design, in another implementation of another aspect of the embodiments of this application, the audio playback control device is applied to the first terminal;
[0021] The storage module is also used to store the received audio data in the first cache space before playing the first audio data stored in the first cache space, during the process of receiving the target audio data sent by the receiving server.
[0022] Among them, the target audio data is obtained by converting the target text data, the target text data is generated based on the original text data, the original text data is obtained by converting the target speech data, and the target speech data is obtained by the second terminal through the speech acquisition device.
[0023] In one possible design, in another implementation of another aspect of the embodiments of this application, the audio playback control device further includes a determining module;
[0024] The acquisition module is also used to acquire a first audio frame sequence, wherein the first audio frame sequence includes at least one audio frame;
[0025] The determining module is configured to determine that the first audio frame sequence originates from the target audio data if each audio frame in the first audio frame sequence carries a source field; or,
[0026] If each audio frame in the first audio frame sequence carries a source field, and the cumulative duration of the audio frames in the first audio frame sequence is greater than or equal to a duration threshold, then the first audio frame sequence is determined to originate from the target audio data.
[0027] In one possible design, in another implementation of another aspect of the embodiments of this application,
[0028] The acquisition module is also used to acquire target voice data through a voice acquisition device before playing the first audio data stored in the first cache space.
[0029] The processing module is also used to convert the target speech data into raw text data;
[0030] The processing module is also used to generate target text data based on the original text data;
[0031] The storage module is also used to store the converted audio data in the first cache space during the process of converting the target text data into target audio data.
[0032] In one possible design, in another implementation of another aspect of the embodiments of this application,
[0033] The acquisition module is also used to acquire target voice data through a voice acquisition device before playing the first audio data stored in the first cache space.
[0034] The sending module is also used to send target voice data to the server so that the server converts the target voice data into raw text data and generates target text data based on the raw text data;
[0035] The acquisition module is also used to receive target text data sent by the server;
[0036] The storage module is also used to store the converted audio data in the first cache space during the process of converting the target text data into target audio data.
[0037] In one possible design, in another implementation of another aspect of the embodiments of this application, the audio playback control device is applied to the first terminal;
[0038] The acquisition module is also used to receive target text data sent by the server before playing the first audio data stored in the first cache space. The target text data is generated based on the original text data, which is obtained by converting the target speech data. The target speech data is obtained by the second terminal through the speech acquisition device.
[0039] The storage module is also used to store the converted audio data in the first cache space during the process of converting the target text data into target audio data.
[0040] In one possible design, in another implementation of another aspect of the embodiments of this application,
[0041] The processing module is specifically used to, during the process of acquiring the second audio data, if the remaining space of the first buffer space is less than or equal to the remaining space threshold, add a target buffer space to the first buffer space to obtain the second buffer space, wherein the size of the target buffer space is preset.
[0042] In one possible design, in another implementation of another aspect of the embodiments of this application,
[0043] The processing module is specifically used to determine the buffer space corresponding to each cycle based on the sampling rate, number of channels and bit depth of the audio when the remaining space of the first buffer space is less than or equal to the remaining space threshold during the process of acquiring the second audio data.
[0044] Based on the first cache space, the cache space corresponding to each cycle is added sequentially to obtain the second cache space, wherein the sum of the cache spaces corresponding to each cycle is equal to the target cache space.
[0045] In one possible design, in another implementation of another aspect of the embodiments of this application,
[0046] The acquisition module is specifically used to acquire a second audio frame sequence, wherein the second audio frame sequence includes at least one audio frame, and each audio frame corresponds to an audio energy;
[0047] All audio frames in the second audio frame sequence whose audio energy is greater than or equal to the audio energy threshold are used as the second audio data.
[0048] In one possible design, in another implementation of another aspect of the embodiments of this application,
[0049] The processing module is also configured to, in response to a continued audio playback operation, after playing the audio data stored in the second buffer space, release the target buffer space to obtain the first buffer space upon completion of audio playback; or,
[0050] During the audio playback process of the audio data stored in the second buffer space, the buffer space corresponding to each cycle is reduced sequentially according to the sampling rate, number of channels and bit depth of the audio until the first buffer space is obtained. The sum of the buffer spaces corresponding to each cycle is equal to the target buffer space.
[0051] In one possible design, in another implementation of another aspect of the embodiments of this application,
[0052] The determination module is also used to determine whether audio playback has ended if, before obtaining the first buffer space, the number of silent frames obtained is greater than or equal to a threshold; or, ...
[0053] If the audio energy of the detected audio frame is less than or equal to the audio energy threshold, then the audio playback is determined to have ended.
[0054] In another aspect, this application provides a computer device including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the methods described above.
[0055] Another aspect of this application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the methods described above.
[0056] Another aspect of this application provides a computer program product, including a computer program that, when executed by a processor, implements the methods described above.
[0057] As can be seen from the above technical solutions, the embodiments of this application have the following advantages:
[0058] This application provides an audio playback control method in which a terminal plays audio data stored in a first cache space. During this process, if the user triggers a pause playback operation, the terminal will still acquire the remaining second audio data from the target audio data while pausing audio playback. Based on this, the terminal can expand the first cache space to obtain a second cache space, that is, the second cache space can not only store the first audio data, but also store the second audio data acquired subsequently. On the one hand, this effectively avoids audio data loss; on the other hand, when the user triggers a resume playback operation, the audio stored in the second cache space will continue playing from the paused position, without having to start playing the audio from the beginning, thereby improving the user experience. Attached Figure Description
[0059] Figure 1 This is a schematic diagram illustrating the application of this application in a voice assistant scenario;
[0060] Figure 2 This is a schematic diagram illustrating the application of this application's embodiments in a translation software scenario;
[0061] Figure 3 This is a schematic diagram illustrating the application of this application's embodiments in a social software scenario;
[0062] Figure 4 This is a schematic diagram illustrating the application of this application's embodiments in a reading software scenario;
[0063] Figure 5 This is a flowchart illustrating an audio playback control method in an embodiment of this application;
[0064] Figure 6 This is a schematic diagram illustrating the storage of audio data in the second cache space in an embodiment of this application;
[0065] Figure 7 This is a schematic diagram of an interaction flow of the audio playback control method in an embodiment of this application;
[0066] Figure 8 This is a schematic diagram of an implementation environment for the audio playback control method in this application.
[0067] Figure 9 This is another interactive flow diagram of the audio playback control method in the embodiments of this application;
[0068] Figure 10 This is a schematic diagram of another implementation environment of the audio playback control method in this application;
[0069] Figure 11 This is another interactive flow diagram of the audio playback control method in the embodiments of this application;
[0070] Figure 12 This is a schematic diagram of another implementation environment of the audio playback control method in this application;
[0071] Figure 13 This is another interactive flow diagram of the audio playback control method in the embodiments of this application;
[0072] Figure 14 This is a schematic diagram of another implementation environment of the audio playback control method in this application;
[0073] Figure 15 This is another interactive flow diagram of the audio playback control method in the embodiments of this application;
[0074] Figure 16 This is a schematic diagram of another implementation environment of the audio playback control method in this application;
[0075] Figure 17 This is another flowchart illustrating the audio playback control method in this application embodiment;
[0076] Figure 18 This is a schematic diagram of the architecture of the audio playback control system in an embodiment of this application;
[0077] Figure 19 This is a schematic diagram of an audio playback control device in an embodiment of this application;
[0078] Figure 20 This is a schematic diagram of the structure of a terminal in an embodiment of this application. Detailed Implementation
[0079] This application provides an audio playback control method, related apparatus, device, and storage medium. It not only effectively prevents audio data loss, but also ensures that when a user triggers a resume playback operation, the audio resumes playback from the paused position, eliminating the need to start playback from the beginning, thereby improving the user experience.
[0080] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a particular order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented, for example, in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “corresponding,” and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0081] It is understood that, in the specific embodiments of this application, user permission or consent is required when data such as voice data is used in specific products or technologies. That is, before collecting user data, users may be prompted through interfaces, pop-ups, or voice prompts to indicate that their data needs to be collected. The process of collecting user data only begins after obtaining user permission or consent. In other words, all user data collected in this application is collected with the user's consent, and the collection, use, and processing of this data must comply with the relevant laws, regulations, and standards of the relevant countries and regions.
[0082] With the rise of natural language processing models, information retrieval capabilities have been further enhanced, leading to an increase in the amount of information retrieved. Since the retrieved information is then broadcast in audio format, a certain playback time is required. When listening to audio information, sometimes a pause is needed before resuming, which can easily trigger audio data acceleration strategies, thus impacting the listening experience.
[0083] Currently, some vendors in the industry deploy audio pause capabilities on the server side. When the server responds to an audio pause request, it needs to retain the context of each interaction, which brings significant maintenance costs to the vendors. Therefore, the current pause strategy of vendors on the server side is to cancel all interaction contexts when the user triggers an audio pause request. This requires the user to re-search, and after re-searching, the voice information will be played from the beginning, which is not user-friendly.
[0084] Based on this, this application provides an audio playback control method that allows users to trigger a pause operation when voice information is being played, and to continue playing at a normal speaking speed when the playback resumes, greatly improving the interactive experience.
[0085] Before introducing the specific methods of this application, the application scenarios of this application will be illustrated by example. It should be understood that the following application scenarios are merely illustrative and are not limited to these examples.
[0086] (1) Voice assistant scenario;
[0087] The terminal runs processes related to a voice assistant, which send real-time voice data collected by the microphone to these processes for processing. For easier understanding, please refer to [link to relevant documentation]. Figure 1 , Figure 1 This is a schematic diagram illustrating the application of this application in a voice assistant scenario. Figure 1 The diagram in Figure (A) shows the voice assistant interface. Users can input voice commands via microphone, such as, "What's the weather like tomorrow?" The voice assistant retrieves relevant answers based on the user's question. Figure 1 As shown in Figure (B), the terminal can play the response content. If the user clicks the pause control indicated by 101 during the playback of the response content, the terminal will pause the playback of the response content and display the following: Figure 1 The interface is shown in Figure (C). When the user clicks the playback control indicated by 102, the operation to continue playback is triggered. Thus, as shown in Figure (C)... Figure 1 As shown in Figure (D), the terminal continues to play the remaining response content.
[0088] (2) Real-time translation scenario;
[0089] The terminal runs translation software processes, which send real-time audio data captured by the microphone to these processes for processing. For better understanding, please refer to [link to relevant documentation]. Figure 2 , Figure 2 This is a schematic diagram illustrating the application of this application's embodiments in a translation software scenario. Figure 2 The diagram in Figure (A) shows the translation software interface. Users can input Chinese voice commands via microphone, such as, "The weather is so nice this year, shall we go see a movie together?" The translation software then translates the user-provided Chinese voice. Figure 2 As shown in Figure (B), the terminal can play audio translated into English. If the user clicks the pause control indicated by 201 while the English audio is playing, the terminal will pause the playback of the English audio and display the following: Figure 2 The interface is shown in Figure (C). When the user clicks the playback control indicated by 202, the operation to continue playback is triggered. Thus, as shown in Figure (C)... Figure 2 As shown in Figure (D), the terminal continues playing the remaining English audio. This enables cross-language communication between the parties.
[0090] (3) Instant messaging scenarios;
[0091] The terminal runs instant messaging (IM) software processes, and sends the voice data collected in real time by the microphone to the IM software processes for processing. For easier understanding, please refer to [link to relevant documentation]. Figure 3 , Figure 3 This is a schematic diagram illustrating the application of this application in a social software scenario. Figure 3 The diagram in Figure (A) shows the interface for a conversation between User B and User A. At this time, User B can input a voice message through the microphone, and the IM software will send User B's voice message to User A. Figure 3 Figure (B) shows the interface for a conversation between User A and User B. User A can click the voice playback control indicated by 301 to start playing the voice message sent by the user. If User A clicks the voice playback control indicated by 301 during voice playback, the terminal pauses voice playback and displays the following: Figure 3 The interface shown in Figure (C) is as follows. When the user clicks the play control indicated by 301, the terminal continues playing the remaining audio.
[0092] (4) Audio reading scenarios;
[0093] The terminal runs processes related to a reading software, which supports audio playback of articles. For easier understanding, please refer to [link / reference]. Figure 4 , Figure 4 This is a schematic diagram illustrating the application of this application's embodiments in a reading software scenario. Figure 4 Figure (A) shows the interface of the reading software. When the user clicks the audio playback control indicated by 401, the terminal begins playing the full text. If the user clicks the pause control indicated by 402 during the playback, the terminal pauses the playback and displays the following: Figure 4 The interface is shown in Figure (B). When the user clicks the playback control indicated by 403, the playback continues. The terminal then continues playing the remaining text.
[0094] It should be noted that the above application scenarios are merely examples, and the audio playback control method provided in this embodiment can also be applied to other scenarios, which are not limited here.
[0095] Based on the above introduction, the audio playback control method in this application will be described below. Please refer to [link / reference]. Figure 5 The audio playback control method in this application embodiment can be completed independently by the terminal or in cooperation with the server. The audio playback control method provided in this application includes:
[0096] S501, Play audio from the first audio data stored in the first buffer space, wherein the first audio data originates from the target audio data and is the audio data that has already been obtained;
[0097] In one or more embodiments, the first cache space is an initial storage space on the terminal side used to store audio data; that is, the terminal stores the acquired audio data in the first cache space. For example, when playing audio corresponding to the target audio data, during playback, the terminal reads the audio data already cached locally from the first cache space. This audio data refers to a portion of the audio data in the target audio data, i.e., the first audio data.
[0098] It should be noted that the cache involved in this application specifically refers to the packet buffer involved in the audio jitter buffer. The audio jitter buffer is a buffer used to store downlink audio data packets. It can predict network jitter and adjust the audio speed based on the actual received data, thus mitigating the impact of network jitter on audio playback and improving the downlink playback quality. The packet buffer, on the other hand, is a buffer that stores the data to be decoded during downlink audio transmission, and the speed-up / slow-down strategy for the audio is determined based on the size of this buffer.
[0099] S502, In response to a pause playback operation for audio, acquire second audio data, wherein the second audio data originates from the target audio data and is the audio data acquired after the audio playback is paused;
[0100] In one or more embodiments, when playing the audio corresponding to the target audio data, if the user triggers a pause operation for that audio, the terminal will pause the audio playback and begin acquiring second audio data from the target audio data. The second audio data refers to the audio data acquired after the audio playback is paused; that is, the second audio data appears after the first audio data.
[0101] It should be noted that the term "in response to" in this application refers to the conditions or states upon which the execution of an operation depends, and one or more operations that can be executed when certain conditions or states are met. These operations can be real-time or have a certain delay.
[0102] S503. Based on the second audio data, the first cache space is expanded to obtain a second cache space, wherein the second cache space includes the first cache space and the target cache space, and the second cache space is used to store the first audio data and the second audio data.
[0103] In one or more embodiments, the terminal determines whether to expand the first cache space based on the remaining capacity of the first cache space and the second audio data. The remaining capacity of the first cache space is equal to the difference between the capacity threshold of the first cache space and the used capacity. The capacity threshold largely determines the overall audio latency. When the capacity threshold increases, the overall latency also increases, but the resistance to network jitter improves. Conversely, when the capacity threshold decreases, the overall latency decreases, but the resistance to weak network conditions is poor, and playback stuttering is more likely.
[0104] Specifically, if the remaining capacity of the first cache space is insufficient to accommodate the second audio data (i.e., the remaining capacity of the first cache space is less than or equal to the remaining capacity threshold), then the first cache space needs to be expanded to obtain the second cache space. The second cache space includes the original first cache space and the newly added target cache space. The purpose of the expansion is to increase the cache capacity to meet current storage needs, ensuring that both the first and second audio data can be stored, thus preventing audio data loss due to insufficient cache space.
[0105] It should be noted that if the remaining capacity of the first cache space is sufficient to accommodate the second audio data, then the second audio data can be stored in the first cache space without requiring expansion. This approach effectively utilizes the first cache space, thereby improving resource utilization.
[0106] Understandably, in practical applications, users can pause the audio multiple times, and therefore, multiple expansion processes may be required, which is not limited here.
[0107] S504. In response to a continue playback operation for audio, audio data stored in a second buffer space is played, wherein the audio data includes at least the second audio data.
[0108] In one or more embodiments, in response to a user's command to continue playback triggered by audio, the terminal can then read audio data from the second buffer space for playback. The second buffer space stores at least the second audio data. Three scenarios involving storing audio data in the second buffer space are described below; for easier understanding, please refer to [link to relevant documentation]. Figure 6 , Figure 6 This is a schematic diagram of storing audio data in the second cache space in an embodiment of this application.
[0109] For example, please refer to Figure 6In Figure (A), when the user triggers a pause playback operation, the first audio data stored in the first buffer space has not yet finished playing, and the first buffer space still has remaining capacity. Based on this, when the terminal obtains the second audio data, it first stores the second audio data in the first buffer space. If the first buffer space has no remaining capacity, it is expanded to obtain the expanded second buffer space. The second buffer space includes the first buffer space and the newly added target buffer space. It can be seen that at this time, part of the second audio data is stored in the first buffer space, and the other part is stored in the target buffer space. When the user triggers a resume playback operation, the terminal first plays the first audio data in the first buffer space, then plays the second audio data in the first buffer space, and then plays the second audio data in the target buffer space.
[0110] For example, please refer to Figure 6 In Figure (B), when the user triggers a pause playback operation, the first audio data stored in the first buffer space has not yet finished playing, and the first buffer space has no remaining capacity. It is necessary to first expand the first buffer space and then store the acquired second audio data in the newly added target buffer space, ultimately obtaining a second buffer space that includes both the first and target buffer spaces. It can be seen that at this point, the second audio data is stored entirely in the target buffer space. When the user triggers a resume playback operation, the terminal first plays the first audio data in the first buffer space, and then plays the second audio data in the target buffer space.
[0111] For example, please refer to Figure 6 In Figure (C), when the user triggers a pause playback operation, all the first audio data stored in the first buffer space has finished playing. Therefore, when the terminal retrieves the second audio data, it first stores the second audio data in the first buffer space. If the first buffer space is insufficient to store all the second audio data, it needs to be expanded, and the second audio data is stored in the newly added target buffer space, ultimately resulting in a second buffer space that includes both the first and target buffer spaces. It can be seen that at this point, all the second audio data is stored in the second buffer space. When the user triggers a resume playback operation, the terminal plays the second audio data stored in the second buffer space.
[0112] This application provides an audio playback control method. By employing the above method, when a user initiates a pause operation, the streaming continues. Therefore, the audio downlink buffer can be increased to cache as much audio data as possible until all relevant audio data can be saved, effectively preventing audio data loss. Furthermore, when the user triggers a resume playback operation, the audio stored in the second buffer space will continue playing from the paused position, without needing to start playing from the beginning, thus improving the user experience.
[0113] Optionally, in the above Figure 5 Based on one or more corresponding embodiments, in another optional embodiment provided by this application, before playing audio from the first audio data stored in the first buffer space, the following may be included:
[0114] Acquire target voice data using a voice acquisition device;
[0115] The target speech data is sent to the server so that the server converts the target speech data into raw text data and generates target text data based on the raw text data. The server is used to convert the target text data into target audio data.
[0116] During the process of receiving target audio data from the receiving server, the received audio data is stored in the first cache space.
[0117] In one or more embodiments, a method for implementing audio streaming based on a network architecture is described. As can be seen from the foregoing embodiments, the method provided in this application is applicable to a server-client (CS) architecture. The following will describe this method in conjunction with... Figure 7 and Figure 8 This paper introduces a process for resuming audio playback after it has been paused, based on a client-server (CS) architecture.
[0118] Specifically, for ease of understanding, please refer to Figure 7 , Figure 7 This is a schematic diagram of an interactive flow of the audio playback control method in an embodiment of this application. As shown in the figure, the audio playback control method can be divided into 5 stages, namely:
[0119] (1) Audio playback preparation stage;
[0120] In step S701, the user inputs voice through a voice acquisition device (e.g., a microphone) to form target voice data. In step S702, the terminal sends the acquired target voice data to the server. In step S703, the server converts the target voice data into raw text data based on automatic speech recognition (ASR). In step S704, the server can perform relevant processing (e.g., information retrieval, or translation) based on the raw text data to obtain processed target text data (e.g., information retrieval results, or translations). In step S705, the server converts the target text data into target audio data based on text-to-speech (TTS) technology. In step S706, the server pushes the target audio data to the terminal. In step S707, the terminal stores the received target audio data, wherein the first audio data is the stored portion of the target audio data.
[0121] (2) Audio playback stage;
[0122] In step S708, the terminal plays a portion of the target audio data (i.e., the first audio data) that has been stored locally.
[0123] (3) Playback pause phase;
[0124] In step S709, the user triggers a pause playback operation for the audio. In step S710, the terminal responds to the pause playback operation and pauses audio playback. In step S711, even while playback is paused, the terminal continues to receive target audio data (i.e., second audio data) sent by the server.
[0125] (4) Cache expansion phase;
[0126] In step S712, if the first buffer space is insufficient to store the target audio data (i.e., the second audio data) obtained after the pause, the terminal begins expansion processing, eventually obtaining the second buffer space. It is evident that since the terminal continues to receive streams from the server after the pause, expanding the first buffer space allows for caching as much target audio data as possible, until all target audio data can be stored.
[0127] (5) Continue playing;
[0128] In step S713, the user triggers a continue playback operation for the audio. In step S714, the terminal responds to the continue playback operation and resumes playing the audio. Furthermore, while playing the audio, the expanded buffer space (i.e., the target buffer space) can be released.
[0129] It's important to note that in one scenario, while the server is converting the target text data into target audio data, the terminal can play the already converted audio data. This allows users to quickly start listening to the audio, reducing waiting time and enhancing the real-time nature of the interaction. In another scenario, the terminal can start playing the audio data only after the server has completely converted the target text data. Playing the audio all at once after conversion ensures audio integrity, thus providing higher-quality voice output.
[0130] based on Figure 7 The described process, further please refer to Figure 8 , Figure 8 This is a schematic diagram of an implementation environment for the audio playback control method in this application. As shown in the figure, this implementation environment includes, but is not limited to, a terminal 801, a network 802, and a server 803. The terminal 801 and the server 803 can communicate via the network 802. The network 802 uses standard communication technologies and / or protocols, typically the Internet, but can also be any network, including but not limited to Bluetooth, a local area network (LAN), a metropolitan area network (MAN), a wide area network (WAN), mobile, private networks, or any combination of virtual private networks. In some embodiments, customized or dedicated data communication technologies can be used to replace or supplement the aforementioned data communication technologies.
[0131] The terminal 801 involved in this application includes, but is not limited to, mobile phones, tablets, laptops, desktop computers, intelligent voice interaction devices, virtual reality devices, smart home appliances, vehicle terminals, and aircraft. The client is deployed on the terminal 801 and can run on the terminal 801 via a browser, a standalone application (APP), or a mini-program. The terminal 801 runs a client. The terminal 801 includes a human-computer interaction screen 8011, a voice acquisition device 8012, a speaker 8013, a processor, and a memory 8014. The human-computer interaction screen 8011 is used to display information and provide a human-computer interaction interface. The voice acquisition device 8012 is used to acquire voice data. The processor is used to respond to interactive commands and perform related processing. The speaker 8013 is used to play audio. The memory 8014 is used to provide cache space.
[0132] The server 803 involved in this application can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDN), and big data and artificial intelligence (AI) platforms.
[0133] The audio playback control method provided in this application will be described below in conjunction with the above-described implementation environment.
[0134] In step A1, the user inputs voice through the voice acquisition device provided by the terminal, for example, the voice is "What are some good suggestions for going out tomorrow", thereby generating the corresponding target voice data.
[0135] In step A2, the terminal sends the target voice data to the server via the network.
[0136] In step A3, the server converts the target speech data based on ASR technology to obtain the original text data. For example, the original text data is "What are some good suggestions for going out tomorrow?"
[0137] In step A4, the server begins information retrieval based on the original text data, that is, obtains the target text data. For example, the target text data is "Tomorrow will be a good day, suitable for outdoor activities such as hiking and camping. The temperature tomorrow will be between 28° and 34°, please take precautions against sunburn and drink plenty of water when outdoors."
[0138] In step A5, the server converts the target text data based on TTS technology to obtain the target audio data.
[0139] In step A6, the server pushes the target audio data to the terminal via the network.
[0140] In step A7, the terminal stores the received target audio data (i.e., the first audio data) into the first buffer space and plays the audio data stored in the first buffer space.
[0141] In step A8, the user triggers a pause playback operation. While pausing audio playback, the terminal continues to receive target audio data pushed by the server.
[0142] In step A9, terminal 801 begins to expand the first buffer space to obtain a second buffer space. This ensures that subsequently received target audio data (i.e., the second audio data) can be stored locally on the terminal.
[0143] In step A10, the user triggers a continue playback operation, and the terminal plays the audio data stored in the second buffer space.
[0144] Secondly, this application provides a method for audio streaming based on a network architecture. In this method, after the user inputs target voice data through a terminal, the terminal requests the server for processing. The server, with its powerful processing capabilities, can generate processing results more efficiently and accurately, and then feed them back to the user's terminal. Simultaneously, the network architecture is easily scalable and can be flexibly adjusted and upgraded according to increases in user numbers and changes in functional requirements, meeting evolving business needs.
[0145] Optionally, in the above Figure 5 Based on one or more corresponding embodiments, in another optional embodiment provided by the present application, it can be applied to a first terminal;
[0146] Before playing audio from the first audio data stored in the first buffer space, the following may also be included:
[0147] During the process of receiving target audio data from the receiving server, the received audio data is stored in the first cache space;
[0148] Among them, the target audio data is obtained by converting the target text data, the target text data is generated based on the original text data, the original text data is obtained by converting the target speech data, and the target speech data is obtained by the second terminal through the speech acquisition device.
[0149] In one or more embodiments, another method for implementing audio streaming based on a network architecture is described. As can be seen from the foregoing embodiments, the method provided in this application is applicable to a client-server architecture, and will be discussed in conjunction with the following... Figure 9 and Figure 10 This paper introduces another process for resuming audio playback after it has been paused, based on a client-server architecture.
[0150] Specifically, for ease of understanding, please refer to Figure 9 , Figure 9 This is another interactive flow diagram of the audio playback control method in this application embodiment. As shown in the figure, the audio playback control method can be divided into 5 stages, namely:
[0151] (1) Audio playback preparation stage;
[0152] In step S901, user B can input voice through the voice acquisition device to form target voice data. In step S902, the second terminal sends the acquired target voice data to the server. In step S903, the server converts the target voice data into raw text data based on ASR (Automatic Speech Recognition). In step S904, the server can perform relevant processing on the raw text data to obtain processed target text data. In step S905, the server converts the target text data into target audio data based on TTS (Text-to-Speech) technology. In step S906, the server pushes the target audio data to the first terminal. In step S907, the first terminal stores the received target audio data, wherein the first audio data is the stored portion of the target audio data.
[0153] (2) Audio playback stage;
[0154] In step S908, the first terminal plays a portion of the target audio data (i.e., the first audio data) that has been stored locally.
[0155] (3) Playback pause phase;
[0156] In step S909, user A triggers a pause playback operation for the audio. In step S910, the first terminal responds to the pause playback operation and pauses audio playback. In step S911, even while playback is paused, the first terminal continues to receive target audio data (i.e., second audio data) sent by the server.
[0157] (4) Cache expansion phase;
[0158] In step S912, if the first buffer space is insufficient to store the target audio data (i.e., the second audio data) obtained after the pause, the first terminal begins expansion processing, eventually obtaining a second buffer space. It is evident that since the first terminal will continue receiving streams from the server after the pause, expanding the first buffer space allows for caching as much target audio data as possible, until all target audio data can be stored.
[0159] (5) Continue playing;
[0160] In step S913, user A triggers a continue playback operation for the audio. In step S914, the first terminal responds to the continue playback operation and resumes playing the audio. Furthermore, while playing the audio, the expanded buffer space (i.e., the target buffer space) can be released.
[0161] based on Figure 9 The described process, further please refer to Figure 10 , Figure 10This is a schematic diagram of another implementation environment for the audio playback control method in this application. As shown in the figure, this implementation environment includes, but is not limited to, a first terminal 1001, a second terminal 1002, a network 1003, and a server 1004. The terminals and the server 1004 can communicate through the network 1003. Both the first terminal 1001 and the second terminal 1002 run client programs, and both include a human-computer interaction screen, a voice acquisition device, a speaker, a processor, and a memory.
[0162] The audio playback control method provided in this application will be described below in conjunction with the above-described implementation environment.
[0163] In step B1, user B inputs voice through the voice acquisition device provided by the second terminal to generate corresponding target voice data.
[0164] In step B2, the second terminal sends the target voice data to the server via the network.
[0165] In step B3, the server converts the target speech data based on ASR technology to obtain the original text data.
[0166] In step B4, the server begins information retrieval or translation processing based on the original text data, that is, to obtain the target text data.
[0167] In step B5, the server converts the target text data based on TTS technology to obtain the target audio data.
[0168] In step B6, the server pushes the target audio data to the first terminal via the network.
[0169] In step B7, the first terminal stores the received target audio data (i.e., the first audio data) into the first buffer space and plays the audio data stored in the first buffer space.
[0170] In step B8, user A triggers a pause playback operation. While pausing audio playback, the first terminal continues to receive the target audio data pushed by the server.
[0171] In step B9, the first terminal begins to expand the first buffer space to obtain a second buffer space. This ensures that subsequently received target audio data (i.e., the second audio data) can be stored locally on the first terminal.
[0172] In step B10, user A triggers the continue playback operation, and the first terminal plays the audio data stored in the second buffer space.
[0173] Secondly, this application provides another method for implementing audio streaming based on a network architecture. Through this method, leveraging the server's powerful processing capabilities, more accurate data processing results can be generated, thereby supporting more business scenarios and better meeting evolving business needs.
[0174] Optionally, in the above Figure 5 In addition to one or more corresponding embodiments, another optional embodiment provided in this application may further include:
[0175] Obtain a first audio frame sequence, wherein the first audio frame sequence includes at least one audio frame;
[0176] If each audio frame in the first audio frame sequence carries a source field indicating the target application, then the first audio frame sequence is determined to originate from the target audio data; or,
[0177] If each audio frame in the first audio frame sequence carries a source field for indicating the target application, and the cumulative duration of the audio frames in the first audio frame sequence is greater than or equal to a duration threshold, then the first audio frame sequence is determined to originate from the target audio data.
[0178] In one or more embodiments, two methods for distinguishing target audio data are described. As can be seen from the foregoing embodiments, some apps (e.g., choral apps) do not need to support pause and resume playback functions. Therefore, in practical applications, it is also necessary to determine the source of the received audio data to ensure that only audio data from specific sources is processed accordingly.
[0179] Specifically, during the acquisition of the first audio frame sequence, the terminal needs to check the source field of each audio frame in the first audio frame sequence. For example, in one scenario, if each audio frame in the first audio frame sequence carries a source field indicating the target application, then the first audio frame sequence is determined to belong to the target audio data. In another scenario, if each audio frame in the first audio frame sequence carries a source field indicating the target application, then it is also necessary to determine whether the cumulative duration of the audio frames in the first audio frame sequence is greater than or equal to a duration threshold (e.g., 5 seconds). If so, then the first audio frame sequence is determined to belong to the target audio data.
[0180] For easier understanding, please refer to Table 1, which shows the packet header information corresponding to the audio frames belonging to the target audio data.
[0181] Table 1
[0182]
[0183]
[0184] The source field is used to identify the source of the audio frame. For example, if the target application is "voice assistant", the source field can indicate that the audio frame belongs to the "voice assistant".
[0185] Secondly, this application provides two methods for distinguishing target audio data. By using these methods, audio frames carrying source fields are used as target audio data, enabling the processing of audio data corresponding to specific services, thereby improving the feasibility of the solution. Based on this, if the cumulative duration of audio frames reaches a duration threshold, it not only indicates that the audio has sufficient duration to support pausing, but also avoids some transmission errors. This improves the accuracy and reliability of determining the source of audio frames.
[0186] Optionally, in the above Figure 5 Based on one or more corresponding embodiments, in another optional embodiment provided by this application, before playing audio from the first audio data stored in the first buffer space, the following may be included:
[0187] Acquire target voice data using a voice acquisition device;
[0188] Convert the target speech data into raw text data;
[0189] Generate target text data based on the original text data;
[0190] During the process of converting target text data into target audio data, the converted audio data is stored in the first cache space.
[0191] In one or more embodiments, a method for implementing audio streaming based on a single-machine architecture is described. As can be seen from the foregoing embodiments, the method provided in this application is applicable to local architectures. The following will describe this method in conjunction with... Figure 11 and Figure 12 This paper introduces a process for resuming audio playback after it has been paused, based on a local architecture.
[0192] Specifically, for ease of understanding, please refer to Figure 11 , Figure 11 This is another interactive flow diagram of the audio playback control method in this application embodiment. As shown in the figure, the audio playback control method can be divided into 5 stages, namely:
[0193] (1) Audio playback preparation stage;
[0194] In step S1101, the user can input voice through the voice acquisition device provided by the terminal to form target voice data. In step S1102, the terminal converts the target voice data into raw text data based on ASR (Automatic Speech Recognition). In step S1103, the terminal can perform relevant processing on the raw text data to obtain processed target text data. In step S1104, the terminal converts the target text data into target audio data based on TTS (Text-to-Speech) technology. In step S1105, the terminal needs to store the generated target audio data, wherein the first audio data is the stored portion of the target audio data.
[0195] (2) Audio playback stage;
[0196] In step S1106, the terminal plays a portion of the target audio data (i.e., the first audio data) that has been stored locally.
[0197] (3) Playback pause phase;
[0198] In step S1107, the user triggers a pause playback operation for the audio. In step S1108, the terminal responds to the pause playback operation and pauses audio playback. In step S1109, even with playback paused, the terminal continues to generate target audio data (i.e., second audio data).
[0199] (4) Cache expansion phase;
[0200] In step S1110, if the first buffer space is insufficient to store the target audio data (i.e., the second audio data) generated after the pause, the terminal begins expansion processing, eventually obtaining the second buffer space. It is evident that since the terminal continues to generate target audio data after the pause, expanding the first buffer space allows for caching as much target audio data as possible, until all target audio data can be stored.
[0201] (5) Continue playing;
[0202] In step S1111, the user triggers a continue playback operation for the audio. In step S1112, the terminal responds to the continue playback operation and resumes playing the audio. Furthermore, while playing the audio, the expanded buffer space (i.e., the target buffer space) can be released.
[0203] It should be noted that, in one scenario, the converted audio data is played while the terminal is converting the target text data into target audio data. In another scenario, the audio data is played only after the terminal has completely converted the target text data into target audio data.
[0204] based on Figure 11The described process, further please refer to Figure 12 , Figure 12 This is a schematic diagram of another implementation environment of the audio playback control method in this application. As shown in the figure, this implementation environment includes, but is not limited to, terminal 1201. Terminal 1201 runs a client and includes a human-computer interaction screen 12011, a voice acquisition device 12012, a speaker 12013, a processor, and a memory 12014.
[0205] The audio playback control method provided in this application will be described below in conjunction with the above-described implementation environment.
[0206] In step C1, the user inputs voice through the voice acquisition device provided by the terminal to generate the corresponding target voice data.
[0207] In step C2, the terminal converts the target speech data based on ASR technology to obtain the original text data.
[0208] In step C3, the terminal begins information retrieval or translation based on the original text data, that is, it obtains the target text data.
[0209] In step C4, the terminal converts the target text data based on TTS technology to obtain the target audio data.
[0210] In step C5, the terminal stores the generated target audio data (i.e., the first audio data) into the first cache space and plays the audio data stored in the first cache space.
[0211] In step C6, the user triggers a pause playback operation, and while the terminal pauses audio playback, it continues to generate target audio data.
[0212] In step C7, the terminal begins to expand the first buffer space to obtain a second buffer space. This ensures that the subsequently generated target audio data (i.e., the second audio data) can be stored locally on the terminal.
[0213] In step C8, the user triggers a continue playback operation, and the terminal plays the audio data stored in the second buffer space.
[0214] Secondly, this application provides a method for implementing audio streaming based on a standalone architecture. Through this method, users can also play audio using the target application without an internet connection. On the one hand, this saves network transmission resources and network traffic. On the other hand, since data storage and processing are performed locally, it typically has lower latency and higher response speed, making it suitable for services with high real-time requirements.
[0215] Optionally, in the above Figure 5 Based on one or more corresponding embodiments, in another optional embodiment provided by this application, before playing audio from the first audio data stored in the first buffer space, the following may be included:
[0216] Acquire target voice data using a voice acquisition device;
[0217] Send the target speech data to the server so that the server can convert the target speech data into raw text data and generate the target text data based on the raw text data;
[0218] Receive target text data sent by the server;
[0219] During the process of converting target text data into target audio data, the converted audio data is stored in the first cache space.
[0220] In one or more embodiments, a method for implementing text streaming based on a network architecture is described. As can be seen from the foregoing embodiments, the method provided in this application is applicable to a client-server architecture. The following will describe this method in conjunction with... Figure 13 and Figure 14 This paper introduces another process for resuming audio playback after it has been paused, based on a client-server architecture.
[0221] Specifically, for ease of understanding, please refer to Figure 13 , Figure 13 This is another interactive flow diagram of the audio playback control method in this application embodiment. As shown in the figure, the audio playback control method can be divided into 5 stages, namely:
[0222] (1) Audio playback preparation stage;
[0223] In step S1301, the user inputs voice through the voice acquisition device to form target voice data. In step S1302, the terminal sends the acquired target voice data to the server. In step S1303, the server converts the target voice data into raw text data based on ASR (Automatic Speech Recognition). In step S1304, the server can perform relevant processing on the raw text data to obtain processed target text data. In step S1305, the server pushes the target text data to the terminal. In step S1306, the terminal converts the target text data into target audio data based on TTS (Text-to-Speech) technology. In step S1307, the terminal stores the generated target audio data, wherein the first audio data is the stored portion of the target audio data.
[0224] (2) Audio playback stage;
[0225] In step S1308, the terminal plays a portion of the target audio data (i.e., the first audio data) that has been stored locally.
[0226] (3) Playback pause phase;
[0227] In step S1309, the user triggers a pause playback operation for the audio. In step S1310, the terminal responds to the pause playback operation and pauses audio playback. In step S1311, even with playback paused, the terminal continues to generate target audio data (i.e., second audio data).
[0228] (4) Cache expansion phase;
[0229] In step S1312, if the first buffer space is insufficient to store the target audio data (i.e., the second audio data) obtained after the pause, the terminal begins expansion processing, eventually obtaining the second buffer space. It is evident that since the terminal continues to generate target audio data after the pause, expanding the first buffer space allows for caching as much target audio data as possible, until all target audio data can be stored.
[0230] (5) Continue playing;
[0231] In step S1313, the user triggers a continue playback operation for the audio. In step S1314, the terminal responds to the continue playback operation and resumes playing the audio. Furthermore, while playing the audio, the expanded buffer space (i.e., the target buffer space) can be released.
[0232] based on Figure 13 The described process, further please refer to Figure 14 , Figure 14 This is a schematic diagram of another implementation environment for the audio playback control method in this application. As shown in the figure, this implementation environment includes, but is not limited to, a terminal 1401, a network 1402, and a server 1403. The terminal 1401 and the server 1403 can communicate through the network 1402. The terminal 1401 runs a client and includes a human-computer interaction screen 14011, a voice acquisition device 14012, a speaker 14013, a processor, and a memory 14014.
[0233] The audio playback control method provided in this application will be described below in conjunction with the above-described implementation environment.
[0234] In step D1, the user inputs voice through the voice acquisition device provided by the terminal to generate the corresponding target voice data.
[0235] In step D2, the terminal sends the target voice data to the server via the network.
[0236] In step D3, the server converts the target speech data based on ASR technology to obtain the original text data.
[0237] In step D4, the server begins information retrieval or translation based on the original text data, that is, to obtain the target text data.
[0238] In step D5, the server sends the target text data to the terminal via the network.
[0239] In step D6, the server converts the target text data based on TTS technology to obtain the target audio data.
[0240] In step D7, the terminal stores the generated target audio data (i.e., the first audio data) into the first cache space and plays the audio data stored in the first cache space.
[0241] In step D8, the user triggers a pause playback operation. While pausing audio playback, the terminal continues to generate target audio data.
[0242] In step D9, the terminal begins to expand the first cache space to obtain a second cache space. This ensures that the subsequently generated target audio data (i.e., the second audio data) can be stored locally on the terminal.
[0243] In step D10, the user triggers a continue playback operation, and the terminal plays the audio data stored in the second buffer space.
[0244] Secondly, this application provides a method for implementing text streaming based on a network architecture. Through this method, leveraging the server's powerful processing capabilities, more accurate data processing results can be generated. Furthermore, compared to transmitting audio data from the server to the terminal, transmitting text data from the server requires lower costs.
[0245] Optionally, in the above Figure 5 Based on one or more corresponding embodiments, in another optional embodiment provided by the present application, it can be applied to a first terminal;
[0246] Before playing the first audio data stored in the first cache space, target text data sent by the server is received. The target text data is generated based on the original text data, which is obtained by converting the target speech data. The target speech data is obtained by the second terminal through the speech acquisition device.
[0247] During the process of converting target text data into target audio data, the converted audio data is stored in the first cache space.
[0248] In one or more embodiments, another method for implementing text streaming based on a network architecture is described. As can be seen from the foregoing embodiments, the method provided in this application is applicable to a client-server architecture, and will be discussed in conjunction with the following... Figure 15 and Figure 16 This paper introduces another process for resuming audio playback after it has been paused, based on a client-server architecture.
[0249] Specifically, for ease of understanding, please refer to Figure 15 , Figure 15 This is another interactive flow diagram of the audio playback control method in this application embodiment. As shown in the figure, the audio playback control method can be divided into 5 stages, namely:
[0250] (1) Audio playback preparation stage;
[0251] In step S1501, user B can input voice through a voice acquisition device (e.g., a microphone) to form target voice data. In step S1502, the second terminal sends the acquired target voice data to the server. In step S1503, the server converts the target voice data into raw text data based on ASR (Automatic Speech Recognition). In step S1504, the server can perform relevant processing based on the raw text data to obtain processed target text data. In step S1505, the server pushes the target text data to the first terminal. In step S1506, the first terminal converts the target text data into target audio data based on TTS (Text-to-Speech) technology. In step S1507, the first terminal stores the generated target audio data, wherein the first audio data is the stored portion of the target audio data.
[0252] (2) Audio playback stage;
[0253] In step S1508, the first terminal plays a portion of the target audio data (i.e., the first audio data) that has been stored locally.
[0254] (3) Playback pause phase;
[0255] In step S1509, user A triggers a pause playback operation for the audio. In step S1510, the first terminal responds to the pause playback operation and pauses audio playback. In step S1511, even while playback is paused, the first terminal continues to generate target audio data (i.e., second audio data).
[0256] (4) Cache expansion phase;
[0257] In step S1512, if the first buffer space is insufficient to store the target audio data (i.e., the second audio data) obtained after the pause, the first terminal begins expansion processing, eventually obtaining the second buffer space. It is evident that since the first terminal continues to generate target audio data after the pause, expanding the first buffer space allows for caching as much target audio data as possible, until all target audio data can be stored.
[0258] (5) Continue playing;
[0259] In step S1513, user A triggers a continue playback operation for the audio. In step S1514, the first terminal responds to the continue playback operation and resumes playing the audio. Furthermore, while playing the audio, the expanded buffer space (i.e., the target buffer space) can be released.
[0260] based on Figure 15 The described process, further please refer to Figure 16 , Figure 16 This is a schematic diagram of another implementation environment for the audio playback control method in this application. As shown in the figure, this implementation environment includes, but is not limited to, a first terminal 1601, a second terminal 1602, a network 1603, and a server 1604. The terminals and the server 1603 can communicate through the network 1603. Both the first terminal 1601 and the second terminal 1602 run client programs, and both include a human-computer interaction screen, a voice acquisition device, a speaker, a processor, and a memory.
[0261] The audio playback control method provided in this application will be described below in conjunction with the above-described implementation environment.
[0262] In step E1, user B inputs voice through the voice acquisition device provided by the second terminal to generate corresponding target voice data.
[0263] In step E2, the second terminal sends the target voice data to the server via the network.
[0264] In step E3, the server converts the target speech data based on ASR technology to obtain the original text data.
[0265] In step E4, the server begins information retrieval or translation based on the original text data, that is, to obtain the target text data.
[0266] In step E5, the server sends the target text data to the first terminal via the network.
[0267] In step E6, the first terminal converts the target text data based on TTS technology to obtain the target audio data.
[0268] In step E7, the first terminal stores the generated target audio data (i.e., the first audio data) into the first cache space and plays the audio data stored in the first cache space.
[0269] In step E8, user A triggers a pause playback operation. While pausing audio playback, the first terminal continues to generate target audio data.
[0270] In step E9, the first terminal begins to expand the first cache space to obtain a second cache space. This ensures that the subsequently generated target audio data (i.e., the second audio data) can be stored locally on the first terminal.
[0271] In step E10, user A triggers the continue playback operation, and the first terminal plays the audio data stored in the second buffer space.
[0272] Secondly, this application provides another method for implementing text streaming based on a network architecture. Through this method, leveraging the server's powerful processing capabilities, more accurate data processing results can be generated. Simultaneously, compared to transmitting audio data from the server to the terminal, transmitting text data from the server requires lower costs. Furthermore, it can support more business scenarios, which is beneficial for meeting ever-evolving business needs.
[0273] Optionally, in the above Figure 5 Based on one or more corresponding embodiments, in another optional embodiment provided by this application, the first buffer space is expanded according to the second audio data to obtain a second buffer space, which may specifically include:
[0274] During the acquisition of the second audio data, if the remaining capacity of the first buffer space is less than or equal to the remaining capacity threshold, a target buffer space is added to the first buffer space to obtain the second buffer space, wherein the size of the target buffer space is preset.
[0275] In one or more embodiments, a method for expanding the first cache space is described. As can be seen from the foregoing embodiments, during the process of acquiring the second audio data, the terminal continuously monitors whether the remaining capacity of the first cache space is sufficient to store the second audio data. That is, if the remaining capacity of the first cache space is less than or equal to the remaining capacity threshold, it indicates that the remaining capacity of the first cache space is insufficient to store the second audio data, and therefore, expansion processing is required.
[0276] Specifically, assuming the target cache space is set to 1920 kilobytes (kb), if the remaining capacity of the first cache space is less than or equal to the remaining capacity threshold, then 1920kb of cache space is directly added to the first cache space. For example, assuming the size of the first cache space is 2000kb, then the second cache space obtained after expansion is 3920kb, and the second audio data can be stored in the newly added target cache space.
[0277] It should be noted that the size of the target cache space can be flexibly set according to the business type, storage resources, etc., and is not limited here.
[0278] Secondly, this application embodiment provides a method for expanding the first cache space. Using this method, if the remaining capacity of the first cache space is less than or equal to a remaining capacity threshold, the cache space can be increased in one go. Therefore, expansion processing can be quickly achieved without performing related calculations, thereby improving expansion efficiency.
[0279] Optionally, in the above Figure 5 Based on one or more corresponding embodiments, in another optional embodiment provided by this application, the first buffer space is expanded according to the second audio data to obtain a second buffer space, which may specifically include:
[0280] During the acquisition of the second audio data, if the remaining capacity of the first buffer space is less than or equal to the remaining capacity threshold, the buffer space corresponding to each cycle is determined according to the sampling rate, number of channels and bit depth of the audio.
[0281] Based on the first cache space, the cache space corresponding to each cycle is added sequentially to obtain the second cache space, wherein the sum of the cache spaces corresponding to each cycle is equal to the target cache space.
[0282] In one or more embodiments, another method for expanding the first cache space is described. As can be seen from the foregoing embodiments, during the process of acquiring the second audio data, the terminal continuously monitors whether the remaining capacity of the first cache space is less than or equal to the remaining capacity threshold. If so, expansion processing needs to be initiated.
[0283] Specifically, the second audio data is represented as a data stream. Therefore, during the acquisition of the second audio data, the required expansion size can be calculated periodically. That is, the buffer space expansion size required for each period can be calculated as follows:
[0284] K cache+ =S1×C1×D1×T1; formula (1)
[0285] Among them, Kcache+ This indicates the size of the buffer space that needs to be expanded. S1 represents the sampling rate. C1 represents the number of channels. D1 represents the bit depth. T1 represents the period duration. The audio duration corresponding to the second audio data can be divided into several period durations. For example, a 10-second audio duration can be divided into 10 period durations, and each period duration is 1 second.
[0286] Based on this, taking a period of 1 second, a sampling rate of 48000 Hz, 2 channels, and a bit depth of 2 as an example, the required buffer space for each period can be calculated to be 192kb using formula (1). During the acquisition of the second audio data, the buffer space can be expanded by 192kb per second until the second audio data is acquired, thus obtaining the second buffer space. The second buffer space includes the first buffer space and the expanded buffer space (i.e., the target buffer space).
[0287] Secondly, this application embodiment provides another method for expanding the first cache space. Using this method, if the remaining capacity of the first cache space is less than or equal to the remaining capacity threshold, it needs to be expanded gradually based on the acquired second audio data, thereby avoiding the waste of resources caused by allocating too much cache space at once.
[0288] Optionally, in the above Figure 5 Based on one or more corresponding embodiments, in another optional embodiment provided by this application, obtaining the second audio data may specifically include:
[0289] Obtain a second audio frame sequence, wherein the second audio frame sequence includes at least one audio frame, and each audio frame corresponds to an audio energy;
[0290] All audio frames in the second audio frame sequence whose audio energy is greater than or equal to the audio energy threshold are used as the second audio data.
[0291] In one or more embodiments, a method for storing audio data based on a delay protection mechanism is described. As described in the foregoing embodiments, assuming the target audio data is 30 seconds long, and the user triggers a pause operation at the 20th second, if the streaming has already finished, the terminal will continue to acquire silence data to keep the current session alive. This silence data will also occupy buffer space. If playback resumes after 5 seconds, the terminal will store 5 seconds of silence data (i.e., 5 seconds of invalid audio data). The next round of audio data will be queued after this invalid audio data, thus causing a delay in new interactions. Therefore, this application employs a delay protection mechanism to process the audio frames acquired after the pause.
[0292] Specifically, after the user triggers a pause playback operation, a second audio frame sequence is acquired, which carries at least one audio frame. An audio frame is the basic unit of audio data, and each audio frame corresponds to an audio energy value, reflecting the volume level. If the audio energy of an audio frame is greater than or equal to an audio energy threshold (e.g., 0), then the audio frame is considered a valid audio frame. Valid audio frames include both audio frames with sound and silent audio frames. If the audio energy of an audio frame is less than the audio energy threshold (e.g., 0), then the audio frame is considered an invalid audio frame. Typically, the audio energy of an invalid audio frame can be set to a special value, such as "-1".
[0293] Based on this, all audio frames in the second audio frame sequence with audio energy greater than or equal to the audio energy threshold are used as second audio data and stored locally on the terminal. All audio frames in the second audio frame sequence with audio energy less than the audio energy threshold are discarded, thus avoiding these invalid audio frames occupying buffer space and therefore not affecting the latency of new voice interactions.
[0294] Secondly, this application provides a method for storing audio data based on a delay protection mechanism. Using this method, the validity of an audio frame is determined based on the audio energy it carries. If the audio frame is invalid, it is discarded. Therefore, audio frames after the end of a voice interaction will not enter the buffer space, thus not affecting the delay of subsequent voice interactions.
[0295] Optionally, in the above Figure 5 Based on one or more corresponding embodiments, in another optional embodiment provided by this application, after playing audio data stored in the second buffer space in response to a continued playback operation for audio, the following may be further included:
[0296] When audio playback ends, release the target buffer space to obtain the first buffer space;
[0297] or,
[0298] During the audio playback process of the audio data stored in the second buffer space, the buffer space corresponding to each cycle is reduced sequentially according to the sampling rate, number of channels and bit depth of the audio until the first buffer space is obtained. The sum of the buffer spaces corresponding to each cycle is equal to the target buffer space.
[0299] In one or more embodiments, two methods for releasing the target buffer space are described. As can be seen from the foregoing embodiments, when the user performs a pause playback operation, based on the effect protection mechanism, the buffer space can be increased starting with the acquired second audio data, and the size of the newly added buffer space (i.e., the target buffer space) is recorded. When the user performs a resume playback operation, the audio continues to play. When the audio playback ends, the size of the buffer space is restored to its initial value (i.e., the size of the first buffer space) to avoid affecting latency.
[0300] The following will introduce two ways to release newly added cache space (i.e., target cache space).
[0301] Method 1: Release all at once;
[0302] Specifically, when audio playback ends, the terminal can release the target cache space all at once. By releasing the target cache space, the cache space can be restored to its initial state, avoiding resource waste and thus improving the overall performance and stability of the system.
[0303] Method two: gradual release;
[0304] Specifically, while the audio continues to play, the audio data stored in the second buffer space is output as a data stream. Therefore, during the output of audio data, the amount of space that needs to be freed can be calculated periodically. That is, the amount of buffer space that needs to be freed in each cycle can be calculated as follows:
[0305] K cache- =S2×C2×D2×T2; Formula (2)
[0306] Among them, K cache- S2 represents the amount of buffer space that needs to be freed. C2 represents the sampling rate. D2 represents the number of channels. T2 represents the bit depth. T2 represents the period duration, for example, 1 second.
[0307] Based on this, taking a period of 1 second, a sampling rate of 48000 Hz, 2 channels, and a bit depth of 2 as an example, the buffer space that needs to be released in each period can be calculated using formula (2) as 192kb. Assuming the target buffer space is 1920kb, then during audio playback, the terminal will reduce the buffer space by 192kb in each period (i.e., 1 second) until the target buffer space is completely released.
[0308] Secondly, this application provides two methods for releasing the target cache space. After triggering the continue playback operation, one method releases the target cache space all at once after the audio playback ends; the other method releases the target cache space gradually during audio playback. Both methods can make reasonable use of resources, optimize resource allocation, and improve system performance.
[0309] Optionally, in the above Figure 5 Based on one or more corresponding embodiments, in another optional embodiment provided by this application, before releasing the target buffer space to obtain the first buffer space after the audio playback ends, the following may be included:
[0310] If the number of silent frames obtained is greater than or equal to the number threshold, then the audio playback is determined to have ended.
[0311] or,
[0312] If the audio energy of the detected audio frame is less than or equal to the audio energy threshold, then the audio playback is determined to have ended.
[0313] In one or more embodiments, two methods for detecting whether audio playback has ended are described. As can be seen from the foregoing embodiments, when audio playback has ended, the target buffer space can be directly released. The following will describe two methods for determining whether audio playback has ended.
[0314] Method 1, based on the number of silent frames;
[0315] Specifically, during audio playback, the terminal can monitor the number of silence frames. If the number of silence frames is greater than or equal to a threshold (e.g., 100 frames), it indicates that audio playback has ended. A silence frame represents an audio frame with zero audio energy; therefore, a large number of consecutive silence frames often indicates that the audio playback has ended. Conversely, a small number of silence frames may indicate a pause in the speech. Therefore, by setting a reasonable threshold, the terminal can determine whether audio playback has ended.
[0316] Method 2, based on audio energy;
[0317] Specifically, during audio playback, the terminal can monitor whether an audio frame with an audio energy less than or equal to an audio energy threshold (e.g., -1) appears. If such a frame exists, it indicates that the audio frame is invalid. Invalid audio frames appear after normal audio playback has finished. Therefore, if an audio frame with an audio energy less than or equal to the audio energy threshold appears, it means that the audio playback has ended.
[0318] Furthermore, this application provides two methods for detecting whether audio playback has ended. One method determines whether audio playback has ended based on the number of silent frames, and the other method determines whether audio playback has ended based on the audio energy of the audio frames. This improves the flexibility and feasibility of the solution.
[0319] Based on the above introduction, the following will combine... Figure 17 This describes the overall flow of the audio playback control method. Please refer to [link / reference]. Figure 17 , Figure 17 This is another flowchart illustrating the audio playback control method in this application, as shown in the figure. Specifically:
[0320] In step S1701, the terminal acquires audio frames.
[0321] In step S1702, the terminal determines whether the audio frame is related to the target application. If yes, step S1703 is executed; otherwise, step S1704 is executed.
[0322] In step S1703, if the audio frame is related to the target application, the terminal determines whether the user has triggered a pause playback operation. If yes, step S1705 is executed; otherwise, step S1706 is executed.
[0323] In step S1704, if the audio frame is unrelated to the target application, the terminal executes a traditional management strategy, that is, there is no need to expand the first buffer space.
[0324] In step S1705, when the user triggers a pause playback operation, the terminal continues to determine whether the audio frame is a valid audio frame. If yes, step S1705 is executed; otherwise, step S1706 is executed.
[0325] In step S1706, if the user triggers an operation that does not pause playback, the terminal executes the traditional management strategy.
[0326] In step S1707, if the audio frame is a valid audio frame, the audio frame should be saved as much as possible.
[0327] In step S1708, if the audio frame is an invalid audio frame, the audio frame is discarded directly.
[0328] The following section will use the "voice assistant" function provided by the CS architecture as an example to introduce the audio playback control system. Please refer to [link / reference]. Figure 18 , Figure 18 This is a schematic diagram of the architecture of the audio playback control system in an embodiment of this application. As shown in the figure, the server and the terminal each include different functional modules, and each functional module is used to perform a corresponding function. Specifically:
[0329] When a user needs to query information through a "voice assistant," the audio input module 1811 in the terminal is used to input audio and send the input audio information to the server. The speech-to-text module 1801 in the server is used to convert the received audio information into text. Then, the information retrieval module 1802 uses the large model capability to perform information retrieval to obtain the search results. The text-to-speech module 1803 then converts the search results into corresponding audio data. Based on this, the audio frame differentiation module 1804 adds a source field indicating the "voice assistant" to the packet header information of each audio frame included in the audio data. The audio streaming module 1805 pushes the audio data to the terminal. After receiving the audio data, the terminal caches the audio data using the audio caching module 1812.
[0330] Based on this, the audio playback module 1814 starts audio playback, at which point the user can hear the search results provided by the "voice assistant". The pause and resume playback module 1815 is used to detect the pause and resume playback operations triggered by the user and notify the audio frame protection module 1813 of the "voice assistant" of the operation instructions. After recognizing the pause playback operation, the audio frame protection module 1813 of the "voice assistant" performs relevant processing on the cached audio data. The audio frame protection module 1813 includes an audio frame recognition module 18131, a cache protection module 18132, a delay protection module 18133, and an effect protection module 18134.
[0331] The audio frame recognition module 18131 is used to identify whether the audio frame comes from the "voice assistant". If so, the cache protection strategy, delay protection strategy and effect protection strategy are continued to be executed.
[0332] The cache protection module 18132 is used to execute a cache protection strategy, that is, in response to the pause playback operation, the original cache space is increased to cache as many audio frames related to the "voice assistant" as possible until all audio frames related to the current voice conversation can be saved.
[0333] The delay protection module 18133 is used to implement a delay protection strategy, which means that after the server pushes the audio frames related to the "voice assistant," it will continue to push silence frames. A large number of silence frames will occupy buffer space, causing subsequent voice interaction responses to be slow. Therefore, the system determines whether the voice interaction has ended by adding audio energy to the audio frames. If the audio energy of the audio frame is less than a threshold, the audio frame is discarded directly.
[0334] The effect protection module 18134 is used to implement an effect protection strategy, namely, in response to a pause playback operation, it increases the buffer space with each received audio frame and records the size of the increase. In response to a resume playback operation, after the audio data playback ends, it restores the buffer space to its initial value to avoid affecting the overall latency.
[0335] The audio playback control device in this application is described in detail below. Please refer to [link / reference]. Figure 19 , Figure 19 This is a schematic diagram of one embodiment of the audio playback control device in this application. The audio playback control device 1900 includes:
[0336] The playback module 1901 is used to play audio data of the first audio data stored in the first buffer space, wherein the first audio data comes from the target audio data and the first audio data is the audio data that has been obtained.
[0337] The acquisition module 1902 is used to acquire second audio data in response to a pause playback operation for audio, wherein the second audio data is derived from the target audio data and is the audio data acquired after the audio playback is paused;
[0338] The processing module 1903 is used to expand the first cache space according to the second audio data to obtain a second cache space, wherein the second cache space includes the first cache space and the target cache space, and the second cache space is used to store the first audio data and the second audio data;
[0339] The playback module 1901 is also configured to play audio data stored in the second buffer space in response to a continued playback operation for audio, wherein the audio data includes at least the second audio data.
[0340] Optionally, in the above Figure 19 Based on the corresponding embodiments, in another embodiment of the audio playback control device 1900 provided in this application, the audio playback control device 1900 further includes a sending module 1904 and a storage module 1905;
[0341] The acquisition module 1902 is also used to acquire target voice data through a voice acquisition device before playing the first audio data stored in the first buffer space;
[0342] The sending module 1904 is used to send target speech data to the server so that the server converts the target speech data into raw text data and generates target text data based on the raw text data. The server is used to convert the target text data into target audio data.
[0343] Storage module 1905 is used to store the received audio data in the first cache space during the process of receiving target audio data sent by the receiving server.
[0344] Optionally, in the above Figure 19 Based on the corresponding embodiments, in another embodiment of the audio playback control device 1900 provided in this application, the audio playback control device 1900 is applied to a first terminal;
[0345] The storage module 1905 is also used to store the received audio data in the first cache space before playing the first audio data stored in the first cache space, during the process of receiving the target audio data sent by the receiving server.
[0346] Among them, the target audio data is obtained by converting the target text data, the target text data is generated based on the original text data, the original text data is obtained by converting the target speech data, and the target speech data is obtained by the second terminal through the speech acquisition device.
[0347] Optionally, in the above Figure 19 Based on the corresponding embodiments, in another embodiment of the audio playback control device 1900 provided in this application, the audio playback control device 1900 further includes a determination module 1906;
[0348] The acquisition module 1902 is further configured to acquire a first audio frame sequence, wherein the first audio frame sequence includes at least one audio frame;
[0349] The determining module 1906 is configured to determine that the first audio frame sequence originates from target audio data if each audio frame in the first audio frame sequence carries a source field indicating the target application; or,
[0350] If each audio frame in the first audio frame sequence carries a source field for indicating the target application, and the cumulative duration of the audio frames in the first audio frame sequence is greater than or equal to a duration threshold, then the first audio frame sequence is determined to originate from the target audio data.
[0351] Optionally, in the above Figure 19 Based on the corresponding embodiments, in another embodiment of the audio playback control device 1900 provided in this application,
[0352] The acquisition module 1902 is also used to acquire target voice data through a voice acquisition device before playing the first audio data stored in the first buffer space;
[0353] Processing module 1903 is also used to convert target speech data into raw text data;
[0354] Processing module 1903 is also used to generate target text data based on the original text data;
[0355] Storage module 1905 is also used to store the converted audio data in the first cache space during the process of converting target text data into target audio data.
[0356] Optionally, in the above Figure 19 Based on the corresponding embodiments, in another embodiment of the audio playback control device 1900 provided in this application,
[0357] The acquisition module 1902 is also used to acquire target voice data through a voice acquisition device before playing the first audio data stored in the first buffer space;
[0358] The sending module 1904 is also used to send target voice data to the server so that the server converts the target voice data into raw text data and generates target text data based on the raw text data;
[0359] The acquisition module 1902 is also used to receive target text data sent by the server;
[0360] Storage module 1905 is also used to store the converted audio data in the first cache space during the process of converting target text data into target audio data.
[0361] Optionally, in the above Figure 19 Based on the corresponding embodiments, in another embodiment of the audio playback control device 1900 provided in this application, the audio playback control device 1900 is applied to a first terminal;
[0362] The acquisition module 1902 is also used to receive target text data sent by the server before playing the first audio data stored in the first cache space. The target text data is generated based on the original text data, which is obtained after converting the target speech data. The target speech data is obtained by the second terminal through the speech acquisition device.
[0363] Storage module 1905 is also used to store the converted audio data in the first cache space during the process of converting target text data into target audio data.
[0364] Optionally, in the above Figure 19 Based on the corresponding embodiments, in another embodiment of the audio playback control device 1900 provided in this application,
[0365] The processing module 1903 is specifically used to, during the process of acquiring the second audio data, if the remaining capacity of the first buffer space is less than or equal to the remaining capacity threshold, add a target buffer space on the basis of the first buffer space to obtain the second buffer space, wherein the size of the target buffer space is preset.
[0366] Optionally, in the above Figure 19 Based on the corresponding embodiments, in another embodiment of the audio playback control device 1900 provided in this application,
[0367] The processing module 1903 is specifically used to determine the buffer space corresponding to each cycle based on the sampling rate, number of channels and bit depth of the audio when the remaining capacity of the first buffer space is less than or equal to the remaining capacity threshold during the process of acquiring the second audio data.
[0368] Based on the first cache space, the cache space corresponding to each cycle is added sequentially to obtain the second cache space, wherein the sum of the cache spaces corresponding to each cycle is equal to the target cache space.
[0369] Optionally, in the above Figure 19 Based on the corresponding embodiments, in another embodiment of the audio playback control device 1900 provided in this application,
[0370] The acquisition module 1902 is specifically used to acquire a second audio frame sequence, wherein the second audio frame sequence includes at least one audio frame, and each audio frame corresponds to an audio energy;
[0371] All audio frames in the second audio frame sequence whose audio energy is greater than or equal to the audio energy threshold are used as the second audio data.
[0372] Optionally, in the above Figure 19 Based on the corresponding embodiments, in another embodiment of the audio playback control device 1900 provided in this application,
[0373] Processing module 1903 is further configured to, in response to a continued audio playback operation, after playing the audio data stored in the second buffer space, release the target buffer space to obtain the first buffer space upon completion of audio playback; or,
[0374] During the audio playback process of the audio data stored in the second buffer space, the buffer space corresponding to each cycle is reduced sequentially according to the sampling rate, number of channels and bit depth of the audio until the first buffer space is obtained. The sum of the buffer spaces corresponding to each cycle is equal to the target buffer space.
[0375] Optionally, in the above Figure 19Based on the corresponding embodiments, in another embodiment of the audio playback control device 1900 provided in this application,
[0376] The determination module 1906 is further configured to, upon the end of audio playback, release the target buffer space before obtaining the first buffer space, and if the number of silent frames obtained is greater than or equal to a threshold, determine that audio playback has ended; or,
[0377] If the audio energy of the detected audio frame is less than or equal to the audio energy threshold, then the audio playback is determined to have ended.
[0378] This application also provides a terminal, such as... Figure 20 As shown, for ease of explanation, only the parts related to the embodiments of this application are shown. For specific technical details not disclosed, please refer to the method section of the embodiments of this application. In the embodiments of this application, a mobile phone is used as an example for illustration:
[0379] Figure 20 This is a block diagram illustrating a portion of the structure of a mobile phone related to the terminal provided in the embodiments of this application. (Reference) Figure 20 The mobile phone includes components such as a radio frequency (RF) circuit 2010, a memory 2020, an input unit 2030, a display unit 2040, a sensor 2050, an audio circuit 2060, a wireless fidelity (WiFi) module 2070, a processor 2080, and a power supply 2090. Those skilled in the art will understand that... Figure 20 The mobile phone structure shown does not constitute a limitation on the mobile phone and may include more or fewer components than shown, or combine certain components, or have different component arrangements. The following section will discuss this further. Figure 20 A detailed introduction to each component of a mobile phone:
[0380] The RF circuit 2010 can be used for receiving and transmitting signals during information transmission or calls. Specifically, it receives downlink information from the base station and processes it with the processor 2080; additionally, it transmits uplink data to the base station. Typically, the RF circuit 2010 includes, but is not limited to, an antenna, at least one amplifier, a transceiver, a coupler, a low-noise amplifier (LNA), and a duplexer. Furthermore, the RF circuit 2010 can also communicate wirelessly with networks and other devices. The aforementioned wireless communication can use any communication standard or protocol, including but not limited to Global System for Mobile Communication (GSM), General Packet Radio Service (GPRS), Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Long Term Evolution (LTE), email, and Short Message Service (SMS).
[0381] The memory 2020 can be used to store software programs and modules. The processor 2080 executes various mobile phone functions and data processing by running the software programs and modules stored in the memory 2020. The memory 2020 may mainly include a program storage area and a data storage area. The program storage area may store the operating system, at least one application program required for a function (such as sound playback function, image playback function, etc.), etc.; the data storage area may store data created according to the use of the mobile phone (such as audio data, phonebook, etc.). In addition, the memory 2020 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device.
[0382] The input unit 2030 can be used to receive input numerical or character information, and to generate key signal inputs related to user settings and function control of the mobile phone. Specifically, the input unit 2030 may include a touch panel 2031 and other input devices 2032. The touch panel 2031, also known as a touch screen, can collect touch operations performed by the user on or near it (such as operations performed by the user using a finger, stylus, or any suitable object or accessory on or near the touch panel 2031), and drive the corresponding connection devices according to a pre-set program. Optionally, the touch panel 2031 may include two parts: a touch detection device and a touch controller. The touch detection device detects the user's touch position and the signal generated by the touch operation, and transmits the signal to the touch controller; the touch controller receives touch information from the touch detection device, converts it into touch point coordinates, and sends it to the processor 2080, and can also receive and execute commands sent by the processor 2080. In addition, the touch panel 2031 can be implemented using various types such as resistive, capacitive, infrared, and surface acoustic wave. In addition to the touch panel 2031, the input unit 2030 may also include other input devices 2032. Specifically, other input devices 2032 may include, but are not limited to, one or more of the following: physical keyboard, function keys (such as volume control buttons, power buttons, etc.), mouse, joystick, etc.
[0383] The display unit 2040 can be used to display information input by the user or information provided to the user, as well as various menus of the mobile phone. The display unit 2040 may include a display panel 2041, which may optionally be configured as a liquid crystal display (LCD), organic light-emitting diode (OLED), or similar device. Further, a touch panel 2031 may cover the display panel 2041. When the touch panel 2031 detects a touch operation on or near it, it transmits the information to the processor 2080 to determine the type of touch event. Subsequently, the processor 2080 provides corresponding visual output on the display panel 2041 based on the type of touch event. Although in Figure 20 In this context, the touch panel 2031 and the display panel 2041 are two separate components for implementing the input and output functions of the mobile phone. However, in some embodiments, the touch panel 2031 and the display panel 2041 can be integrated to achieve the input and output functions of the mobile phone.
[0384] The mobile phone may also include at least one sensor 2050, such as a light sensor, a motion sensor, and other sensors. Specifically, the light sensor may include an ambient light sensor and a proximity sensor. The ambient light sensor can adjust the brightness of the display panel 2041 according to the ambient light level, and the proximity sensor can turn off the display panel 2041 and / or backlight when the phone is moved to the ear. As a type of motion sensor, an accelerometer sensor can detect the magnitude of acceleration in various directions (generally three axes). When stationary, it can detect the magnitude and direction of gravity, and can be used for applications that recognize the phone's posture (such as landscape / portrait switching, related games, magnetometer posture calibration), vibration recognition-related functions (such as pedometer, taps), etc. Other sensors that may be configured in the mobile phone, such as gyroscopes, barometers, hygrometers, thermometers, and infrared sensors, will not be described in detail here.
[0385] Audio circuit 2060, speaker 2061, and microphone 2062 provide an audio interface between the user and the mobile phone. Audio circuit 2060 converts received audio data into electrical signals and transmits them to speaker 2061, where speaker 2061 converts them into sound signals for output. On the other hand, microphone 2062 converts collected sound signals into electrical signals, which are received by audio circuit 2060, converted into audio data, and then processed by processor 2080 before being transmitted via RF circuit 2010 to, for example, another mobile phone, or the audio data can be output to memory 2020 for further processing.
[0386] WiFi is a short-range wireless transmission technology. Through the WiFi module 2070, mobile phones can help users send and receive emails, browse web pages, and access streaming media, providing users with wireless broadband internet access. Although Figure 20 WiFi module 2070 is shown, but it is understood that it is not an essential component of a mobile phone and can be omitted as needed without changing the essence of the invention.
[0387] The processor 2080 is the control center of the mobile phone, connecting various parts of the phone through various interfaces and lines. It performs various functions and processes data by running or executing software programs and / or modules stored in the memory 2020 and calling data stored in the memory 2020. Optionally, the processor 2080 may include one or more processing units; optionally, the processor 2080 may integrate an application processor and a modem processor, wherein the application processor mainly handles the operating system, user interface, and applications, and the modem processor mainly handles wireless communication. It is understood that the aforementioned modem processor may also not be integrated into the processor 2080.
[0388] The mobile phone also includes a power supply 2090 (such as a battery) to power various components. Optionally, the power supply can be logically connected to the processor 2080 through a power management system, thereby enabling functions such as charging, discharging, and power consumption management through the power management system. Although not shown, the mobile phone may also include a camera, Bluetooth module, etc., which will not be described in detail here.
[0389] The steps performed by the terminal in the above embodiments can be based on this Figure 20 The terminal structure shown.
[0390] This application also provides a computer device, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the steps of the methods described in the foregoing embodiments.
[0391] This application also provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the steps of the methods described in the foregoing embodiments.
[0392] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the methods described in the foregoing embodiments.
[0393] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0394] In this application embodiment, the terms "module" or "unit" refer to a computer program or part of a computer program that has a predetermined function and works with other related parts to achieve a predetermined goal, and can be implemented wholly or partially using software, hardware (such as processing circuitry or memory), or a combination thereof. Similarly, a processor (or multiple processors or memory) can be used to implement one or more modules or units. Furthermore, each module or unit can be part of an overall module or unit that includes the functionality of that module or unit.
[0395] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be an indirect coupling or communication connection between apparatuses or units through some interfaces, and may be electrical, mechanical, or other forms.
[0396] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0397] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0398] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a server or terminal device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing computer programs, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0399] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.
Claims
1. A method for controlling audio playback, characterized in that, include: Play audio data from the first audio data stored in the first cache space, wherein the first audio data is derived from the target audio data and is already acquired audio data. In response to a pause playback operation on the audio, second audio data is acquired, wherein the second audio data originates from the target audio data and is audio data acquired after the audio playback is paused; Based on the second audio data, the first cache space is expanded to obtain a second cache space, wherein the second cache space includes the first cache space and the target cache space, and the second cache space is used to store the first audio data and the second audio data; In response to a continue playback operation for the audio, audio data stored in the second buffer space is played, wherein the audio data includes at least the second audio data.
2. The control method according to claim 1, characterized in that, Before playing the first audio data stored in the first buffer space, the method further includes: Acquire target voice data using a voice acquisition device; The server sends the target speech data to the server so that the server converts the target speech data into raw text data and generates target text data based on the raw text data, wherein the server is used to convert the target text data into the target audio data; During the process of receiving the target audio data sent by the server, the received audio data is stored in the first cache space.
3. The control method according to claim 1, characterized in that, The control method is applied to the first terminal; Before playing the first audio data stored in the first buffer space, the method further includes: During the process of receiving the target audio data sent by the receiving server, the received audio data is stored in the first cache space; The target audio data is obtained by converting the target text data, the target text data is generated based on the original text data, the original text data is obtained by converting the target speech data, and the target speech data is obtained by the second terminal through the speech acquisition device.
4. The control method according to any one of claims 1 to 3, characterized in that, The method further includes: Obtain a first audio frame sequence, wherein the first audio frame sequence includes at least one audio frame; If each audio frame in the first audio frame sequence carries a source field indicating the target application, then the first audio frame sequence is determined to originate from the target audio data; or, If each audio frame in the first audio frame sequence carries a source field for indicating the target application, and the cumulative duration of the audio frames in the first audio frame sequence is greater than or equal to a duration threshold, then the first audio frame sequence is determined to originate from the target audio data.
5. The control method according to claim 1, characterized in that, Before playing the first audio data stored in the first buffer space, the method further includes: Acquire target voice data using a voice acquisition device; Convert the target speech data into raw text data; Generate target text data based on the original text data; During the process of converting the target text data into the target audio data, the converted audio data is stored in the first cache space.
6. The control method according to claim 1, characterized in that, Before playing the first audio data stored in the first buffer space, the method further includes: Acquire target voice data using a voice acquisition device; The target speech data is sent to the server so that the server converts the target speech data into raw text data and generates target text data based on the raw text data. Receive the target text data sent by the server; During the process of converting the target text data into the target audio data, the converted audio data is stored in the first cache space.
7. The control method according to claim 1, characterized in that, The control method is applied to the first terminal; Before playing the first audio data stored in the first buffer space, the method further includes: The server receives target text data, wherein the target text data is generated based on the original text data, the original text data is obtained after converting the target speech data, and the target speech data is obtained by the second terminal through a speech acquisition device. During the process of converting the target text data into the target audio data, the converted audio data is stored in the first cache space.
8. The control method according to any one of claims 1 to 7, characterized in that, The step of expanding the first buffer space according to the second audio data to obtain a second buffer space includes: During the acquisition of the second audio data, if the remaining capacity of the first cache space is less than or equal to the remaining capacity threshold, the target cache space is added to the first cache space to obtain the second cache space, wherein the size of the target cache space is preset.
9. The control method according to any one of claims 1 to 7, characterized in that, The step of expanding the first buffer space according to the second audio data to obtain a second buffer space includes: During the acquisition of the second audio data, if the remaining capacity of the first buffer space is less than or equal to the remaining capacity threshold, the buffer space corresponding to each cycle is determined according to the sampling rate, number of channels and bit depth of the audio. Based on the first cache space, the cache space corresponding to each cycle is added sequentially to obtain the second cache space, wherein the sum of the cache spaces corresponding to each cycle is equal to the target cache space.
10. The control method according to any one of claims 1 to 9, characterized in that, The acquisition of the second audio data includes: Obtain a second audio frame sequence, wherein the second audio frame sequence includes at least one audio frame, and each audio frame corresponds to an audio energy; All audio frames in the second audio frame sequence whose audio energy is greater than or equal to the audio energy threshold are used as the second audio data.
11. The control method according to any one of claims 1 to 10, characterized in that, After playing the audio data stored in the second buffer space in response to the continued playback operation for the audio, the method further includes: When the audio playback ends, the target buffer space is released to obtain the first buffer space; or, During the audio playback process of the audio data stored in the second cache space, the cache space corresponding to each cycle is reduced sequentially according to the sampling rate, number of channels and bit depth of the audio until the first cache space is obtained, wherein the sum of the cache spaces corresponding to each cycle is equal to the target cache space.
12. The control method according to claim 11, characterized in that, Before releasing the target buffer space to obtain the first buffer space after the audio playback ends, the method further includes: If the number of silent frames obtained is greater than or equal to the number threshold, then the audio playback is determined to have ended; or, If the audio energy of the detected audio frame is less than or equal to the audio energy threshold, then the audio playback is determined to have ended.
13. An audio playback control device, characterized in that, include: The playback module is used to play audio data of the first audio data stored in the first cache space, wherein the first audio data comes from the target audio data and the first audio data is already obtained audio data. The acquisition module is used to acquire second audio data in response to a pause playback operation for audio, wherein the second audio data originates from the target audio data and is audio data acquired after the audio playback is paused; The processing module is used to expand the first cache space according to the second audio data to obtain a second cache space, wherein the second cache space includes the first cache space and the target cache space, and the second cache space is used to store the first audio data and the second audio data; The playback module is further configured to play audio data stored in the second buffer space in response to a continue playback operation for the audio, wherein the audio data includes at least the second audio data.
14. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the control method according to any one of claims 1 to 12.
15. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the control method according to any one of claims 1 to 12.
16. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the control method according to any one of claims 1 to 12.