Audio playing method and electronic equipment
By adopting streaming reception and splitting decoding processing of audio playback client in the browser, the dual storage strategy of audio data is realized, which solves the compatibility and real-time problems of audio playback in the browser, and improves user experience and system efficiency.
Patent Information
- Application Number
- CN202510752154.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-05
- Publication Date
- 2025-09-05
AI Technical Summary
When the prior art realizes audio playback in web applications such as browsers, there are problems such as poor compatibility, high development complexity, and insufficient real-time performance, resulting in poor user experience.
The audio playback client uses the audio data to receive audio data through streaming, perform split processing and pre-decoding, realizes a dual storage strategy, generates audio sources while playing, reduces interactive steps with the server, and improves playback real-time.
Significantly reduce user waiting time, lower system power consumption, reduce development and maintenance costs, and improve user experience and playback effects.
Smart Images

Figure CN120596054A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of audio playback technology, and in particular to an audio playback method and electronic equipment. Background Art
[0002] For example, when implementing audio playback in a web application like a browser, the audio playback system typically consists of an audio playback client and an audio playback server. The audio playback client corresponds to the web application, acting as the front-end for audio playback to the user, and is used to play audio. The audio playback server is the server corresponding to the web application, acting as the back-end for audio playback, and providing audio to the browser.
[0003] With the continuous development of computer technology, the types of web applications such as browsers are increasing, and the types of audio playback scenarios supported by browsers and other web applications are also increasing. How to better implement audio playback in web applications such as browsers is a problem that is currently being explored in the field. Summary of the Invention
[0004] The embodiments of the present application provide an audio playback method and an electronic device, which can better implement audio playback in web applications such as browsers.
[0005] To solve the above technical problems, in the first aspect, an embodiment of the present application provides an audio playback method, which can be applied to an audio playback client, and the method includes: receiving first audio data, segmenting the first audio data to obtain multiple second audio data, and storing the second audio data in a target data storage array respectively, where the first audio data is any one of the multiple streaming audio data corresponding to the target audio sent by the audio playback server; pre-decoding the unplayed second audio data stored in the target data storage array to obtain third audio data, and creating an audio source, setting the third audio data as the audio data of the audio source for storage; determining the target audio source to be played, and playing the target audio source according to the audio data of the target audio source.
[0006] By adopting the audio playback method provided in this embodiment, the audio playback client receives the streaming audio data corresponding to the target audio in a streaming manner, slices the streaming audio data, realizes slice caching, and pre-decodes the stored audio data to obtain the audio source to be played for playback. This can play the audio more timely, improve the real-time performance of the audio playback, thereby improving the audio playback effect and thus improving the user experience.
[0007] Furthermore, based on the audio playback method, the second audio data is stored in the target data storage array to store the audio data for the first time. The third audio data is set to be stored as the audio data of the audio source to store the audio data for the second time. Thus, a dual storage strategy or double buffering strategy for audio data is implemented in the audio playback method. Therefore, in the process of playing audio based on the audio playback method, while playing the audio source of the current segment, the third audio data of the audio source of the next segment will be pre-decoded to obtain the next audio source, thereby better achieving seamless playback of the audio source. This achieves the generation of the audio source while playing, which effectively solves the problem of poor user experience caused by long user waiting time compared to the method of waiting for all streaming audio data to be received and decoded before playing. In addition, the double buffering strategy adopted in the audio playback method can reduce the delay of the first frame audio source playback time to less than 300ms in some scenarios, compared to the existing method of waiting for the complete audio source to be generated before playing, significantly reducing user waiting time.
[0008] Furthermore, the audio playback method is applied to the audio playback client, so that the audio playback client performs all the processing from receiving the streaming audio data corresponding to the audio to implementing the audio playback process, without the participation of the audio playback server such as the server, and the implementation logic is simpler. In this way, on the one hand, the interaction steps between the audio playback client and the audio playback server during the audio playback process are effectively reduced, the real-time performance of the playback is improved, and the system power consumption is reduced. On the other hand, it can also reduce the complexity of system development, reduce the cost of system development, deployment and maintenance, and improve the efficiency of system development and reduce the implementation cost. On the other hand, it can also better adapt to different systems and applications.
[0009] In summary, the audio playback method provided in the embodiments of the present application can better implement audio playback in web applications such as browsers. In addition, the audio playback method provided in the present application implements streaming transmission of audio data. Therefore, the method can also be called an audio streaming media playback method, and the target audio can be called target audio streaming media.
[0010] In a possible implementation of the first aspect above, the streaming audio data corresponding to the target audio can be generated by a multimodal large language model, and the multimodal large language model can be deployed on an audio playback service end such as a server.
[0011] In one possible implementation of the first aspect, the audio playback client may receive multiple streamed audio data generated by the multimodal large language model, namely, the first audio data, in a streaming manner. The audio playback client receives the multiple first audio data gradually in a streaming manner, rather than receiving all of the first audio data at once. This improves the efficiency and performance of network requests, particularly when processing large files or on slow networks.
[0012] In a possible implementation of the first aspect, the first audio data and the second audio data may be binary data.
[0013] In a possible implementation of the first aspect, pre-decoding the unplayed second audio data stored in the target data storage array to obtain third audio data includes: sequentially acquiring the unplayed second audio data from the target data storage array according to the value of a current audio playback position identifier variable, and accumulating data amounts of the acquired second audio data; if the sum of the data amounts of the acquired second audio data is greater than or equal to a first data amount threshold, determining whether the acquired second audio data includes file header information of the target audio; if the acquired second audio data includes the file header information of the target audio, synthesizing the acquired second audio data to obtain fourth audio data; if the acquired second audio data does not include the file header information of the target audio, synthesizing the acquired second audio data with first second audio data corresponding to the target audio stored in the target data storage array to obtain fourth audio data; and decoding the fourth audio data to obtain third audio data.
[0014] With the audio playback method of this embodiment, when the sum of the acquired second audio data amounts is greater than or equal to the first data amount threshold, data synthesis processing is performed to obtain fourth audio data. This avoids the problem of the acquired fourth audio data being too small, resulting in a smaller generated audio source size, a shorter audio source playback time, interrupted audio playback, and tearing sounds, which affects the user experience.
[0015] Furthermore, in the audio playback method of this embodiment, it is determined whether the acquired second audio data includes the file header information of the target audio. If the acquired second audio data does not include the file header information of the target audio, the acquired second audio data and the second audio data including the header information corresponding to the target audio are synthesized. Since the header information of the target audio is usually in the header of the first streaming audio data, the acquired second audio data and the first second audio data are synthesized to obtain fourth audio data, so as to ensure that the fourth audio data contains the file header information of the target audio, thereby avoiding errors when decoding the fourth audio data and affecting the decoding of the fourth audio data.
[0016] In a possible implementation of the first aspect described above, the identifier of the second audio data corresponding to the currently playing audio can be determined based on the value of the current audio playback position identifier variable. Therefore, the specific steps of obtaining the unplayed second audio data may include: sequentially intercepting the unplayed second audio data from the plurality of second audio data following the identifier corresponding to the second audio data corresponding to the currently playing audio in the target data storage array, in the arrangement direction of the plurality of second audio data stored in the target data storage array, to obtain the unplayed second audio data in the target data storage array, and accumulating the amount of the obtained second audio data until the sum of the amount of the obtained second audio data is greater than or equal to the first data amount threshold.
[0017] In one possible implementation of the first aspect, after obtaining the sum of the data volumes of the acquired second audio data, it is determined whether the sum of the data volumes of the acquired second audio data is greater than or equal to a first data volume threshold. If it is determined that the sum of the data volumes of the acquired second audio data is less than the first data volume threshold, the generation of new second audio data is waited for until the sum of the data volumes of the acquired second audio data is greater than or equal to the first data volume threshold, and then subsequent steps are performed to ensure that the data volume of the acquired fourth audio data is sufficient, thereby ensuring audio playback quality.
[0018] In a possible implementation of the first aspect described above, pre-decoding the unplayed second audio data stored in the target data storage array to obtain third audio data further includes: if a target pre-decoding condition is met, pre-decoding the unplayed second audio data stored in the target data storage array to obtain third audio data.
[0019] In a possible implementation of the first aspect above, the audio playback method also includes determining that the target pre-decoding condition is met when it is determined that one or more of the following conditions are met: the amount of unplayed second audio data in the target data storage array is greater than or equal to the second data amount threshold; audio playback processing is currently being performed; decoding processing is not currently being performed; reception of multiple streaming audio data corresponding to the target audio has not been completed; the complete audio source variable has no value; and there is unplayed second audio data in the target data storage array.
[0020] The audio playback method of this embodiment determines whether the target pre-decoding condition is met. If the target pre-decoding condition is met, pre-decoding is performed on the unplayed second audio data in the target data storage array, thereby avoiding decoding errors and better ensuring the execution of pre-decoding.
[0021] In a possible implementation of the first aspect above, before performing pre-decoding processing on the unplayed second audio data stored in the target data storage array, the method further includes: setting the audio decoding state variable to decoding and clearing the value of the next audio source variable.
[0022] In the audio playback method of this embodiment, the audio playback client sets the audio decoding state variable of the audio decoder in the audio playback client to "decoding", which can prevent other audio data to be decoded from occupying the audio decoder and ensure smooth decoding. The audio playback client clears the value of the next audio source variable, which can facilitate the use of the third audio data decoded by the audio decoder as the next audio source variable, that is, as the next audio source to be played.
[0023] In a possible implementation of the first aspect above, after creating the audio source, the method further includes: setting the created audio source to a next audio source variable, and updating the value of the current audio playback position identification variable to the value of the next audio playback position identification variable.
[0024] In the audio playback method of this embodiment, the created audio source is set as the next audio source variable, that is, the audio source is used as the next audio source to be played, so that the next audio source to be played can be played later. The value of the current audio playback position identifier variable is updated to the value of the next audio playback position identifier variable, so that the audio playback client can obtain the unplayed second audio data from the target data storage array based on the value of the next audio playback position identifier variable, and execute subsequent steps to complete the playback of the target audio.
[0025] In a possible implementation of the first aspect above, determining whether the acquired second audio data includes file header information in the target audio includes: determining whether the acquired second audio data includes file header information in the target audio based on the value of a current audio playback position identification variable.
[0026] Using the audio playback method in this embodiment, according to the value of the current audio playback position identification variable, it is determined whether the acquired second audio data includes the file header information of the target audio, so as to ensure that the obtained fourth audio data contains the file header information of the target audio, so as to avoid errors when decoding the fourth audio data, affecting the decoding of the fourth audio data.
[0027] In one possible implementation of the first aspect described above, whether the acquired second audio data includes the file header information of the target audio can be determined based on the value of the current audio playback position identification variable. If the value of the current audio playback position identification variable is the value of the audio playback position identification variable corresponding to the first and second audio data of the target audio stored in the target data storage array, the acquired second audio data includes the file header information of the target audio. If the value of the current audio playback position identification variable is greater than the value of the audio playback position identification variable corresponding to the first and second audio data of the target audio stored in the target data storage array, the acquired second audio data does not include the file header information of the target audio.
[0028] In a possible implementation of the first aspect, decoding the fourth audio data to obtain the third audio data includes decoding the fourth audio data using a target decoding method corresponding to the audio context to obtain the third audio data.
[0029] By adopting the audio playback method in this embodiment, the fourth audio data is decoded and processed through the target decoding method corresponding to the audio context, thereby improving the efficiency of audio data decoding processing and enhancing the operational flexibility of audio data decoding processing.
[0030] In a possible implementation of the first aspect above, when the reception of multiple streaming audio data corresponding to the target audio is completed, the unplayed second audio data stored in the target data storage array is pre-decoded to obtain third audio data, including: pre-decoding all unplayed second audio data stored in the target data storage array to obtain third audio data; and the method also includes: setting the corresponding audio source created based on the third audio data to the complete audio source variable, and setting the value of the data loading completion mark variable corresponding to the target audio to loading completion.
[0031] In the audio playback method of this embodiment, when the reception of the multiple streaming audio data corresponding to the target audio is completed, the value of the data loading completion flag variable corresponding to the target audio is set to loading completed. In this way, it can be subsequently determined whether the multiple streaming audio data corresponding to the current target audio have all been received based on the value of the data loading completion flag variable.
[0032] In a possible implementation of the first aspect, determining the target audio source to be played includes: initializing an audio source variable to be played corresponding to the audio source to be played; determining whether a complete audio source variable has a value; if so, setting the value of the audio source variable to be played to the value of the complete audio source variable to determine the target audio source, clearing the value of the next audio source variable, and updating the value of the current audio playback position identifier variable to the length of the target data storage array; if there is no value, determining whether the next audio source variable has a value; if so, setting the value of the audio source variable to be played to the value of the next audio source variable to determine the target audio source; if there is no value, determining Determine whether decoding processing is currently in progress; if so, wait for the decoding processing to be completed, set the value of the audio source variable to be played to the value of the next audio source variable to determine the target audio source; if not, after waiting for the sum of the data amount of the unplayed second audio data stored in the target data storage array to be greater than or equal to the second data amount threshold, take out all the currently stored unplayed second audio data from the target data storage array, perform decoding processing to obtain third audio data, create an audio source, set the third audio data as the audio data of the audio source for storage, and set the audio source to the audio source variable to be played to determine the target audio source.
[0033] The audio playback method of this embodiment uses multiple determinations to determine the target audio source, ensuring that the target audio source is more accurate and that the playback order of the audio sources matches the order in which the first audio data corresponding to the audio sources are received. Furthermore, multiple determinations to determine the target audio source ensure uninterrupted playback of the audio corresponding to the audio source.
[0034] In a possible implementation of the first aspect above, an initialization operation is performed on the audio source variable to be played corresponding to the audio source to be played to clear the value of the audio source variable to be played corresponding to the audio source to be played, so that in the process of playing the target audio based on the audio playback method, the audio source variable to be played is assigned a value to ensure accurate playback of the audio.
[0035] In a possible implementation of the first aspect, playing the target audio source includes setting a global current audio source variable to the audio source and binding a play end monitor to play the target audio source.
[0036] By adopting the audio playback method in this embodiment, the audio source is bound to the playback monitoring to monitor the playback of the audio source in real time and ensure the stability of the audio source playback.
[0037] In a possible implementation of the first aspect above, before determining the target audio source to be played, the audio playback method further includes: if it is determined that audio source playback processing is not currently being performed, executing the audio source playback method to determine whether the audio context is currently in a working state; if the audio context is not currently in a working state, setting the audio context to a working state; if the audio context is currently in a working state, entering the audio source playback process to determine the target audio source to be played.
[0038] The audio playback method of this embodiment determines whether an audio source is currently being played, thereby determining whether the created audio source to be played is the target audio source to be played. This ensures that the audio sources are played in the order they were generated, preventing errors in the order in which the audio sources are played. Furthermore, by setting the audio context to a working state, the target audio source to be played can enter the audio source playback process.
[0039] In a possible implementation of the first aspect above, after the target audio source is played, the audio playback method further includes: determining whether the reception of the multiple streaming audio data corresponding to the target audio has been completed, and determining whether there is any unplayed second audio data in the target data storage array through the value of the data loading completion mark variable; if the reception of the multiple streaming audio data corresponding to the target audio has been completed, and there is no unplayed second audio data in the target data storage array, then ending the playback processing for the target audio; if the reception of the multiple streaming audio data corresponding to the target audio has not been completed, or there is any unplayed second audio data in the target data storage array, then continuing the playback processing for the target audio to play the next target audio source to ensure that the target audio to be played is completed.
[0040] Using the audio playback method in this embodiment, the data loading completion mark variable is used to determine whether the streaming audio data corresponding to the target audio has been received, so as to determine whether the playback processing of the target audio has been ended to ensure the complete playback of the target audio.
[0041] In a possible implementation of the first aspect above, the audio playback method further includes: when determining to play the target audio, or before playing the target audio, or before receiving the first audio data, performing audio playback initialization processing, creating variable information required in the audio playback process, that is, creating variable information for audio playback processing, the variable information at least includes a target data storage array, and a first data volume threshold, a second data volume threshold, an audio context, a to-be-played audio source variable, a current audio playback position identification variable, a next audio playback position identification variable, an audio decoding state variable, a current audio source variable, a next audio source variable, a complete audio source variable, and a data loading completion mark variable. Of course, other variable information can also be created.
[0042] The audio playback method in this embodiment is used to establish the variable information required during the audio playback process, so as to facilitate the updating of the required variable information during the audio playback process, and to make it easier for the audio playback client to play different audio sources when the target audio to be played is in different playback stages, so as to facilitate the playback processing of the audio source corresponding to the target audio.
[0043] In one possible implementation of the first aspect, the audio playback client is a browser, and the target audio is audio generated by a multimodal large language model provided by the audio playback server. Of course, the audio playback client may also be another application or electronic device, and the target audio may also be other audio.
[0044] In a second aspect, an embodiment of the present application provides an electronic device, comprising: a memory for storing a computer program, the computer program including program instructions; a processor for executing the program instructions so that the electronic device performs an audio playback method as described in any one of the above embodiments.
[0045] In a third aspect, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program. The computer program includes program instructions, and the program instructions are executed by an electronic device to enable the electronic device to execute the audio playback method provided in the first aspect and / or any possible implementation of the first aspect.
[0046] In a fourth aspect, the implementation of the present application provides a computer program product, including a computer program, which, when executed by an electronic device, implements the audio playback method provided by the above-mentioned first aspect and / or any possible implementation of the first aspect.
[0047] It can be understood that the beneficial effects of the second to fourth aspects mentioned above can also be found in the relevant description of the first aspect mentioned above, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] In order to more clearly illustrate the technical solution of the present application, the following is a brief introduction to the drawings used in the description of the implementation methods.
[0049] Figure 1 A schematic structural diagram of an audio playback system disclosed in an embodiment of the present application;
[0050] Figure 2 A flowchart of the audio playback method disclosed in the embodiment of this application;
[0051] Figure 3 Another flowchart of the audio playback method disclosed in the embodiment of this application;
[0052] Figure 4 for Figure 2 A schematic diagram of an implementation flow of step S1200 in the audio playback method shown;
[0053] Figure 5 for Figure 2 A schematic diagram of an implementation flow of step S1300 in the audio playback method shown;
[0054] Figure 6 Another flowchart of the audio playback method disclosed in the embodiment of this application;
[0055] Figure 7 A schematic diagram of the structure of an electronic device disclosed in an embodiment of the present application. DETAILED DESCRIPTION
[0056] The technical solution of this application will be described in further detail below with reference to the accompanying drawings.
[0057] Take the scenario where a browser plays audio generated by a multimodal large language model as an example. Figure 1 As shown, the audio playback system involved in this scenario includes an audio playback client and an audio playback server. As mentioned above, the audio playback client corresponds to the browser, serving as the front end of audio playback facing the user, and is used to play audio. The audio playback server is the server corresponding to the browser, serving as the back end of audio playback, and is used to provide audio to the browser. Among them, a multimodal large language model is deployed in the server to generate audio. A multimodal large language model refers to a large language model that can understand, process, and output content in multiple modalities such as text, sound, and pictures.
[0058] The process of a browser playing audio generated by a multimodal large language model can be, for example, as follows: a user opens an artificial intelligence (AI) question-and-answer interface through a browser application. After entering a desired "question" through the AI Q&A interface, the backend server corresponding to the front-end browser application quickly responds, generating an "answer" related to the "question" through the multimodal large language model, and the browser, as the front-end, presents the "answer." The "answer" generated by the large language model includes content such as sound, text, and images. Therefore, the browser presents the "answer" by playing sound, displaying text, and images, etc.
[0059] In this scenario, since it takes time for the multimodal large language model to generate the "response," that is, to generate the audio data, the browser also needs time to load the audio data and generate the corresponding audio playback data. Therefore, if the browser waits for all the audio playback data to be generated before playing it to the user, the user will wait for a long time, which can easily lead to a poor user experience. Therefore, there is a need for an audio playback method that can play audio while loading audio data, that is, an audio playback method that supports real-time audio playback.
[0060] The existing technology provides a variety of audio playback methods that support real-time audio playback, such as the Media Source Extensions (MSE) method, or other more complex but relatively compatible methods such as the HyperText Transfer Protocol (HTTP) Live Streaming (HLS) method and the Web Real-Time Communication (WebRTC) method.
[0061] The MSE method is a method based on the Application Programming Interface (API) that can extend the functions of media elements, allowing developers to dynamically load media data streams through the JavaScript programming language to achieve more flexible audio and video playback control.
[0062] Using MSE, developers can implement advanced features such as customized data storage and dynamic adjustment of playback content. This offers significant advantages in implementing adaptive streaming media, such as Dynamic Adaptive Streaming over HTTP (DASH) and HLS. Therefore, the MSE approach offers advantages such as flexible audio loading and management, support for multiple audio formats, and adaptive playback.
[0063] However, on the one hand, the MSE method has different support conditions on different devices and browsers. For example, it has poor support on mobile devices (especially iPhone Operating System, iOS), and has poor compatibility, which limits its use in cross-platform applications. On the other hand, the MSE method requires writing custom logic to manage data storage, process decoders and network data, etc., which has high development complexity and increases the difficulty of development and maintenance. On the other hand, the MSE method is not suitable for the transmission of ultra-low latency real-time audio data, and has the problem of insufficient real-time performance. Therefore, the MSE method has problems such as poor compatibility, high development complexity, high maintenance difficulty and poor real-time performance.
[0064] The HLS method is an HTTP-based streaming media transmission protocol. It describes the audio data path by indexing plain text files (for example, MediaPlaylist, m3u8), supports adaptive bitrate playback, and enables the client to dynamically adjust the playback quality according to network conditions. It is widely used as an audio playback method for video on demand and live broadcast services.
[0065] The HLS method plays audio based on the architecture of HTTP and m3u8 files, making it directly applicable to existing Content Delivery Networks (CDNs) and World Wide Web (Web) servers. The HLS method can dynamically adjust the audio playback quality according to network conditions to improve the smoothness of audio playback. It is natively supported by iOS and the desktop operating system (Macintosh Operating System, macOS). Other platforms can also play it through MSE or third-party tools to achieve good cross-device compatibility. It uses the standard HTTP protocol and is suitable for broadcast and long-playing audio content. Therefore, when playing audio based on the HLS method, it has the advantages of easy deployment, support for adaptive bitrates, good cross-device compatibility, and suitability for large-scale distribution.
[0066] However, the HLS approach requires ensuring that index files are correctly processed, which increases configuration workload. Playback on certain browsers and devices may require integration with MSE or third-party tools. Consequently, the HLS approach requires complex format support and reliance on additional tools.
[0067] WebRTC is an open method for real-time audio and video communication. It enables peer-to-peer audio data transmission between web applications and devices. It provides a low-latency, high-quality communication experience through built-in encoding, decoding, and encryption capabilities. It supports end-to-end encryption to protect data security and is widely used in modern browsers.
[0068] The WebRTC method supports a point-to-point transmission mechanism, achieving a low-latency communication experience. This makes the WebRTC method suitable for real-time audio transmission, such as voice chat, video calls, online collaboration, and real-time interaction. It provides audio data encryption during end-to-end transmission, ensuring the security of audio data transmission. It can dynamically optimize the transmission path and quality of streaming audio data through protocols such as the Real-time Transport Control Protocol (RTCP) and the Interactive Connectivity Establishment Protocol (ICE). In point-to-point transmission, there is no need for relay servers, reducing bandwidth and server costs. Therefore, the WebRTC method has the advantages of low latency, end-to-end encryption, support for dynamic network adjustments, and no need for relay servers.
[0069] However, the WebRTC method requires setting up a signaling server, dealing with issues such as Network Address Translation (NAT) penetration and ICE candidate paths, as well as the compatibility of audio encoders and audio decoders, which increases the difficulty of development and deployment. In the case of poor network conditions or network congestion, the sound quality will be poor when playing the audio source. Therefore, the network bandwidth requirements are relatively high. Mainstream browsers all support the WebRTC method, but there may be inconsistent performance in specific devices or scenarios, and it is not suitable for large-scale broadcast scenarios. It is more suitable for one-to-one or real-time communication with a small number of participants, and cannot meet the real-time communication needs of multiple participants, that is, its scalability is limited. Therefore, the WebRTC method has problems such as poor compatibility, high bandwidth requirements, complex implementation, poor browser compatibility, limited scalability, and poor real-time performance.
[0070] To sum up, how to better implement audio playback in browsers is a technical problem that needs to be solved urgently.
[0071] Based on the above problems, the embodiment of the present application proposes an audio playback method, which can be applied to the above-mentioned audio playback client. Among them, the multimodal large language model included in the audio playback server is used to generate a plurality of streaming audio data corresponding to the "reply" audio (as an example of the target audio) in a streaming manner, and send it to the audio playback client in a streaming manner. The audio playback client receives the plurality of streaming audio data in a streaming manner, and when the streaming audio data is received, the audio playback processing is realized by the audio playback method provided by the embodiment of the present application.
[0072] like Figure 2 As shown, in one implementation of the present application, the audio playback method proposed in the present application includes the following steps.
[0073] Step S1100, receive the first audio data, split the first audio data to obtain multiple second audio data, and store the second audio data in the target data storage array respectively. The first audio data is any one of the multiple streaming audio data corresponding to the target audio sent by the audio playback server.
[0074] Step S1200: Pre-decode the unplayed second audio data stored in the target data storage array to obtain third audio data, create an audio source, and set the third audio data as the audio data of the audio source for storage.
[0075] Step S1300: determine the target audio source to be played, and perform playback processing on the target audio source according to the audio data of the target audio source.
[0076] In this way, based on the audio playback method, the audio playback client receives audio data in a streaming manner during the process of playing the target audio, splits or slices each audio data, stores the second audio data obtained by the splitting or slicing in the target data storage array to realize slice caching, and pre-decodes the stored second audio data to obtain the audio source to be played for playback. The audio can be played more timely, the real-time performance of the audio playback is improved, thereby improving the audio playback effect and thus improving the user experience.
[0077] Furthermore, in the audio playback method, the second audio data is stored in the target data storage array to store the audio data for the first time. The third audio data is set to be stored as the audio data of the audio source to store the audio data for the second time. Thus, a dual storage strategy or double buffering strategy for audio data is implemented in the audio playback method. Therefore, in the process of playing audio based on the audio playback method, while playing the audio source of the current segment, the background will pre-decode to obtain the third audio data of the next audio source to obtain the next audio source, so as to achieve seamless playback of the audio source. Thus, the audio source is generated while playing, which can effectively solve the problem of poor user experience caused by waiting for all streaming audio data to be received and decoded before playing. In addition, the double buffering strategy adopted in the audio playback method can reduce the delay of the playback time of the first frame of the audio source to less than 300ms in some scenarios compared to the existing method of waiting for the complete audio source to be generated before playing, significantly reducing the user waiting time.
[0078] Furthermore, the audio playback method is applied to an audio playback client, so that all processing from receiving the streaming audio data corresponding to the target audio to realizing the target audio playback process is implemented by the audio playback client, without the participation of an audio playback server such as a server, and the implementation logic is simpler. In this way, on the one hand, the interaction steps between the audio playback client and the audio playback server during the audio playback process are effectively reduced, the real-time performance of the playback is improved, and the system power consumption of the audio playback client is reduced. On the other hand, it can also reduce the complexity of system development, reduce the cost of system development, deployment and maintenance, and improve the efficiency of system development and reduce implementation costs. On the other hand, it can also better adapt to different systems and applications.
[0079] The audio playback client may be the aforementioned browser, and the target audio may be the audio generated by the multimodal large language model provided by the aforementioned audio playback server. In summary, the audio playback method provided in the embodiment of the present application can better implement audio playback in a browser.
[0080] Furthermore, in the audio playing method, it is also necessary to execute an initialization process related to audio playing.
[0081] Therefore, if Figure 3 As shown, in one implementation of the present application, the audio playback client needs to execute the following step S1000 before executing step S1100.
[0082] Step S1000: Perform audio playback initialization processing to create variable information required during the audio playback process.
[0083] For example, when it is determined that the target audio needs to be played, or before the target audio is played, or before the aforementioned streaming audio data is received, some variable information required for the subsequent playback of the target audio can be initialized, created, and preset for subsequent playback processing.
[0084] The specific content of the variable information can be determined according to the target audio to be played and / or the subsequent playback process.
[0085] In one implementation of the present application, the variable information may include: a target data storage array, a first data volume threshold, a second data volume threshold, an audio context, a to-be-played audio source variable, a current audio playback position identification variable, a next audio playback position identification variable, an audio decoding state variable, a current audio source variable, a next audio source variable, a complete audio source variable, and a data loading completion mark variable.
[0086] The target data storage array, the first data volume threshold, and the second data volume threshold can be set to fixed default values or determined based on the target audio data volume. Furthermore, the first data volume threshold and the second data volume threshold can be set to, for example, a size that facilitates decoding or a data volume that avoids tearing during audio playback.
[0087] Furthermore, during initialization, the values of the variable for the audio source to be played, the variable for identifying the current audio playback position, the variable for identifying the next audio playback position, the audio decoding state variable, the variable for the current audio source, the variable for identifying the next audio source, the variable for identifying the complete audio source, and the variable for marking the completion of data loading can be set to initial values, such as 0. Then, during audio playback, these variables are assigned corresponding values.
[0088] Furthermore, the first data volume threshold may be a maximum playback segment size limit (maxSize), which may be set to a value greater than 15 KB (i.e., 150,000 bytes) as needed. The second data volume threshold may be a minimum playback segment size limit (minSize), which may be understood as a preset minimum playback threshold, for example, set to 15 KB (i.e., 150,000 bytes) by default.
[0089] In addition, the audio playback position identification variable can be the audio playback index offsetIndex, the audio decoding status variable can be the decoding status lock, the current audio source variable can be the current audio source source, the next audio source variable can be the next audio source nextSource, the complete audio source variable can be the complete audio source allSource, the data loading completion mark variable can be the data loading completion mark loadEnd, etc.
[0090] Furthermore, a corresponding audio context may be determined based on the current playback scene for use in the audio playback process.
[0091] Of course, you can also create variables such as playback duration and the upper limit of the size of the aforementioned audio data piece. Based on these variables, you can achieve normal audio playback.
[0092] The settings of the above variable information are only examples. The settings of the above variable information can be set to other variables and corresponding values based on the specific application scenarios and user needs of the audio playback method.
[0093] After the audio playback client initializes the variables required for audio playback based on the above step S1000 to complete the preliminary preparations for audio playback, it executes step S1100 to enter the audio data input process to segment and store the streaming audio data.
[0094] In one implementation of the present application, regarding the above-mentioned step S1100, first audio data is received, the first audio data is segmented and processed to obtain multiple second audio data, and the second audio data are respectively stored in the target data storage array, which can be specifically implemented based on the following process.
[0095] The first audio data is continuously received in a streaming manner, and the first audio data is segmented and processed according to the preset audio data slice size upper limit maxSize to obtain multiple second audio data, so that the size of the multiple second audio data does not exceed the audio data slice size upper limit maxSize. The second audio data can be understood as an audio data slice, and the audio data slice size upper limit maxSize can be set based on the specific application scenario and requirements. For example, the audio data slice size upper limit maxSize can be set to a fixed value of 65536 bytes to avoid the second audio data being too large, resulting in the subsequent second audio data decoding taking too long. That is, during the pre-decoding process, the second audio data can be decoded quickly, effectively reducing the decoding time of the second audio data.
[0096] Then, the obtained multiple second audio data fragments are stored in a pre-created target data cache array in sequence.
[0097] Furthermore, in the process of receiving the first audio data, the audio playback client needs to determine whether the last streaming audio data among the multiple streaming audio data corresponding to the target audio has been received. When the last streaming audio data among the multiple streaming audio data corresponding to the target audio has been received, the value of a pre-created data loading completion flag variable needs to be set to loading completion (e.g., true).
[0098] In the process of dividing the received first audio data into second audio data based on the above steps and storing the second audio data in the target data storage array, the audio playback client can also execute step S1200 to enter the pre-decoding process to perform pre-decoding processing on the stored audio data.
[0099] like Figure 4 As shown, in one implementation of the present application, with respect to the above-mentioned step S1200, the unplayed second audio data stored in the target data storage array is pre-decoded to obtain third audio data, and an audio source is created, and the third audio data is set as the audio data of the audio source for storage, which can be specifically implemented based on the following steps.
[0100] Step S1201: Determine whether a target pre-decoding condition is met.
[0101] If the target pre-decoding condition is met, step S1202 is executed. If the target pre-decoding condition is not met, the process continues to wait for input of streaming audio data and continues to execute step S1201 to determine whether the target pre-decoding condition is met.
[0102] In one implementation of the present application, determining whether the target pre-decoding condition is met may be determining whether one or more of the following six conditions are currently met: the amount of unplayed second audio data in the target data storage array is greater than or equal to the second data amount threshold; audio streaming playback processing is currently being performed; decoding processing is not currently being performed; reception of multiple streaming audio data corresponding to the target audio streaming media has not been completed; the complete audio source variable has no value; or there is unplayed second audio data in the target data storage array.
[0103] Exemplary, when starting to receive the streaming first audio data of the target audio, it can be determined whether the data amount of the unplayed second audio data present in the target data storage array is greater than or equal to the second data amount threshold value, whether audio streaming media playback processing is currently being performed, and whether decoding processing is not currently being performed, to determine whether to enter the pre-decoding process. The judgment order of these 3 conditions can be, for example, first determining whether the data amount of the unplayed second audio data present in the target data storage array is greater than or equal to the second data amount threshold value, if greater than, then determining whether audio streaming media playback processing is currently being performed, and whether decoding processing is not currently being performed. If it is determined that the data amount of the unplayed second audio data present in the target data storage array is greater than or equal to the second data amount threshold value, audio streaming media playback processing is currently being performed, and decoding processing is not currently being performed, then determining that pre-decoding processing can be performed and enter the pre-decoding process.
[0104] Therefore, when the amount of unplayed second audio data in the target data storage array is greater than or equal to the second data amount threshold and is not being decoded, a pre-decoding process can be triggered.
[0105] However, in most cases, the data has been received before much of the audio data has been played. Therefore, after the reception of the streaming first audio data of the target audio is completed, there may be a situation where the data amount of the unplayed second audio data present in the target data storage array is less than the second data amount threshold. At this time, if the pre-decoding process is triggered only by the storage data amount, there may be a playback interruption caused by not performing the pre-decoding process in time to obtain the audio source. Therefore, it can be further determined whether the reception process of the multiple streaming audio data corresponding to the target audio streaming media has ended, whether the complete audio source variable has no value, and whether the target data storage array still has the unplayed second audio data, so as to determine whether to enter the pre-decoding process. If it is determined that the reception of the multiple streaming audio data corresponding to the target audio streaming media has not ended, the complete audio source variable has no value, and the target data storage array still has the unplayed second audio data, then it is determined that the pre-decoding process can be performed and the pre-decoding process is entered. In this way, it can be guaranteed that the remaining audio data can be played after the data reception is completed.
[0106] Therefore, it is possible to determine to enter the pre-decoding process only when it is determined that all the above six conditions are met, so as to ensure normal processing of the pre-decoding.
[0107] Among them, the complete audio source variable has no value, for example, the value of the complete audio source variable allSource may be 0, indicating that the audio source has not been generated yet, that is, the audio playback of the target audio has not been completed.
[0108] In the implementation of the present application, by judging whether the current playback state meets the target pre-decoding conditions to determine whether to execute step S1202, the unplayed second audio data in the target data storage array can be better pre-decoded to ensure that the target audio can be played completely.
[0109] Of course, the setting of the target pre-decoding condition may be determined based on a specific application scenario of the audio playback method, and may include, but is not limited to, the conditions described in the above possible implementations.
[0110] After determining based on the above steps that the pre-decoding conditions are met and pre-decoding processing is required, the current audio decoding state variable and some variables required for decoding can be set based on step S1202 to complete the pre-decoding preparation.
[0111] Step S1202: Set the audio decoding state variable to decoding, and clear the value of the next audio source variable.
[0112] Setting the audio decoding state variable lock of the audio decoder in the audio playback client to "decoding" prevents the audio decoder from being occupied by other audio data to be decoded, ensuring smooth decoding. Clearing the value of the next audio source variable nextSource facilitates setting the next audio source to be played.
[0113] After the pre-decoding preparation is completed based on step S1202, subsequent pre-decoding processing is performed through the following steps.
[0114] After the preparation is completed, the unplayed second audio data stored in the target data storage array is subjected to subsequent pre-decoding processing to obtain the third audio data, which can be achieved through the following steps.
[0115] Step S1203 : According to the value of the current audio playback position identification variable, the unplayed second audio data is sequentially acquired from the target data storage array, and the amount of the acquired second audio data is accumulated.
[0116] Furthermore, the specific steps of step S1203 may include: calculating the next audio playback position identification variable nextIndex based on the current audio playback position identification variable offsetIndex and the upper limit of the playback segment size maxSize. The specific calculation method can be to define a calculation function, and create a final audio playback position identification variable resultIndex inside the function so that it is equal to the current audio playback position identification variable offsetIndex. Starting from the final audio playback position identification variable resultIndex, the data size of the unplayed second audio data is accumulated to obtain the unplayed audio data segment. Each time an unplayed second audio data is accumulated, the corresponding final audio playback position identification variable resultIndex will be increased by 1.
[0117] In step S1203, the data amounts of the acquired second audio data are accumulated to obtain the sum of the data amounts of the acquired second audio data. When the sum of the data amounts of the acquired second audio data is obtained, the first data amount threshold is compared with the sum of the data amounts of the acquired second audio data, i.e., step S1204 is executed to determine whether to execute subsequent steps.
[0118] Step S1204 determines whether the sum of the acquired second audio data volumes is greater than or equal to the first data volume threshold. If not, step S1203 is continued to acquire the second audio data until the sum of the acquired second audio data volumes is greater than or equal to the first data volume threshold. If so, that is, if it is determined that the sum of the acquired second audio data volumes is greater than or equal to the first data volume threshold, step S1205 is directly executed.
[0119] Furthermore, step S1204 specifically includes the following steps: when the sum of the acquired second audio data amounts is greater than or equal to the upper limit maxSize of the playback segment size (i.e., the first data amount threshold), returning the current final audio playback position identifier variable resultIndex. When the sum of the acquired second audio data amounts is less than the upper limit maxSize of the playback segment size, returning the final audio playback position identifier variable resultIndex+1, and returning to step S1203.
[0120] In this way, the audio data between nextIndex is the audio data of the next audio source, and these audio data segments can be merged into one audio segment as the fourth audio data. Based on step S1204, it is possible to effectively avoid the problem that the audio data segments obtained by accumulating the acquired second audio data are too small, resulting in a small size of the generated audio source, making the audio source playback time too short, resulting in audio playback interruptions and tearing sound during playback, and affecting the user experience.
[0121] Step S1205: Determine whether the acquired second audio data includes the file header information of the target audio. If so, execute step S1206. If not, execute step S1207.
[0122] Furthermore, the specific steps of step S1205 may include: determining whether the acquired second audio data includes the file header information of the target audio according to the value of the current audio playback position identification variable offsetIndex. For example, determining whether the acquired second audio data includes second audio data whose corresponding current audio playback position identification variable offsetIndex is 0, and the current audio playback position identification variable offsetIndex corresponding to the second audio data is less than the next audio playback position identification variable nextIndex. If so, it is considered that the acquired second audio data includes the file header information of the target audio; if not, it is considered that the acquired second audio data does not include the file header information of the target audio.
[0123] Because audio decoding requires audio file header information, which is present in the first audio data, failing to include this information will result in an error during audio decoding. Therefore, determining whether the acquired second audio data includes the file header information of the target audio can ensure that the audio data segment obtained by accumulating the acquired second audio data contains the file header information of the target audio, thus avoiding errors during audio data segment decoding and affecting the decoding of the audio data segment.
[0124] Step S1206: If the acquired second audio data includes the file header information of the target audio, synthesize the acquired second audio data to obtain fourth audio data.
[0125] Step S1207: If the acquired second audio data does not include the file header information of the target audio, the acquired second audio data and the first second audio data of the target audio stored in the target data storage array are synthesized to obtain fourth audio data.
[0126] Furthermore, synthesizing the acquired second audio data and the first second audio data of the target audio stored in the target data storage array may include: splicing the first second audio data of the target audio stored in the target data storage array in front of the acquired second audio data in the arrangement direction of the acquired second audio data.
[0127] Furthermore, data synthesis based on the above steps can ensure that the obtained fourth audio data contains the file header information in the target audio, so as to avoid errors when decoding the fourth audio data, which affects the decoding of the fourth audio data.
[0128] Furthermore, the fourth audio data is decoded through the following steps.
[0129] Step S1208: Decode the fourth audio data using a target decoding method corresponding to the audio context to obtain third audio data.
[0130] For example, decoding is performed using a target decoding method (eg, decodeAudioData method) corresponding to the audio context to obtain decoded data, namely, the third audio data.
[0131] Then, the following step S1209 is executed to create an audio source.
[0132] Step S1209: Create an audio source, and set the third audio data as the audio data of the audio source.
[0133] For example, a new audio source is created through the target audio source creation method corresponding to the audio context (for example, the createBufferSource method), and the decoded data, that is, the third audio data, is set as the audio data of the audio source.
[0134] Then, the audio source is set to the next audio source variable nextSource, and the current audio playback position identification variable offsetIndex is updated to the next audio playback position identification variable nextIndex, that is, the value of the current audio playback position identification variable offsetIndex is updated to the value of the next audio playback position identification variable nextIndex.
[0135] Furthermore, the created audio source is the audio source to be played, and the audio source can be stored in the audio source list to be played. Setting the third audio data as the audio data of the audio source can be assigning the third audio data to the audio source variable.
[0136] Furthermore, if it is determined that the last streaming audio data among the multiple streaming audio data corresponding to the target audio is received, the audio source obtained based on the above steps is set to the complete audio source variable. If it is determined that the last streaming audio data among the multiple streaming audio data corresponding to the target audio is not received, the audio source obtained based on the above steps is set to the next audio source variable.
[0137] In the above steps, the fourth audio data is decoded using the target decoding method corresponding to the audio context, thereby improving the efficiency of the audio data decoding process and enhancing the operational flexibility of the audio data decoding process.
[0138] Further, in another implementation of the present application, when the multiple streaming audio data corresponding to the target audio have finished receiving, that is, when determining to receive the last streaming audio data in the multiple streaming audio data corresponding to the target audio, the unplayed second audio data stored in the target data storage array are pre-decoded to obtain the third audio data, or all unplayed second audio data stored in the target data storage array are pre-decoded to obtain the third audio data. Then a corresponding audio source is created, and the created audio source is set to a complete audio source variable. In addition, as previously mentioned, when the multiple streaming audio data corresponding to the target audio have finished receiving, the value of the data loading completion flag variable corresponding to the target audio is also required to be set to loading completion (e.g., true). In this way, final audio source can be generated in a disposable manner to improve audio playback efficiency when all streaming audio data are received.
[0139] Furthermore, when there is an audio source to be played, the audio playback client may also execute the aforementioned step S1300, enter the audio source playback process, confirm the audio source, play the audio source, and perform playback processing on the obtained audio source.
[0140] Furthermore, before executing step S1300, that is, before determining the target audio source to be played, if it is determined that the audio source playback process is not currently in progress, an audio source playback method (e.g., the play method) is executed to determine whether the audio context is currently in an active state. If the audio context is not currently in an active state, the audio context is enabled; if the audio context is currently in an active state, the audio source playback process is entered to determine the target audio source to be played.
[0141] In this way, the audio context can be adjusted to the working state. After the audio context is adjusted to the working state based on the above steps, step S1300 is executed to directly enter the audio source playback process.
[0142] like Figure 5 As shown, in one implementation of the present application, the target audio source to be played is determined, and the target audio source is played according to the audio data of the target audio source, that is, the specific steps of executing step S1300 may include the following steps.
[0143] Step S1301: Initialize the audio source variable corresponding to the audio source to be played.
[0144] By initializing the audio source variable source corresponding to the audio source to be played, the value of the audio source variable corresponding to the audio source to be played is cleared, so that in the process of playing the target audio based on the audio playback method, the audio source variable to be played can be assigned a value to ensure accurate playback of the audio.
[0145] After executing step S1301, it is determined whether the current complete audio source variable has a value. That is, step S1302 is executed. If the value of the data loading completion flag variable is loading completed, all unplayed audio sources corresponding to the target audio, that is, the complete audio sources, are played as the audio sources to be played.
[0146] Step S1302: Determine whether the complete audio source variable has a value. If so, proceed to step S1303. If not, proceed to step S1304.
[0147] After executing step S1302, if the complete audio source variable allSource has a value, it means that all audio sources corresponding to the current target audio have been generated and there is no need to generate new audio sources. Therefore, if the complete audio source variable has a value, step S1303 is directly executed. If the complete audio source variable has no value, it means that the first audio data corresponding to the current target audio has not been received, the audio source corresponding to the target audio is still being generated, and new audio sources are being generated. Therefore, if the complete audio source variable has no value, step S1304 is directly executed to determine whether the next audio source variable has a value.
[0148] Step S1303: If there is a value, the value of the audio source variable to be played is set to the value of the complete audio source variable to determine the target audio source, and the value of the next audio source variable is cleared, and the value of the current audio playback position identification variable is updated to the length of the target data storage array.
[0149] During step S1303, the value of the variable "source" (the audio source to be played) is set to the value of the variable "allSource" (the complete audio source). This indicates that all the streaming audio data in the target audio has been received and decoded. Therefore, no new audio source is generated for the target audio. Therefore, the value of the variable "next audio source" (the source of the audio source to be played) is cleared, and the value of the variable "current audio playback position identifier" (the position of the audio playback position) is updated to the length of the target data storage array to prevent audio source playback errors.
[0150] After executing step S1303 , the target audio source has been determined, so step S1309 is directly executed to play the target audio source.
[0151] In step S1304, if there is no value, it is determined whether the next audio source variable has a value. If so, step S1305 is executed. If not, step S1306 is executed.
[0152] After executing step S1304, if the next audio source variable nextSource has a value, it means that there is an audio source to be played, and step S1305 is executed. If the next audio source variable has no value, it means that there is no audio source to be played, and step S1306 is executed.
[0153] Step S1305: If there is a value, the value of the audio source variable source to be played is set to the value of the next audio source variable nextSource to determine the target audio source.
[0154] After executing step S1305, the target audio source has been determined, so step S1315 is directly executed to play the target audio source.
[0155] If there is no value in step S1306, it is determined whether decoding is currently in progress. If so, step S1307 is executed. If not, step S1308 is executed.
[0156] If the next audio source variable has no value, indicating that there is no audio source to be played, then a determination is made based on step S1306 as to whether decoding is currently in progress. If decoding is currently in progress, step S1307 is performed to determine the target audio source. If decoding is not currently in progress, step S1308 is performed to determine the target audio source.
[0157] Step S1307: If yes, wait for the decoding process to be completed, set the value of the audio source variable source to be played to the value of the next audio source variable nextSource to determine the target audio source.
[0158] In step S1307, wait for the decoding process to be completed to obtain audio data, obtain a new audio source based on the audio data obtained after the decoding process, use the new audio source as the next audio source, set the value of the audio source variable to be played to the value of the next audio source variable to determine the target audio source.
[0159] After executing step S1307, the target audio source has been determined, so step S1309 is directly executed to play the target audio source.
[0160] Step S1308, if not, then after waiting for the sum of the data amounts of the unplayed second audio data stored in the target data storage array to be greater than or equal to the second data amount threshold, all the currently stored unplayed second audio data are taken out from the target data storage array, decoded and processed to obtain the third audio data, create an audio source, set the third audio data as the audio data of the audio source, and set the audio source to the audio source variable to be played to determine the target audio source.
[0161] For example, if the total amount of data from the current audio playback position identifier variable offsetIndex to the last data in the target data cache array is greater than the aforementioned playback segment lower limit minSize, then all data from offsetIndex onwards is taken to generate an audio source. Then, the variable source of the audio source to be played is set to this new audio source.
[0162] Based on this step, it can be ensured that when the network conditions are poor, enough data will be accumulated before playback, avoiding the situation where the tearing sound occurs when the data volume is too small.
[0163] After executing step S1308, the target audio source to be played can be determined, so step S1309 is directly executed to play the target audio source.
[0164] Step S1309: Set the global current audio source variable to the audio source, bind the playback end monitoring, and perform playback processing on the target audio source.
[0165] In the process of determining the target audio source to be played based on the above steps and playing the target audio source based on the audio data of the target audio source, multiple determinations are performed to determine the target audio source, thereby ensuring that the obtained target audio source is more accurate and that the playback order of the audio sources matches the reception order of the first audio data corresponding to the audio sources. Furthermore, multiple determinations to determine the target audio source ensure uninterrupted playback of the audio corresponding to the audio source.
[0166] Furthermore, by initializing the audio source variable to be played corresponding to the audio source to be played, the value of the audio source variable to be played corresponding to the audio source to be played is cleared, so that the audio source variable to be played can be assigned a value during the process of playing the target audio based on the audio playback method.
[0167] After playing the target audio source based on the above steps, the audio playback client can also determine whether the target audio is played.
[0168] like Figure 6 As shown, in a possible implementation of the first aspect above, after playing the target audio source, the audio playing method further includes the following steps.
[0169] Step S1400, after playing the target audio source, determine whether the streaming audio data corresponding to the target audio has been received through the data loading completion flag variable, and determine whether there is any unplayed audio data in the target data storage array.
[0170] In step S1500 , if the reception of the streaming audio data corresponding to the target audio is completed and there is no unplayed audio data in the target data storage array, the playback process for the target audio is terminated.
[0171] Step S1600: If the reception of the streaming audio data corresponding to the target audio has not been completed, or there is unplayed audio data in the target data storage array, then the audio playback process for the target audio continues.
[0172] In this way, after a single audio source finishes playing, the data loading completion flag variable loadEnd is used to determine whether data reception has been completed and whether the buffered audio data in the target data storage array has been played. If so, streaming playback of the target audio has ended. Otherwise, the next audio source will be played, and the audio source playback process will be repeated until streaming playback of the target audio is completed. This ensures the complete playback of the target audio.
[0173] The process of audio playback client to implement audio playback mainly includes the above-mentioned initialization process, data input process, pre-decoding process and audio source playback process.
[0174] The following takes the example of a web browser in a computer playing audio based on the audio playback method in the above embodiment to further illustrate the audio playback provided by the embodiment of the present application.
[0175] The user opens the artificial intelligence (AI) question-and-answer interface through a web browser. After entering the "question" they want to ask through the AI question-and-answer interface, the back-end server corresponding to the front-end web browser will respond quickly, and continuously generate "answer" audio related to the "question" in a streaming manner through a multimodal large language model, and continuously send the "answer" audio to the front-end web browser in a streaming manner until all "answer" audio is sent.
[0176] In addition, when the user opens the AI question-and-answer interface or enters a "question", the web browser can enter the initialization process and complete the initialization processing through the aforementioned step S1000 to create the aforementioned variable information.
[0177] Furthermore, after the front-end web browser is initialized, it enters the data input process, and through the aforementioned step S1000, for example, using the fetch application programming interface (English: Application Programming Interface, API), it continuously receives the streaming "reply" audio blocks sent by the multimodal large language model in a streaming manner, and the streaming "reply" audio blocks are binary data. Then, the received streaming "reply" audio blocks are passed to the input method of the streaming playback object (i.e., the browser) in the format of, for example, Uint8Array. The player uses a buffer management mechanism to fragment and cache each received streaming "reply" audio. In the process of receiving the streaming "reply" audio, when the cumulative received audio data reaches the aforementioned preset minimum playback threshold (default 15KB), the system can automatically trigger the decoding and playback process, first entering the pre-decoding process, and performing pre-decoding processing on the cached "reply" audio through the aforementioned steps S1201-step S1209 to create an audio source. Then, the audio source playback process is entered, and the playback processing of the audio source is implemented according to the aforementioned S1301-step S1309. In this way, a double-buffering strategy is adopted. While the current audio source segment is playing, the background will pre-decode the next audio data segment, achieving seamless playback. In other words, audio data can be cached while pre-decoding processing is performed while the audio source is played, achieving a "generation and playback" effect until all audio data is played.
[0178] Furthermore, the process of implementing audio playback by the audio playback client may further include: a playback completion determination process, that is, after playing the target audio source, determining whether the target audio is played completely.
[0179] In summary, this solution is implemented purely on the front end, and is a method for front-end streaming audio data playback. This method does not require special back-end configuration and can achieve real-time reception and playback of audio data. Compared with the traditional method of waiting for the complete audio to be generated before playback, this solution can reduce the first frame playback delay to less than 300ms, significantly reducing user waiting time, which significantly improves the user experience. Compared with the aforementioned streaming technologies such as WebRTC and HLS, this solution has the advantage of being lightweight, does not require additional server support, and is specifically optimized for audio scenarios, avoiding the additional overhead brought by video streaming. In addition, the system has a built-in exception handling mechanism. When the amount of audio data received is small due to network reasons, it can wait until the audio data volume meets the conditions before decoding, playing, and other processing. It can effectively deal with abnormal situations such as network fluctuations and decoding failures, and ensure the stability of playback. Through this solution, developers can easily implement the real-time playback function of audio output by multimodal models, providing users with a smooth interactive experience while reducing system deployment and maintenance costs.
[0180] The method for implementing streaming audio data playback on the front end provided by the implementation of this application can be applied to the audio playback of the above-mentioned multimodal large language model, and can also be applied to audio playback in other scenarios.
[0181] In one implementation of the present application, the audio playback client can be set in a mobile terminal (such as a mobile phone), computer and other electronic devices. It is a web application such as a web browser used to play audio in these electronic devices, or a first-level related component of a web application such as a web browser. The audio playback server can be set in the server.
[0182] In other embodiments of the present application, the audio playback client and the audio playback server can also be set in mobile terminals (for example, mobile phones), computers and other electronic devices. For example, they can be different applications or functional modules in electronic devices, which can be set as needed.
[0183] Of course, in other implementations of the present application, the audio playback client may also be a mobile terminal (such as a mobile phone), a computer or other electronic device, and the audio playback server may be a server, etc.
[0184] like Figure 7 As shown, an embodiment of the present application further provides an electronic device, comprising: a memory 123 for storing a computer program, the computer program including program instructions; a processor 122 for executing program instructions so that the electronic device executes an audio playback method as any one of the above embodiments.
[0185] The processor 122 executes the computer-executable instructions stored in the memory, so that the processor 122 performs the technical solution of the audio playback method in the above embodiment. The processor 122 can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital data processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.
[0186] The memory 123 is connected to the processor 122 via a system bus and communicates with the processor 122. The memory 123 is used to store computer program instructions.
[0187] The transceiver 121 may be used to obtain tasks to be executed and configuration information of the tasks to be executed.
[0188] The system bus can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, among others. The system bus can be divided into an address bus, a data bus, a control bus, and so on. For ease of illustration, the figure shows only one thick line, but this does not imply that there is only one bus or only one type of bus. Transceivers are used to enable communication between the database access device and other computers (e.g., clients, read-write libraries, and read-only libraries). Memory may include random access memory (RAM) and non-volatile memory.
[0189] The implementation of the present application also provides a chip for running computer instructions / programs, which is used to execute the technical solution of the above-mentioned audio playback method.
[0190] Another embodiment of the present application further discloses a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the audio playback method in any of the above embodiments.
[0191] Another embodiment of the present application further discloses a computer program product, including a computer program, which, when executed by a processor, implements the audio playback method in any of the above embodiments.
[0192] It should be noted that the terms "first", "second", etc. are only used for distinction and description, and cannot be understood as indicating or implying relative importance.
[0193] It should be noted that in the accompanying drawings, some structural or method features may be shown in a specific arrangement and / or order. However, it should be understood that such a specific arrangement and / or order may not be required. Rather, in some embodiments, these features may be arranged in a manner and / or order different from that shown in the illustrative drawings. In addition, the inclusion of structural or method features in a particular figure does not imply that such features are required in all embodiments, and in some embodiments, these features may not be included or may be combined with other features.
[0194] Although the present application has been illustrated and described with reference to certain preferred embodiments of the present application, those skilled in the art will appreciate that the above descriptions are provided as further details of the present application in conjunction with specific embodiments, and that the specific implementation of the present application should not be limited to these descriptions. Those skilled in the art may make various changes in form and detail, including simple deductions or substitutions, without departing from the spirit and scope of the present application.
Claims
1. An audio playback method, characterized in that: Applied to an audio playback client, the method includes: Receive first audio data, segment the first audio data to obtain multiple second audio data, and store the second audio data in a target data storage array, wherein the first audio data is any one of multiple streaming audio data corresponding to the target audio sent by the audio playback server; Performing pre-decoding processing on the unplayed second audio data stored in the target data storage array to obtain third audio data, creating an audio source, and setting the third audio data as the audio data of the audio source for storage; A target audio source to be played is determined, and playback processing is performed on the target audio source according to audio data of the target audio source.
2. The audio playback method according to claim 1, wherein: Performing pre-decoding processing on the unplayed second audio data stored in the target data storage array to obtain third audio data, including: sequentially acquiring the unplayed second audio data from the target data storage array according to the value of the current audio playback position identification variable, and accumulating the amount of the acquired second audio data; If the sum of the data amounts of the acquired second audio data is greater than or equal to a first data amount threshold, determining whether the acquired second audio data includes the file header information of the target audio; If the acquired second audio data includes the file header information of the target audio, synthesizing the acquired second audio data to obtain fourth audio data; If the acquired second audio data does not include the file header information of the target audio, synthesizing the acquired second audio data with the first second audio data corresponding to the target audio stored in the target data storage array to obtain fourth audio data; The fourth audio data is decoded to obtain the third audio data.
3. The audio playback method according to claim 2, wherein: Performing pre-decoding processing on the unplayed second audio data stored in the target data storage array to obtain third audio data, further comprising: If the target pre-decoding condition is met, the unplayed second audio data stored in the target data storage array is pre-decoded to obtain third audio data, wherein: When it is determined that one or more of the following conditions are met, it is determined that the target pre-decoding condition is met: The amount of the unplayed second audio data in the target data storage array is greater than or equal to a second data amount threshold; Currently processing audio playback; Decoding is not currently in progress; The receiving of the plurality of streaming audio data corresponding to the target audio is not completed; The complete audio source variable has no value; The target data storage array contains the second audio data that has not been played.
4. The audio playback method according to claim 3, characterized in that: Before performing pre-decoding processing on the unplayed second audio data stored in the target data storage array, the method further includes: Set the audio decoding state variable to decoding and clear the value of the next audio source variable; After creating the audio source, the method further includes: The created audio source is set to the next audio source variable, and the value of the current audio playback position identification variable is updated to the value of the next audio playback position identification variable.
5. The audio playback method according to claim 4, characterized in that: Determining whether the acquired second audio data includes file header information of the target audio includes: determining, according to the value of the current audio playback position identifier variable, whether the acquired second audio data includes the file header information of the target audio; Decoding the fourth audio data to obtain the third audio data includes: The fourth audio data is decoded using a target decoding method corresponding to the audio context to obtain the third audio data.
6. The audio playback method according to claim 5, characterized in that: When the reception of the plurality of streaming audio data corresponding to the target audio is completed, Performing pre-decoding processing on the unplayed second audio data stored in the target data storage array to obtain third audio data, including: performing the pre-decoding process on all the unplayed second audio data stored in the target data storage array to obtain the third audio data; Furthermore, the method further comprises: Set the created corresponding audio source to the complete audio source variable, and The value of the data loading completion flag variable corresponding to the target audio is set to loading completion.
7. The audio playback method according to claim 6, characterized in that: Determine the target audio source to be played, including: Initialize the audio source variable corresponding to the audio source to be played; Determine if the complete audio source variable has a value; If there is a value, the value of the to-be-played audio source variable is set to the value of the complete audio source variable to determine the target audio source, and the value of the next audio source variable is cleared, and the value of the current audio playback position identifier variable is updated to the length of the target data storage array; If there is no value, determine whether the next audio source variable has a value; If there is a value, setting the value of the to-be-played audio source variable to the value of the next audio source variable to determine the target audio source; If there is no value, determine whether decoding is currently in progress; If so, wait for the decoding process to be completed, and set the value of the to-be-played audio source variable to the value of the next audio source variable to determine the target audio source; If not, after waiting for the sum of the data amounts of the unplayed second audio data stored in the target data storage array to be greater than or equal to the second data amount threshold, all the currently stored unplayed second audio data are taken out from the target data storage array, decoded and processed to obtain third audio data, and the audio source is created. The third audio data is set as the audio data of the audio source for storage, and the audio source is set to the audio source variable to be played to determine the target audio source.
8. The audio playback method according to claim 7, characterized in that: Before determining the target audio source to be played, the method further includes: If it is determined that the audio source playback process is not currently in progress, the audio source playback method is executed to determine whether the audio context is currently in a working state; If the audio context is not currently in a working state, setting the audio context to a working state; If the audio context is currently in a working state, then entering the audio source playing process to determine the target audio source to be played; Playing the target audio source includes: Set the global current audio source variable to the target audio source and bind the playback end monitor to play the target audio source; After playing the target audio source, the method further includes: Determining whether the plurality of streaming audio data corresponding to the target audio have been received through the value of a data loading completion flag variable, and determining whether unplayed second audio data exists in the target data storage array; If the reception of the plurality of streaming audio data corresponding to the target audio is completed and there is no unplayed second audio data in the target data storage array, then the playback process for the target audio is terminated; If the reception of the multiple streaming audio data corresponding to the target audio has not been completed, or there is unplayed second audio data in the target data storage array, the playback processing for the target audio is continued to play the next target audio source.
9. The audio playback method according to any one of claims 1 to 8, characterized in that: The method further comprises: Before receiving the first audio data, audio playback initialization processing is performed, and variable information for audio playback processing is created, wherein the variable information at least includes the target data storage array, as well as a first data volume threshold, a second data volume threshold, an audio context, a to-be-played audio source variable, a current audio playback position identification variable, a next audio playback position identification variable, an audio decoding state variable, a current audio source variable, a next audio source variable, a complete audio source variable, and a data loading completion mark variable.
10. An electronic device, characterized in that: include: a memory for storing a computer program, wherein the computer program includes program instructions; The processor is configured to execute the program instructions so that the electronic device implements the audio playback method according to any one of claims 1 to 9.