Audio playing method, device, electronic device and storage medium

By dynamically adjusting the audio code stream and splicing according to the required parameters of the receiving end, the noise problem caused by the switching of the code stream during audio playback is solved, and the user experience and device utilization is improved.

CN111798858BActive Publication Date: 2025-07-18TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202010635495.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-07-03
Publication Date
2025-07-18
Estimated Expiration
2040-07-03

AI Technical Summary

Technical Problem

In the existing audio playback technology, due to the differences in network quality and playback equipment at the receiving end, the audio code stream has noise problems during switching, which affects the user experience and cannot effectively meet the audio needs of different receiving ends.

Method used

Dynamically adjust the audio code stream according to the audio demand parameters of the receiver, and generate the switched audio code stream through splicing processing to avoid noise caused by direct switching, and ensure that the code stream matches the demand of the receiver.

Benefits of technology

It realizes dynamic adjustment of audio code stream without affecting the user experience, improves network bandwidth and utilization of playback devices, meets the audio playback needs of the receiver, and avoids bandwidth waste and long response time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN111798858B_ABST
    Figure CN111798858B_ABST
Patent Text Reader

Abstract

The present application relates to the field of audio technologies, and discloses an audio playing method, apparatus, electronic device, and storage medium. Among them, the audio playing method includes: determining an audio bitstream to be sent to a receiving end according to audio requirement parameters of the receiving end; when it is detected that the audio requirement parameters are updated during the playing of the audio bitstream, determining an updated audio bitstream according to the updated audio requirement parameters; performing bitstream splicing processing based on the audio bitstream and the updated audio bitstream to obtain a switched audio bitstream; and sending the switched audio bitstream to the receiving end, so that the receiving end receives the switched audio bitstream and decodes and plays the switched audio bitstream. By using the solution provided by the present application, it is possible to dynamically adjust the audio bitstream according to changes in audio requirement parameters, and achieve smooth switching of the audio bitstream under different audio requirement parameters.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of audio technology. Specifically, this application relates to an audio playback method, apparatus, electronic device, and storage medium. Background Art

[0002] With the development of the Internet, users can obtain and play audio at any time through audio sharing websites, such as news, film and television works, etc. Currently, the processing flow of audio playback is as follows: The sending end collects audio signals through a microphone, and the sending end performs audio encoding according to a set of preset encoding parameters. The generated audio bitstream is sent to the server through the network, and the server sends the corresponding bitstream to each receiving client. After receiving the bitstream, the receiving client decodes and plays it.

[0003] In the existing audio playback process, an audio bitstream is generated according to fixed encoding parameters, and all receiving clients receive the same bitstream. However, there are differences in network quality among different receiving clients. In the case of good network quality, the network bandwidth is relatively sufficient and there are basically no packet losses. However, in the case of poor network quality, the network bandwidth is limited and packet losses often occur. Moreover, there are differences in the audio playback capabilities of the playback devices of the receiving clients. Some devices can support high-sampling-rate high-quality audio playback, while some devices have limited playback capabilities and can only support the playback of lower audio sampling rate signals. With the adjustment of factors such as network bandwidth and playback devices, the audio signal that was originally suitable for adapting to factors such as network bandwidth and playback devices may not be able to play normally on the adjusted device. If the bitstream generated by forcibly switching different encoding parameters is used, obvious noise problems may occur, affecting the user's listening experience. Summary of the Invention

[0004] The purpose of this application is to at least solve one of the above technical defects, and the following technical solutions are specifically proposed:

[0005] In one aspect of this application, an audio playback method is provided, including:

[0006] Determine the audio bitstream to be sent to the receiving end according to the audio requirement parameters of the receiving end;

[0007] When it is detected that the audio requirement parameters are updated during the playback of the audio bitstream, determine an updated audio bitstream according to the updated audio requirement parameters;

[0008] Perform bitstream splicing processing based on the audio bitstream and the updated audio bitstream to obtain a switched audio bitstream;

[0009] Send the switched audio bitstream to the receiving end, so that the receiving end receives the switched audio bitstream and decodes and plays the switched audio bitstream.

[0010] Another aspect of the present application provides an audio playback device, which includes:

[0011] An audio bitstream module, configured to determine an audio bitstream to be sent to the receiving end according to the audio requirement parameters of the receiving end;

[0012] An updated audio bitstream module, configured to determine an updated audio bitstream according to the updated audio requirement parameters when it is detected that the audio requirement parameters are updated during the playback of the audio bitstream;

[0013] A switched audio bitstream module, configured to perform bitstream splicing processing based on the audio bitstream and the updated audio bitstream to obtain a switched audio bitstream;

[0014] A played audio bitstream module, configured to send the switched audio bitstream to the receiving end, so that the receiving end receives the switched audio bitstream and decodes and plays the switched audio bitstream.

[0015] Another aspect of the present application provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the audio playback method shown in the first aspect of the present application.

[0016] Another aspect of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the audio playback method shown in the first aspect of the present application.

[0017] The beneficial effects brought by the technical solution provided by the present application are:

[0018] The audio playback method provided by the present application performs splicing processing based on the audio bitstream before the update of the audio requirement parameters and the updated audio bitstream to obtain a switched audio bitstream. The switched audio bitstream can be realized by splicing, such as at an appropriate timing and in a special splicing manner. The spliced switched audio bitstream is different from the updated audio bitstream, avoiding directly switching from the audio bitstream to the updated audio bitstream, and thus avoiding problems such as noise caused by forced bitstream switching, and improving the user experience.

[0019] The audio playback method provided by the present application determines the finally played audio bitstream according to the audio requirement parameters of the receiving end. The finally decoded and played audio bitstream is adapted to the audio requirement parameters of the receiving end, that is, the sent audio bitstream meets the audio playback requirements of the receiving end, such as meeting the playback parameters, network bandwidth, user-defined audio parameters, etc. of the receiving end, realizing dynamic adjustment of the sent audio bitstream according to the playback requirements of the receiving end, being able to avoid bandwidth waste and too long response time of the audio bitstream to be sent, improving the utilization rate of the playback device and the bandwidth, and being able to meet the real-time requirements of audio playback, and being applicable to the audio live broadcast scenario.

[0020] The audio playing method provided by this application, when it detects that the audio requirement parameters change, detects the audio frame sample values in the audio bitstream before the update. If the audio frame sample values meet the switching condition, the updated audio bitstream is used as the audio bitstream to be sent, realizing the switching of the audio bitstream. It realizes that during the audio playing process, the audio bitstream to be sent is dynamically switched according to the adjustment of the audio requirement parameters at the receiving end. Moreover, the switching condition is that the energy value represented by the audio frame sample values is lower than a preset threshold. The audio frames with energy values lower than the preset threshold are low-energy frames or silent frames. This application switches the audio bitstream at low-energy frames or silent frames, which can avoid obvious noise during the audio bitstream switching process and realize the smooth switching of the audio bitstream.

[0021] The additional aspects and advantages of this application will be partially given in the following description, which will become obvious from the following description or be understood through the practice of this application. Brief Description of the Drawings

[0022] The above-mentioned and / or additional aspects and advantages of this application will become obvious and easy to understand from the following description of the embodiments in conjunction with the drawings, where:

[0023] Figure 1 is an application scenario diagram of the audio playing method provided by an embodiment of this application;

[0024] Figure 2 is a flowchart of the audio playing method provided by an embodiment of this application;

[0025] Figure 3 is a flowchart of the audio bitstream switching provided by an embodiment of this application;

[0026] Figure 4-1 is a spectrogram of switching audio bitstreams corresponding to different coding parameters under unconstrained conditions provided by an embodiment of this application;

[0027] Figure 4-2 is a spectrogram of switching audio bitstreams corresponding to different coding parameters under constraints provided by an embodiment of this application;

[0028] Figure 5 is a flowchart of the audio playing method provided by another embodiment of this application;

[0029] Figure 6 is a schematic diagram of the audio bitstream matrix provided by an embodiment of this application;

[0030] Figure 7 is a flowchart of switching the audio bitstream provided by another embodiment of this application;

[0031] Figure 8A structural schematic diagram of an audio playback device provided by an embodiment of the present application;

[0032] Figure 9 A structural schematic diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0033] The embodiments of the present application will be described in detail below. The examples of the embodiments are shown in the accompanying drawings, in which the same or similar reference numerals indicate the same or similar elements or elements having the same or similar functions throughout. The embodiments described below by referring to the accompanying drawings are exemplary and are only used to explain the present application, and cannot be construed as a limitation to the present application.

[0034] Those skilled in the art of the present technology can understand that unless specifically stated otherwise, the singular forms "a", "an", "the", and "said" used herein may also include the plural forms. It should be further understood that the term "including" used in the specification of the present application means the presence of features, integers, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or their groups. It should be understood that when we say that an element is "connected" or "coupled" to another element, it can be directly connected or coupled to other elements, or there may also be intermediate elements. In addition, the "connection" or "coupling" used herein may include wireless connection or wireless coupling. The phrase "and / or" used herein includes all or any unit and all combinations of one or more related listed items.

[0035] Those skilled in the art can understand that unless otherwise defined, all terms (including technical terms and scientific terms) used herein have the same meaning as the general understanding of those of ordinary skill in the art to which the present invention belongs. It should also be understood that terms such as those defined in a general dictionary should be understood to have a meaning consistent with the meaning in the context of the prior art, and will not be interpreted with an idealized or overly formal meaning unless specifically defined as here.

[0036] Audio live broadcast is to listen to the same audio program on different communication platforms through the Internet network system, and listeners can communicate with the anchor in voice through a wired connection.

[0037] Sampling rate: The number of samples extracted from a continuous signal per second and composed into a discrete signal, and the unit is expressed in hertz (HZ).

[0038] Frame rate is a measure for measuring the number of displayed frames, and the measurement unit is the number of frames displayed per second or hertz.

[0039] Bit rate is the data traffic used by an audio file per unit time, and the unit of this parameter is usually kilobits per second (kbps).

[0040] In the process of research, the inventors found that even if the server stores bitstreams of different qualities, due to the encoding principle of most audio encoders, the generated bitstreams have inter-frame correlation, that is, the bitstreams generated under different encoding parameters cannot be switched randomly. If the audio bitstreams with different encoding parameters are forced to be switched, the decoder will produce abnormal audio signals, and obvious noise may occur, affecting the user's listening experience.

[0041] Regarding the technical problems existing in the prior art, the audio playback method, device, electronic device, and storage medium provided in this application aim to solve at least one of the above technical problems in the prior art. It can determine the corresponding audio bitstream according to the audio requirements of the receiving end, so as to improve the utilization rate of network bandwidth and the receiving-end playback device, and can also switch the audio bitstream according to the update of the audio requirement parameters without affecting the user experience.

[0042] The following uses specific embodiments to elaborate in detail on the technical solution of this application and how the technical solution of this application solves the above technical problems. These several specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of this application will be described below in conjunction with the drawings.

[0043] Figure 1 It is an application scenario diagram of the audio playback method provided in an embodiment of this application. This scenario includes: a server side, a sending end, and a receiving end. The sending end and the receiving end can both be set on the client side. The sending end uploads the locally collected source audio signal to the server. The server side processes the source audio signal to obtain the processed audio bitstream, determines the audio bitstream to be sent according to the audio requirements uploaded by the receiving end, and sends the audio bitstream to be sent to the receiving end. The receiving end receives the audio bitstream sent by the server and decodes and plays it. When it detects that the audio requirement parameters are updated, it obtains the updated audio bitstream corresponding to the updated audio requirement parameters, performs splicing processing on the audio bitstream and the updated audio bitstream to obtain the switched audio bitstream, and sends the switched audio bitstream to the receiving end, so as to determine the corresponding audio bitstream according to the audio requirements of the receiving end, and also realize dynamically adjusting the audio bitstream sent to the receiving end according to the adjustment of the audio requirement parameters, improving the utilization rate of network bandwidth, the receiving-end playback device, and the real-time performance of switching.

[0044] A possible implementation manner is provided in the embodiments of this application, as Figure 2 shown. A kind of audio playback method is provided. This solution can be executed on the server side and includes the following steps:

[0045] S210, determining the audio bitstream sent to the receiving end according to the audio requirement parameters of the receiving end;

[0046] S220, when it is detected that the audio requirement parameters are updated during the playback of the audio bitstream, determine the updated audio bitstream according to the updated audio requirement parameters;

[0047] S230, perform splicing processing on the audio bitstream based on the audio bitstream and the updated audio bitstream to obtain a switched audio bitstream;

[0048] S240, send the switched audio bitstream to the receiving end so that the receiving end receives the switched audio bitstream and decodes and plays the switched audio bitstream.

[0049] First, determine the audio bitstream sent to the receiving end according to the audio requirement parameters of the receiving end, and this audio bitstream is adapted to the current audio requirement parameters of the receiving end. Then, continuously detect the audio requirement parameters of the receiving end. If it is detected that the audio requirement parameters of the receiving end are updated during the playback of the audio bitstream, determine the updated audio bitstream according to the updated audio requirement parameters, and this updated audio bitstream is adapted to the updated audio requirement parameters.

[0050] Continuously detect the audio requirement parameters of the receiving end, and the audio requirement parameters include: playback parameters of the receiving end, network bandwidth, custom audio parameters, etc.

[0051] Optionally, the audio requirement parameters of the receiving end can be obtained according to a preset period, and compare the current audio requirement parameters with the previous audio requirement parameters. If any parameter in the two audio requirement parameters is different, it indicates that the audio requirement parameters of the receiving end have been updated.

[0052] Optionally, the audio requirement parameters of the receiving end can also be determined to be updated in the following way, specifically as follows:

[0053] The audio requirement parameters of the receiving end can be actively uploaded by the receiving end, or the server side can set an instruction to regularly obtain the audio requirement parameters of the receiving end, and respond to this instruction to regularly obtain the audio requirement parameters of the receiving end. The playback parameters of the receiving end include the playback device and its playback ability. After the playback device is replaced, the receiving end can actively upload the playback device and its playback ability. For example, when switching from a playback speaker to a Bluetooth headset, after the switch is completed, the receiving end actively detects and uploads the identification and playback ability of the Bluetooth headset. The network bandwidth of the receiving end can include: bandwidth forms corresponding to WIFI, 5G, 4G, 3G, etc. When the network bandwidth is switched, the network bandwidth can be reported regularly or in real time in response to the acquisition instruction sent by the server. The custom audio parameters of the receiving end are user-defined. When the user triggers the custom audio parameters, the receiving end can actively upload the custom audio parameters.

[0054] When any of the above audio requirement parameters is detected to change, that is, when the audio requirement parameters are updated, obtain the updated audio requirement parameters, and determine the updated audio stream from multiple groups of audio streams according to the updated audio requirement parameters.

[0055] Since forced switching of the audio stream may cause problems such as noise, therefore, in this application, splicing processing is performed based on the audio stream before the update of the audio requirement parameters and the updated audio stream after the update to obtain a switched audio stream. The switched audio stream is achieved by splicing, for example, it can be realized at an appropriate time and in a special splicing manner. The spliced switched audio stream is different from the updated audio stream. That is to say, there will be no direct switching from the audio stream to the updated audio stream, avoiding problems such as noise caused by forced switching and improving the user experience.

[0056] To more clearly understand the audio playback solution provided by this application and its technical effects, next, multiple embodiments will be used to elaborate on its specific implementation methods in detail.

[0057] In one embodiment, the method of determining the audio stream to be sent to the receiving end according to the audio requirement parameters of the receiving end provided in S210 can be implemented in the following manner, including the following sub-steps:

[0058] A1, obtain the original audio signal, and perform audio encoding on the original audio signal according to multiple sets of preset encoding parameters respectively to obtain audio streams corresponding to each set of encoding parameters;

[0059] A2, obtain the audio requirement parameters of the receiving end, and determine the audio stream to be sent from multiple groups of audio streams according to the audio requirement parameters; wherein, the audio requirement parameters include at least one of the following: playback parameters of the receiving end, network bandwidth, and custom audio parameters;

[0060] A3, send the audio stream to be sent to the receiving end, so that the receiving end receives the audio stream and decodes and plays the audio stream.

[0061] The server obtains the original audio signal, which can be the audio signal uploaded by the sending end or the signal obtained by further processing the audio signal uploaded by the sending end, such as the audio signal after decoding the audio signal uploaded by the sending end.

[0062] Obtain multiple sets of preset encoding parameters, and perform audio encoding on the original audio signal by using the multiple sets of encoding parameters to obtain audio streams corresponding to each set of encoding parameters respectively; wherein, the encoding parameters include audio sampling rate, encoding bit rate, packet size, etc.

[0063] The process of obtaining encoding parameters is as follows: A bitrate list containing multiple encoding bitrates and a sampling rate list containing multiple sampling rates are preset. Among them, an encoding bitrate is extracted from the bitrate list, and a sampling rate is extracted from the sampling rate list to form a set of encoding parameters; the encoding bitrate and encoding parameters are determined according to the support capabilities of specific audio encoders. For example, common sampling rates are: 11025Hz, 22.05kHz, 24kHz, 44.1kHz, 48kHz, and common encoding bitrates are: 6000kbps, 1100kbps, 2000kbps, 192kbps, 12kbps, etc. For example, encoding bitrate is 6000 kbps + sampling rate is 24khz, encoding bitrate is 2000kbps + sampling rate is 24 khz, etc.

[0064] Use the multiple sets of encoding parameters obtained above to perform audio encoding on the original audio signal, obtain the audio bitstreams corresponding to each set of encoding parameters respectively, and cache the generated audio bitstreams on the server.

[0065] Since the types of audio encoders are limited, the number of encoding parameters determined according to the support capabilities of the audio encoder is limited. Therefore, the space for the server to cache the audio bitstreams corresponding to each encoding parameter is limited, and it will not cause a huge waste of the server's storage resources.

[0066] Obtain the audio demand parameters of the receiving end. The audio demand parameters include: playback parameters of the receiving end, network bandwidth, custom audio parameters, etc. Among them, the playback parameters of the receiving end are such as the playback device, the sampling rate of the playback device, etc., and the custom audio parameters are such as the playback parameters of the audio signal customized by the user according to their own needs.

[0067] Determine the audio bitstream to be sent according to the audio demand parameters from multiple sets of audio bitstreams, that is, select the audio bitstream that matches the audio demand parameters of the receiving end from multiple sets of audio bitstreams, and determine it as the audio bitstream to be sent. The audio bitstream to be sent meets the audio demand parameters of the receiving end. This audio demand is also an audio playback demand. The audio demand parameters include: playback parameters of the receiving end, network bandwidth, custom audio parameters. The audio bitstream to be sent needs to meet at least one of the audio demand parameters, and can also meet all audio demand parameters. For example, the playback device is a high-fidelity audio, the network bandwidth is 10Mbit / s, but the sampling rate and bitrate in the custom audio parameters are both low, then only the audio bitstream corresponding to the encoding parameters that match the custom audio parameters in the audio demand parameters can be sent for this audio demand parameter.

[0068] Send the audio bitstream to be sent to the receiving end so that the receiving end can receive the audio bitstream and decode and play it.

[0069] For audio playback requirements with good network quality and good playback devices, an audio bitstream corresponding to the encoding parameters that is proportional to the audio requirement parameters is sent down to achieve high-fidelity and high-quality audio playback. For devices with poor playback capabilities and bad network quality that cannot perform high-quality playback, only the audio bitstream that matches the audio requirement parameters is sent down for decoding and playback.

[0070] The audio requirement parameters are proportional to the encoding parameters, that is, the higher the playback parameters, network bandwidth, and custom audio parameters in the audio requirement parameters, the higher the encoding bitstream and sampling rate in the corresponding encoding parameters. That is, under the audio requirement parameters of good network quality, good playback devices, and high custom audio parameters, the receiving end receives an audio bitstream with a high encoding bitrate and sampling rate to achieve high-fidelity and high-quality audio playback.

[0071] The audio playback method provided in this application determines the finally played audio bitstream according to the audio requirement parameters of the receiving end. The finally decoded and played audio bitstream is adapted to the audio requirement parameters of the receiving end, that is, the sent-down audio bitstream meets the audio playback requirements of the receiving end, such as: meeting the playback parameters, network bandwidth, user-defined audio parameters, etc. of the receiving end, and realizing dynamic adjustment of the sent-down audio bitstream according to the playback requirements of the receiving end, which can avoid bandwidth waste and improve the utilization rate of the receiving-end playback device and network bandwidth. Moreover, it can also avoid the response time of the audio bitstream to be sent down from being too long and meet the real-time requirements of audio playback. The method provided by this solution is applicable to the audio live broadcast scenario.

[0072] In an alternative embodiment, A2 determines the audio bitstream to be sent down from multiple groups of audio bitstreams according to the audio requirement parameters, which can be implemented in the following ways, including:

[0073] B1, using the preset mapping table of audio requirement parameters and encoding parameters and the audio requirement parameters to determine the encoding parameters corresponding to the audio requirement parameters;

[0074] B2, determining the audio bitstream corresponding to the encoding parameters as the audio bitstream to be sent down.

[0075] Pre-build a mapping table of audio requirement parameters and encoding parameters, that is, each audio requirement parameter in the mapping table corresponds to a set of encoding parameters. For example: for good network quality and high-fidelity playback devices, the corresponding encoding parameters can be set as: 48khz sampling rate + 192kbps bitrate. In this way, an association relationship between the audio requirement parameters and the encoding parameters is constructed to form a mapping table of audio requirement parameters and encoding parameters.

[0076] Find the associated encoding parameter in the mapping table according to the audio requirement parameter, then encode the audio requirement parameter based on the obtained encoding parameter to obtain the corresponding audio bitstream, and use this audio bitstream as the audio bitstream to be sent. This method can accurately and quickly determine the audio bitstream to be sent.

[0077] Since the audio requirement parameters at the receiving end are diverse, and the audio requirement parameters stored in the mapping table are limited, there are some audio requirement parameters that are not stored in the mapping table. For these audio requirement parameters, the embodiments of the present application provide a solution as follows:

[0078] Based on the previous embodiment, in an optional embodiment, when the encoding parameter corresponding to the audio requirement parameter of the receiving end is not stored in the mapping table, A2 determines the audio bitstream to be sent from multiple groups of audio bitstreams according to the audio requirement parameter, which can be implemented in the following ways, including:

[0079] C1, obtain the audio requirement parameter in the mapping table that is closest to the audio requirement parameter from the receiving end, and use the closest audio requirement parameter in the mapping table as the approximate audio requirement parameter of the receiving end;

[0080] C2, use the encoding parameter corresponding to this approximate audio requirement parameter as the encoding parameter corresponding to the audio requirement parameter of the receiving end, and determine the audio bitstream corresponding to this encoding parameter as the audio bitstream to be sent.

[0081] Find the audio requirement parameter in the mapping table that is closest to the audio requirement parameter of the receiving end, use the closest audio requirement parameter as the approximate audio requirement parameter of the receiving end, find the encoding parameter associated with this approximate audio requirement parameter in the mapping table, and use the audio bitstream obtained by audio encoding with this encoding parameter as the audio bitstream to be sent, which can quickly and accurately determine the audio bitstream to be sent.

[0082] The solution provided by this embodiment can solve the problem of determining the audio bitstream to be sent when the audio requirement parameter of the receiving end is not stored in the mapping table. Moreover, in this solution, there is no need to store a large number of parameter pairs in the mapping table, and only a limited number of parameter pairs are stored to reduce the memory space occupied by the mapping table.

[0083] Optionally, the mapping table can be sorted according to the level of demand represented by the audio requirement parameter. The present application can also use the audio requirement parameter in the mapping table whose demand is lower than the audio requirement parameter of the receiving end as the closest audio requirement parameter to meet the smooth playback requirement of the playback device at the receiving end.

[0084] If the audio requirement parameters change during playback, the audio bitstream corresponding to the audio requirement parameters will also change, and the corresponding encoding parameters will be different. That is, in this case, a switch of encoding parameters will occur. In response to this situation, the present application performs bitstream splicing processing based on the currently playing audio bitstream and the updated audio bitstream to obtain a switched audio bitstream. The present application provides an optional implementation manner. One implementation manner for obtaining the switched audio bitstream is shown in the process Figure 3 , including:

[0085] S310, obtaining the audio frame that is being played when the audio requirement parameters are updated;

[0086] S320, detecting the audio frame sample values of the audio frame and each subsequent audio frame in the audio bitstream;

[0087] S330, obtaining the first audio frame corresponding to when the audio frame sample value is first detected to meet the switching condition, obtaining the audio bitstream segment from the audio frame being played in the audio bitstream to this first audio frame, and the updated audio bitstream segment after the first audio frame in the updated audio bitstream; wherein, the switching condition includes: the energy value represented by the audio frame sample value is lower than a preset threshold;

[0088] S340, performing splicing processing on the audio bitstream segment and the updated audio bitstream segment to obtain a switched audio bitstream.

[0089] After obtaining the updated audio bitstream, switching can only be performed at an appropriate time and under specific rule constraints to avoid abnormalities in the decoded audio bitstream. The spectrograms of the audio bitstream switching for different processing operations are shown in Figure 4. Figure 4-1 is the spectrogram of switching the audio bitstreams corresponding to different encoding parameters without constraint conditions. Figure 4-2 is the spectrogram of switching the audio bitstreams corresponding to different encoding parameters with constraints. By comparison, it can be seen that Figure 4-1 obviously vertical stripes appear, and obvious noise will appear at the positions where these vertical stripes appear, while Figure 4-2 no obvious vertical stripes appear, and the spectrogram is natural and smooth. Therefore, switching the audio bitstreams corresponding to different encoding parameters under constraint conditions can reduce the generation of noise, thereby improving the user experience.

[0090] In an optional embodiment, switching the audio bitstream under constraint conditions can be implemented through the steps provided by S320 to S340, specifically as follows:

[0091] Detect the audio frame sample values for each audio frame in the current audio bitstream, where the audio frame sample value is the amplitude of a preset sampling point on the audio frame; when it is detected that the audio frame sample value meets the switching condition, it indicates that the switching timing is appropriate. Obtain the first audio frame corresponding to when the audio frame sample value is first detected to meet the switching condition, and perform the switching at the first audio frame. Before the first audio frame, play using the audio bitstream, and after the first audio frame, play using the updated audio bitstream. Therefore, the switching audio bitstream to be played is composed of two parts. Specifically, it is composed of the audio bitstream segment from the audio frame being played to the first audio frame in the audio bitstream, and the updated audio bitstream segment after the first audio frame in the updated audio bitstream.

[0092] Before S330, it further includes: S321, determining whether the audio frame sample value meets the switching condition. When it meets the switching condition, execute S330; when it does not meet the switching condition, execute S331. When the audio frame sample value does not meet the switching condition, perform audio encoding according to the encoding parameters before the update. When playing this audio frame, use the audio bitstream information corresponding to this audio encoding. For the next audio frame, loop and execute S320 - S321 until the audio sample value meets the switching condition, and use the updated audio bitstream as the audio bitstream to be sent for the audio segment after the first audio frame corresponding to the audio sample value that meets the switching condition.

[0093] For the audio playback method provided in this embodiment, when it is detected that the audio requirement parameter changes, detect the audio frame sample values in the audio bitstream before the update. If the audio frame sample value meets the switching condition, that is, the constraint condition, then use the updated audio bitstream as the audio bitstream to be sent, realizing the switching of the audio bitstream, and realizing the dynamic switching of the audio bitstream to be sent according to the adjustment of the audio requirement parameter at the receiving end during the audio playback process.

[0094] For the solution provided in this embodiment, judge the switching timing according to the result of whether the audio frame sample value meets the switching condition. If the audio frame sample value does not meet the switching condition, no switching is performed. Only when the audio frame sample value meets the switching condition is the audio bitstream switched. That is to say, the switching audio bitstream after the audio frame being played includes: the audio bitstream segment before the first audio frame and the updated audio bitstream segment after the first audio frame in the updated audio bitstream. Performing the switching of the audio stream at this switching timing is conducive to realizing the smooth switching of the audio bitstream.

[0095] Optionally, the above switching condition is: the energy value represented by the audio frame sample value is lower than a preset threshold. Among them, the preset threshold can be customized according to the actual situation. An audio frame with an energy value lower than the preset threshold is a low-energy frame or a silent frame. In this application, switching the audio bitstream at a low-energy frame or a silent frame can avoid obvious noise during the switching of the audio bitstream, realize the smooth switching of the audio bitstream, and will not affect the user experience.

[0096] For each audio frame in the audio bitstream before update, perform audio frame sample value detection frame by frame. Assuming the audio sampling rate = 8000, the sampling channels = 2, the bit depth = 8, and the sampling interval = 20 ms, then the size of each audio frame data is 320 bits, the number of samples per channel is 160 bits. Set 320 sample points on each audio frame, obtain the amplitude of each sample point, and calculate the average of the sum of the squares of the amplitudes of all sample points included in the audio frame to obtain the energy value of the audio frame.

[0097] Optionally, when the amplitudes of all sample points on the audio frame do not exceed 50, the audio frame is a silent frame; when the energy value of the audio frame is less than 10000, the audio frame is a low-energy frame. The judgment logic for the audio bitstream to switch when the current audio frame is a silent frame or a low-energy frame is as follows:

[0098] Lowenerflag = 1;

[0099] for i = 0:Framesize-1

[0100] If (abs(x[i]>thrd))

[0101] Lowenerflag = 0;

[0102] end

[0103] End

[0104] Among them, thrd is the sampling value judgment threshold, and Lowenerflag is the low-energy frame flag. When Lowenerflag is 1, the switch can be performed.

[0105] The audio playback method provided by this application can perform the switching of the audio bitstream only when the current audio frame is a silent frame or a low-energy frame, which can avoid obvious noise during the switching of the audio bitstream and achieve smooth switching of the audio bitstream.

[0106] The solution provided in the above embodiment is a method for switching the old and new configured audio bitstreams on the server based on the audio frame sample values meeting the switching conditions. The finally determined switched audio bitstream will be decoded and played on the client. This method can achieve functional compatibility for different versions of the client and smoothly play audio information. In addition, this application also provides another embodiment to obtain the switched audio bitstream in real time, including:

[0107] D1. Obtain the audio frame being played when the audio requirement parameter is updated and the corresponding new audio frame in the updated audio bitstream;

[0108] D2, perform a fade-in and fade-out splicing process on the audio frame and the new audio frame to obtain a switched audio frame;

[0109] D3, use the switched audio frame to replace the new audio frame in the updated audio stream, and determine the updated audio stream containing the switched audio frame as the switched audio stream.

[0110] Among them, performing a fade-in and fade-out splicing process on the audio frame and the new audio frame to obtain a switched audio frame can be carried out through the following formula:

[0111]

[0112] Among them, d(n) is the switched audio frame signal obtained after its fade-in and fade-out processing, d1(n) and d2(n) are respectively the pulse code modulation signals decoded from the audio stream before the update of the audio demand parameters and the pulse code modulation signals decoded from the updated audio stream after the update of the audio demand parameters, n is the audio sample number, the value range of n is 1 to N, win is the half window of the Hanning window. For example, the size of the Hanning window is 2*N, and the win used here is the window signal of the first N audio samples.

[0113] In the solution provided in this embodiment, when the server receives a change in the audio parameter sending, it does not need to wait for the switching condition to be met, but instead sends the old code stream frame and the new code stream frame of the same audio frame corresponding to the switching moment to the client at the same time. Here, the old code stream frame is the audio frame at the switching moment in the audio stream, and the new code stream frame is the new audio frame at the switching moment in the updated audio stream. After receiving the two old and new code stream frames of the same audio frame, the client decodes them separately, and after decoding, two frames of audio data at the same moment are obtained. The audio contents of the old and new code stream frames are similar, only with some differences in quality. After performing a fade-in and fade-out splicing process on these two frames of signals, an integrated switched audio frame is obtained. This integrated switched audio frame is the first frame of audio after switching. After that, the server only needs to send the updated audio stream segment after this frame to the client, and the client only needs to decode and play the received updated audio stream segment.

[0114] The solution provided in this embodiment can immediately switch the audio stream when it detects an update in the audio demand parameters. Moreover, since two audio frames are processed for the currently playing audio frame, there will be no problems such as noise during the switching process. While smoothly switching the audio stream, it is beneficial to improve the real-time performance of the audio stream switching.

[0115] In an alternative embodiment, detecting an update in the audio demand parameters during the playback of the audio stream provided by S320 can be achieved through the following methods, including:

[0116] If any one of the trigger operation of the custom audio parameters from the receiving end, the update message of the network bandwidth, and the update operation of the playback parameters is detected, it is determined that the audio requirement parameters of the receiving end have been updated.

[0117] Specifically, the playback parameters of the receiving end include the playback device and its playback capabilities. After the playback device is replaced, the receiving end can actively upload the playback device and its playback capabilities. For example, when switching from a playback speaker to a Bluetooth headset, after the switch is completed, the receiving end actively uploads the identification and playback capabilities of the Bluetooth headset, where the playback capabilities are such as the sampling rate. The network bandwidth of the receiving end can include: bandwidth forms corresponding to WIFI, 5G, 4G, 3G, etc. When the network bandwidth is switched, it can report the network bandwidth regularly or in real time in response to the acquisition instruction issued by the server. The custom audio parameters of the receiving end are user-defined. When the user triggers the custom audio parameters, the receiving end actively uploads the adjusted custom audio parameters.

[0118] When it is detected that any one of the above audio requirement parameters has changed, it is determined that the audio requirement parameters are updated. This method can accurately judge whether the audio requirement parameters of the receiving end have changed, and the execution process is simple, avoiding occupying too many resources for this judgment.

[0119] In an alternative embodiment, after obtaining the audio bitstreams corresponding to each group of encoding parameters disclosed in A1, the following steps are further included. The flowchart of the audio playback method including the following steps is as Figure 5 shown.

[0120] S510, segment the audio bitstream according to a preset time period to obtain multiple audio segments, and sort the audio segments in chronological order;

[0121] S520, form an audio bitstream matrix for the multiple audio segments sorted in time according to the way that the audio segments corresponding to the same time period are placed in the same row and the audio segments corresponding to different encoding parameters are placed in different columns.

[0122] Multiple groups of encoding parameters correspond to multiple audio bitstreams. Each audio bitstream can be divided into multiple audio segments according to a preset time period, and then these divided audio segments are sorted in chronological order. The audio segments corresponding to different encoding parameters in the same time period are placed in the same row, and the audio segments corresponding to the same encoding parameter in different time periods are placed in the same column. In this way, multiple audio bitstreams corresponding to multiple groups of encoding parameters form an audio bitstream matrix, and each audio segment in the matrix corresponds to different times and encoding parameters. As Figure 6As shown, the audio bitstream matrix includes audio bitstreams corresponding to encoding parameter 1, encoding parameter 2, encoding parameter 3, …, encoding parameter n. The audio bitstream corresponding to encoding parameter 1 includes the following audio segments: Stream_1(t), Stream_1(t + 1), Stream_1(t + 2), … The audio bitstream corresponding to encoding parameter n includes the following audio segments: Stream_n(t), Stream_n(t + 1), Stream_n(t + 2), …. The audio segments corresponding to each set of encoding parameters are arranged in a column in chronological order. The encoding capabilities corresponding to each set of encoding parameters are different. The encoding capabilities corresponding to encoding parameters 1 to n can be from large to small or from small to large, which is not restricted here.

[0123] For example: What encoding parameter 1 corresponds to is a combination of a sampling rate of 48 kHz + a bitrate of 192 kbps. The audio bitstream corresponding to this combination is Figure 6 the queue of encoding parameter 1 in []. However, for users with low requirements for audio quality, their network quality is poor and their playback devices are inferior, unable to meet the needs of high-quality playback. If smooth playback is required, their audio demand parameters correspond to a lower encoding parameter configuration, such as a combination of a sampling rate of 8 kHz + a bitrate of 12 kbps. Since the bitstream packets in this combination are smaller, it can adapt to scenarios with a weaker network. Combining certain packet loss resistance technologies can better meet the user's need for smooth playback. If encoding parameters 1 to n are sorted in descending order of encoding ability, then in this case, the bitstream queue of encoding parameter n can be selected.

[0124] Furthermore, if the preset time period of the above segmentation is the sampling interval of the audio bitstream, such as 20 ms, etc., then the audio segment corresponding to one preset time period is an audio frame.

[0125] In this case, that is, when the preset time period is equal to the sampling interval of the audio bitstream, when it is detected that the audio demand parameters at the receiving end are updated, the updated audio bitstream can be determined according to the updated audio demand parameters, and it can be obtained through the following methods, including:

[0126] S3101, obtain the update time of the audio demand parameters, and determine the corresponding updated encoding parameters according to the updated audio demand parameters;

[0127] S3102, locate the position of the updated encoding parameters in the audio bitstream matrix, and determine the audio segments corresponding to the time period after the update time in the column where the updated encoding parameters are located as the updated audio bitstream.

[0128] Specifically, for a receiving end with dynamically changing audio demand parameters, the required bitstream may come from audio bitstream queues corresponding to different encoding parameters. For example, it may switch from the audio bitstream queue of encoding parameter 1 with the highest audio quality to the audio bitstream queue of encoding parameter 3 with slightly lower audio quality. If it is detected that the update time of the audio demand parameter is (t + 1), then in the audio bitstream queue corresponding to encoding parameter 3, the audio segment corresponding to the time period after (t + 1) is the updated audio bitstream.

[0129] The audio playback method provided by the embodiments of the present application divides the audio bitstream into multiple audio segments and forms an audio bitstream matrix in sequence, which can intuitively determine the position of the current audio frame or audio segment in the audio bitstream matrix. When using the audio bitstream matrix to determine the updated audio bitstream, only the update time of the audio demand parameter needs to be obtained, and then the updated encoding parameter can be accurately obtained by combining this update time, the updated encoding parameter, and the audio bitstream matrix.

[0130] In one embodiment of the present application, the scheme for switching the audio bitstream can be implemented through a Figure 7 flowchart as shown below. The specific process is as follows: The server side performs the following steps: Obtain the audio demand parameters of the receiving client, including: user-defined audio quality configuration information, network bandwidth and quality detection information, and playback device capability detection information; Based on the audio demand parameters of the receiving client, combine with a preset audio quality configuration mapping table to determine the configuration parameters, and here the configuration parameters are also called encoding parameters; Then, perform frame sample value detection on the unplayed audio frames in the audio bitstream, and thus judge whether the frame sample values of each audio frame meet the switching conditions; If so, configure the new encoding parameter; If not, configure the existing encoding parameter, that is, the audio demand parameter before the update, obtain the next audio frame in the audio bitstream corresponding to the audio demand parameter before the update, judge whether the frame sample value of the next audio frame meets the switching conditions, if not, continue to judge in a loop until the frame sample value of the next audio frame meets the switching conditions, configure the new encoding parameter, obtain the updated audio bitstream corresponding to the new encoding parameter, and forward it to the receiving client through the distributed CDN server, realizing the audio bitstream switching when the frame sample value of the audio frame meets the switching conditions, and avoiding phenomena such as noise caused by forced audio bitstream switching.

[0131] In an alternative embodiment, the acquisition of the original audio signal disclosed in step A1 can be achieved in the following ways, including:

[0132] E1, Receive the source audio signal collected and uploaded by the sending end at a preset sampling rate; where the preset sampling rate is greater than a preset threshold;

[0133] E2, Decode the source audio signal to obtain the original audio signal.

[0134] The sending end collects the source audio signal, which can be collected using a preset sampling rate. The preset sampling rate can be a sampling rate greater than a preset threshold. The preset threshold can be obtained based on the sampling rate of high-quality audio signals divided by the industry, or it can be the maximum sampling rate of the sending end, capable of uploading audio signals with the highest fidelity. Among them, the sampling rate can be 48 kHz, 24 kHz, 16 kHz, 8 kHz, etc. If the preset threshold is 16 kHz, then the preset sampling rate can be 48 kHz, 24 kHz, to ensure that the source audio signal collected by this application using a preset sampling rate greater than the preset threshold is a high-quality audio signal. So that after the server processes based on the high-quality audio signal, the receiving end can play a high-quality audio stream when the audio demand parameters allow, improving the user experience.

[0135] Based on the same principle as the method provided in the embodiments of this application, the embodiments of this application also provide an audio playback device 800, as Figure 8 shown. The device may include: an audio stream module 810, an updated audio stream module 820, a switched audio stream module 830, and a played audio stream module 840, where:

[0136] The audio stream module 810 is used to determine the audio stream sent to the receiving end according to the audio demand parameters of the receiving end;

[0137] The updated audio stream module 820 is used to determine an updated audio stream according to the updated audio demand parameters when it is detected that the audio demand parameters are updated during the playback of the audio stream;

[0138] The switched audio stream module 830 is used to perform stream splicing processing based on the audio stream and the updated audio stream to obtain a switched audio stream;

[0139] The played audio stream module 840 is used to send the switched audio stream to the receiving end, so that the receiving end receives the switched audio stream and decodes and plays the switched audio stream.

[0140] The audio playback device provided by this application performs splicing processing based on the audio stream before the update of the audio demand parameters and the updated audio stream after the update to obtain a switched audio stream. The switched audio stream is achieved by splicing, such as at an appropriate time and in a special splicing method. The spliced switched audio stream is different from the updated audio stream, avoiding directly switching from the audio stream to the updated audio stream. Therefore, problems such as noise caused by forced stream switching will not occur, improving the user experience.

[0141] Optionally, the switched audio stream module 830 includes:

[0142] An audio frame obtaining unit, configured to obtain an audio frame that is being played when the audio requirement parameter is updated;

[0143] A detection unit, configured to perform audio frame sample value detection on the audio frame and each subsequent audio frame in the audio bitstream;

[0144] A bitstream segment obtaining unit, configured to obtain a first audio frame corresponding to when the audio frame sample value is first detected to meet the switching condition, obtain an audio bitstream segment from the audio frame being played to the first audio frame in the audio bitstream, and an updated audio bitstream segment after the first audio frame in the updated audio bitstream; wherein, the switching condition includes: the energy value represented by the audio frame sample value is lower than a preset threshold;

[0145] A first splicing unit, configured to perform splicing processing on the audio bitstream segment and the updated audio bitstream segment to obtain a switched audio bitstream.

[0146] Optionally, the switched audio bitstream module 830 further includes:

[0147] A new audio frame obtaining unit, configured to obtain an audio frame that is being played when the audio requirement parameter is updated and a new audio frame corresponding to the audio frame in the updated audio bitstream;

[0148] A second splicing unit, configured to perform signal fade-in and fade-out splicing processing on the audio frame and the new audio frame to obtain a switched audio frame;

[0149] A switched audio bitstream determining unit, configured to replace the new audio frame in the updated audio bitstream with the switched audio frame, and determine the updated audio bitstream including the switched audio frame as the switched audio bitstream.

[0150] Optionally, the audio bitstream module 810 includes:

[0151] An audio encoding unit, configured to obtain an original audio signal, perform audio encoding on the original audio signal according to a preset multiple groups of encoding parameters respectively, and obtain audio bitstreams corresponding to each group of encoding parameters respectively;

[0152] A to-be-transmitted bitstream determining unit, configured to obtain audio requirement parameters of a receiving end, and determine a to-be-transmitted audio bitstream from multiple groups of audio bitstreams according to the audio requirement parameters; the audio requirement parameters include at least one of: playback parameters of the receiving end, network bandwidth, and custom audio parameters;

[0153] A decoding and playing unit, configured to send the to-be-transmitted audio bitstream to the receiving end, so that the receiving end receives the audio bitstream and decodes and plays the audio bitstream.

[0154] The audio playback device provided by this application determines the finally played audio bitstream according to the audio requirement parameters of the receiving end. The finally decoded and played audio bitstream is adapted to the audio requirement parameters of the receiving end, that is, the sent audio bitstream meets the audio playback requirements of the receiving end, such as: the playback parameters of the receiving end, network bandwidth, user-defined audio parameters, etc., realizing dynamically adjusting the sent audio bitstream according to the playback requirements of the receiving end, which can avoid bandwidth waste and avoid too long response time of the audio bitstream to be sent.

[0155] Optionally, determining the bitstream unit to be sent is specifically used for:

[0156] Using the mapping table of preset audio requirement parameters and encoding parameters, and the audio requirement parameters, to determine the encoding parameters corresponding to the audio requirement parameters;

[0157] Determining the audio bitstream corresponding to the encoding parameters as the audio bitstream to be sent.

[0158] Optionally, when the encoding parameters corresponding to the audio requirement parameters are not stored in the mapping table, determining the bitstream unit to be sent is specifically used for:

[0159] Obtaining the audio requirement parameter in the mapping table that is closest to the audio requirement parameter from the receiving end, and using the closest audio requirement parameter in the mapping table as the approximate audio requirement parameter of the receiving end;

[0160] Using the encoding parameters corresponding to the approximate audio requirement parameter as the encoding parameters corresponding to the audio requirement parameter of the receiving end, and determining the audio bitstream corresponding to the encoding parameters as the audio bitstream to be sent.

[0161] Optionally, updating the audio bitstream module 820 is specifically used for:

[0162] Detecting any one of the triggering operation of the user-defined audio parameter from the receiving end, the update message of the network bandwidth, and the update operation of the playback parameter, then determining that the audio requirement parameter of the receiving end has been updated.

[0163] Optionally, the audio playback device 800 further includes:

[0164] A segment sorting module, used to segment the audio bitstream according to a preset time period to obtain multiple audio segments, and sort each audio segment in chronological order;

[0165] An audio bitstream matrix module, used to form an audio bitstream matrix for the multiple audio segments sorted by time in such a way that the audio segments corresponding to the same time period are placed in the same row, and the audio segments corresponding to different encoding parameters are placed in different columns;

[0166] In this embodiment, when the preset time period is equal to the sampling interval of the audio code stream, the audio code stream module is updated, specifically for:

[0167] Obtaining an update time of the audio requirement parameter, and determining a corresponding update encoding parameter according to the updated audio requirement parameter;

[0168] The updated coding parameter is located in the audio bitstream matrix, and an audio segment corresponding to a time period after the update time in the column where the updated coding parameter is located is determined as an updated audio bitstream.

[0169] Optionally, the audio encoding unit further includes:

[0170] A source audio signal receiving subunit is used to receive a source audio signal collected and uploaded by a transmitting end at a preset sampling rate; wherein the preset sampling rate is greater than a preset threshold;

[0171] The original audio signal obtaining subunit is used to decode the source audio signal to obtain the original audio signal.

[0172] The audio playback device of the embodiment of the present application can execute the audio playback method provided by the embodiment of the present application, and the implementation principle is similar. The actions performed by each module and unit in the audio playback device in each embodiment of the present application correspond to the steps in the audio playback method in each embodiment of the present application. For the detailed functional description of each module of the audio playback device, please refer to the description in the corresponding audio playback method shown in the previous text, which will not be repeated here.

[0173] Based on the same principle as the method shown in the embodiment of the present application, an electronic device is also provided in the embodiment of the present application, which may include but is not limited to: a processor and a memory; a memory for storing a computer program; a processor for executing the audio playback method shown in any optional embodiment of the present application by calling a computer program. Compared with the prior art, the present application avoids directly switching from an audio code stream to an updated audio code stream, thereby avoiding problems such as noise caused by forced switching of code streams. The audio code stream sent meets the audio playback requirements of the receiving end, and dynamically adjusts the audio code stream sent according to the playback requirements of the receiving end, which can avoid bandwidth waste, avoid excessively long response time of the audio code stream to be sent, and improve the utilization rate of the playback device and bandwidth.

[0174] In an alternative embodiment, an electronic device is provided, such as Figure 9 As shown, Figure 9The electronic device 4000 shown may be a server, including: a processor 4001 and a memory 4003. Among them, the processor 4001 and the memory 4003 are connected, such as through a bus 4002. Optionally, the electronic device 4000 may further include a transceiver 4004. It should be noted that in actual applications, the transceiver 4004 is not limited to one, and the structure of the electronic device 4000 does not constitute a limitation to the embodiments of the present application.

[0175] The processor 4001 may be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute various exemplary logical blocks, modules, and circuits described in connection with the disclosure of the present application. The processor 4001 may also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, etc.

[0176] The bus 4002 may include a path for transmitting information between the above components. The bus 4002 may be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, etc. The bus 4002 may be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, Figure 9 only a thick line is used to represent it in the figure, but it does not mean that there is only one bus or one type of bus.

[0177] The memory 4003 can be a ROM (Read Only Memory), or other types of static storage devices that can store static information and instructions, a RAM (Random Access Memory), or other types of dynamic storage devices that can store information and instructions. It can also be an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read Only Memory), or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media, or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto.

[0178] The memory 4003 is used to store the application program code for executing the solution of this application, and is controlled by the processor 4001 for execution. The processor 4001 is used to execute the application program code stored in the memory 4003 to implement the content shown in the foregoing method embodiments.

[0179] Among them, the electronic device includes but is not limited to: mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), PMPs (Portable Multimedia Players), in-vehicle terminals (such as in-vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 9 The electronic device shown is only an example and should not impose any limitations on the functions and usage scope of the embodiments of this application.

[0180] The embodiments of this application provide a computer-readable storage medium, on which a computer program is stored. When it runs on a computer, it enables the computer to execute the corresponding content in the foregoing method embodiments.

[0181] It should be understood that although the steps in the flowchart of the accompanying drawings are shown in sequence according to the indication of the arrows, these steps do not necessarily have to be executed in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps has no strict order limit, and they can be executed in other orders. Moreover, at least a part of the steps in the flowchart of the accompanying drawings may include multiple sub-steps or multiple stages. These sub-steps or stages do not necessarily have to be executed at the same time, but can be executed at different times, and their execution order does not necessarily have to be sequential, but can be executed alternately or in turn with at least a part of other steps or sub-steps or stages of other steps.

[0182] It should be noted that the above-mentioned computer-readable medium in this application can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this application, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device. And in this application, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and this computer-readable signal medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted by any suitable medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.

[0183] The above-mentioned computer-readable medium can be included in the above-mentioned electronic device; it can also exist separately and not be assembled into the electronic device.

[0184] The above-mentioned computer-readable medium carries one or more programs, and when the above-mentioned one or more programs are executed by the electronic device, the electronic device is caused to execute the method shown in the above-mentioned embodiments.

[0185] Computer program code for performing the operations of this application can be written in one or more programming languages or combinations thereof. The above-mentioned programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any kind of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).

[0186] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, as well as the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0187] The modules described in the embodiments of this application can be implemented in software or in hardware. Among them, the name of the module does not constitute a limitation to the module itself in some cases. For example, the audio code stream module can also be described as the "audio code stream module for determining the one sent to the receiving end".

[0188] The above description is only for the preferred embodiments of this application and the explanation of the applied technical principles. Those skilled in the art should understand that the scope of disclosure involved in this application is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosure concept. For example, the technical solutions formed by mutually replacing the above features with the (but not limited to) technical features with similar functions disclosed in this application.

Claims

1. An audio playback method, characterized in that, Including: Determine the audio bitstream to be sent to the receiving end according to the audio requirement parameters of the receiving end; When it is detected that the audio requirement parameters are updated during the playback of the audio bitstream, determine the updated audio bitstream according to the updated audio requirement parameters; Perform bitstream splicing processing based on the audio bitstream and the updated audio bitstream to obtain a switched audio bitstream, including: obtaining the audio frame being played when the audio requirement parameters are updated; detecting the audio frame sample values of each audio frame after the audio frame and this frame in the audio bitstream; obtaining the first audio frame corresponding to when the audio frame sample value is first detected to meet the switching condition, obtaining the audio bitstream segment from the audio frame being played to the first audio frame in the audio bitstream, and the updated audio bitstream segment after the first audio frame in the updated audio bitstream; wherein, the switching condition includes: the energy value represented by the audio frame sample value is lower than a preset threshold; perform splicing processing on the audio bitstream segment and the updated audio bitstream segment to obtain a switched audio bitstream; Send the switched audio bitstream to the receiving end so that the receiving end receives the switched audio bitstream and decodes and plays the switched audio bitstream.

2. The method according to claim 1, wherein The performing bitstream splicing processing based on the audio bitstream and the updated audio bitstream to obtain a switched audio bitstream includes: Obtain the audio frame being played when the audio requirement parameters are updated and the corresponding new audio frame of this audio frame in the updated audio bitstream; Perform signal fade-in and fade-out splicing processing on the audio frame and the new audio frame to obtain a switched audio frame; Use the switched audio frame to replace the new audio frame in the updated audio bitstream, and determine the updated audio bitstream containing the switched audio frame as the switched audio bitstream.

3. The method according to claim 1, wherein The determining the audio bitstream to be sent to the receiving end according to the audio requirement parameters of the receiving end includes: Obtain the original audio signal, perform audio encoding on the original audio signal according to a preset set of encoding parameters respectively to obtain audio bitstreams corresponding to each set of encoding parameters; Obtain the audio requirement parameters of the receiving end, and determine the audio bitstream to be sent from the multiple sets of audio bitstreams according to the audio requirement parameters; the audio requirement parameters include at least one of the following: the playback parameters of the receiving end, the network bandwidth, and the custom audio parameters; Send the audio bitstream to be sent to the receiving end so that the receiving end receives the audio bitstream and decodes and plays the audio bitstream.

4. The method according to claim 3, wherein The determining the audio bitstream to be sent from the multiple sets of audio bitstreams according to the audio requirement parameters includes: Use a preset mapping table of audio requirement parameters and encoding parameters and the audio requirement parameters to determine the encoding parameters corresponding to the audio requirement parameters; Determine the audio bitstream corresponding to the encoding parameters as the audio bitstream to be sent.

5. The method according to claim 4, characterized in that, When the mapping table does not store the encoding parameters corresponding to the audio requirement parameters, the determining the audio bitstream to be sent from the multiple sets of audio bitstreams according to the audio requirement parameters includes: Obtain the audio requirement parameter in the mapping table that is closest to the audio requirement parameter from the receiving end, and use the closest audio requirement parameter in the mapping table as the approximate audio requirement parameter of the receiving end; Use the encoding parameter corresponding to the approximate audio requirement parameter as the encoding parameter corresponding to the audio requirement parameter of the receiving end, and determine the audio bitstream corresponding to this encoding parameter as the audio bitstream to be sent down.

6. The method according to claim 1, wherein The detecting that the audio requirement parameter is updated during the playing of the audio bitstream includes: If any one of a triggering operation of a custom audio parameter from the receiving end, an update message of the network bandwidth, and an update operation of the playing parameter is detected, it is determined that the audio requirement parameter of the receiving end is updated.

7. The method according to claim 3, characterized in that, After obtaining the audio bitstreams corresponding to each group of encoding parameters, it further includes: Segment the audio bitstream according to a preset time period to obtain a plurality of audio segments, and sort the audio segments in chronological order; Form an audio bitstream matrix for the plurality of audio segments sorted in time in such a way that the audio segments corresponding to the same time period are placed in the same row, and the audio segments corresponding to different encoding parameters are placed in different columns; When the preset time period is equal to the sampling interval of the audio bitstream, when it is detected that the audio requirement parameter of the receiving end is updated, determining the updated audio bitstream according to the updated audio requirement parameter includes: Obtain the update time of the audio requirement parameter, and determine the corresponding updated encoding parameter according to the updated audio requirement parameter; Locate the position of the updated encoding parameter in the audio bitstream matrix, and determine the audio segments corresponding to the time period after the update time in the column where the updated encoding parameter is located as the updated audio bitstream.

8. The method according to claim 3, wherein The obtaining the original audio signal includes: Receive the source audio signal collected and uploaded by the sending end at a preset sampling rate; wherein, the preset sampling rate is greater than a preset threshold; Decode the source audio signal to obtain the original audio signal.

9. An audio playback device, characterized in that, It includes: An audio bitstream module, configured to determine the audio bitstream sent to the receiving end according to the audio requirement parameter of the receiving end; An updated audio bitstream module, configured to determine the updated audio bitstream according to the updated audio requirement parameter when it is detected that the audio requirement parameter is updated during the playing of the audio bitstream; A switched audio bitstream module, configured to perform bitstream splicing processing based on the audio bitstream and the updated audio bitstream to obtain a switched audio bitstream; A playing audio bitstream module, configured to send the switched audio bitstream to the receiving end, so that the receiving end receives the switched audio bitstream and decodes and plays the switched audio bitstream; Wherein, the switched audio bitstream module includes: An obtaining audio frame unit, configured to obtain the audio frame being played when the audio requirement parameter is updated; A detecting unit, configured to detect the audio frame sample point values of the audio frame and each audio frame after this frame in the audio bitstream; An obtaining bitstream segment unit, configured to obtain the first audio frame corresponding to when the audio frame sample point value is first detected to meet the switching condition, obtain the audio bitstream segment from the audio frame being played in the audio bitstream to the first audio frame, and the updated audio bitstream segment after the first audio frame in the updated audio bitstream; wherein, the switching condition includes: the energy value represented by the audio frame sample point value is lower than a preset threshold. The first splicing unit is used to splice the audio code stream segment and the updated audio code stream segment to obtain a switched audio code stream.

10. The device according to claim 9, characterized in that, The switched audio code stream module further includes: The new audio frame obtaining unit is used to obtain the audio frame being played when the audio requirement parameter is updated and the corresponding new audio frame of the audio frame in the updated audio code stream; The second splicing unit is used to perform signal fade-in and fade-out splicing processing on the audio frame and the new audio frame to obtain a switched audio frame; The switched audio code stream determining unit is used to replace the new audio frame in the updated audio code stream with the switched audio frame, and determine the updated audio code stream including the switched audio frame as the switched audio code stream.

11. The device according to claim 9, characterized in that, The audio code stream module includes: The audio encoding unit is used to obtain the original audio signal, and perform audio encoding on the original audio signal according to a preset multiple groups of encoding parameters respectively to obtain the audio code streams corresponding to the respective groups of encoding parameters; The to-be-transmitted code stream determining unit is used to obtain the audio requirement parameters of the receiving end, and determine the to-be-transmitted audio code stream from the multiple groups of audio code streams according to the audio requirement parameters; the audio requirement parameters include at least one of the playback parameters of the receiving end, the network bandwidth, and the custom audio parameters; The decoding and playback unit is used to send the to-be-transmitted audio code stream to the receiving end, so that the receiving end receives the audio code stream and decodes and plays the audio code stream.

12. The device according to claim 11, characterized in that, The to-be-transmitted code stream determining unit is specifically used for: Using a preset mapping table of audio requirement parameters and encoding parameters and the audio requirement parameters, determine the encoding parameters corresponding to the audio requirement parameters; Determine the audio code stream corresponding to the encoding parameters as the to-be-transmitted audio code stream.

13. The device according to claim 12, characterized in that, When the mapping table does not store the encoding parameters corresponding to the audio requirement parameters, the to-be-transmitted code stream determining unit is specifically used for: Obtain the audio requirement parameter in the mapping table that is closest to the audio requirement parameter from the receiving end, and use the audio requirement parameter closest to the audio requirement parameter in the mapping table as the approximate audio requirement parameter of the receiving end; Use the encoding parameters corresponding to the approximate audio requirement parameters as the encoding parameters corresponding to the audio requirement parameters of the receiving end, and determine the audio code stream corresponding to the encoding parameters as the to-be-transmitted audio code stream.

14. The device according to claim 9, characterized in that, The updated audio code stream module is specifically used for: Detecting any one of a trigger operation of the custom audio parameters from the receiving end, an update message of the network bandwidth, and an update operation of the playback parameters, and determining that the audio requirement parameters of the receiving end have been updated.

15. The device according to claim 11, characterized in that, The audio playback device further includes: The segmentation and sorting module is used to segment the audio code stream according to a preset time period to obtain a plurality of audio segments, and sort the audio segments in chronological order; The audio code stream matrix module is used to form an audio code stream matrix for the plurality of audio segments sorted in time according to the manner that the audio segments corresponding to the same time period are placed in the same row and the audio segments corresponding to different encoding parameters are placed in different columns; When the preset time period is equal to the sampling interval of the audio code stream, the updated audio code stream module is specifically used for: Obtain the update time of the audio requirement parameter, and determine the corresponding updated coding parameter according to the updated audio requirement parameter; Locate the position of the updated coding parameter in the audio bitstream matrix, and determine the audio segment corresponding to the time period after the update time in the column where the updated coding parameter is located as the updated audio bitstream.

16. The device according to claim 11, wherein The audio coding unit further includes: A source audio signal receiving subunit, configured to receive a source audio signal collected and uploaded by a sending end at a preset sampling rate; wherein, the preset sampling rate is greater than a preset threshold; An original audio signal obtaining subunit, configured to decode the source audio signal to obtain an original audio signal.

17. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, the audio playback method according to any one of claims 1-8 is implemented.

18. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium, and when the program is executed by a processor, the audio playback method according to any one of claims 1-8 is implemented.

19. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, the audio playback method according to any one of claims 1 to 8 is implemented.

Citation Information

Patent Citations

  • Encoding / decoding device and method

    CN101231850A

  • Audio data encoding method and device, audio data decoding method and device, electronic equipment and storage medium

    CN111128203A