Audio Data Processing Method, Device, Storage Medium and Electronic Device

By segmenting and processing the live audio data, the noise problem during HTML5 streaming audio is solved, achieving better user experience and delay control.

CN114566172BActive Publication Date: 2025-06-13BEIJING KANSHI HIGH-TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210179834.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-02-25
Publication Date
2025-06-13
Estimated Expiration
2042-02-25

AI Technical Summary

Technical Problem

When streaming audio with HTML5, the decoded PCM data will generate noise during continuous playback, affecting the user experience. The existing methods reduce noise but increase delay by increasing the size of the data block.

Method used

By obtaining the audio data sequence generated during the live broadcast process, calling the audio context interface to decode the PCM audio data sequence, dividing it into multiple PCM audio data subsequences, and determining the fade-in or fade-out processing operation based on the PCM audio data at the head and tail positions, the PCM audio data subsequence is processed to reduce noise.

Benefits of technology

Effectively reduces the noise when streaming audio in HTML5, improves the user experience, while avoiding the disadvantage of increasing latency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114566172B_ABST
    Figure CN114566172B_ABST
Patent Text Reader

Abstract

The present disclosure relates to an audio data processing method, apparatus, storage medium and electronic device. The method includes: obtaining an audio data sequence generated during a live broadcast, and calling a decoding method through an audio context interface to decode each audio data block in the audio data sequence to obtain a PCM audio data sequence; segmenting the PCM audio data sequence to obtain a plurality of PCM audio data subsequences; for each PCM audio data subsequence, determining a processing operation for the PCM audio data subsequence according to the PCM audio data at the head and tail positions in the PCM audio data subsequence, where the processing operation is a fade-in processing operation or a fade-out processing operation; for each PCM audio data subsequence, processing each PCM audio data in the PCM audio data subsequence according to the corresponding processing operation for the PCM audio data subsequence to obtain a target PCM audio data sequence, so as to avoid the appearance of noise when playing the audio in a streaming manner using HTML5.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of electronic information technology, and in particular, to an audio data processing method, apparatus, storage medium, and electronic device. Background Art

[0002] In the related art, when using the audio context interface dedicated to streaming audio in HTML5 (HyperText Markup Language 5), the PCM (Pulse Code Modulation) data decoded by the decoding method corresponding to the audio context interface will generate noise during continuous playback, seriously affecting the user experience when watching live classes or listening to teaching audio for a long time.

[0003] Currently, by increasing the size of the data block pulled from the live end to reduce the probability of noise occurrence, but this method does not completely avoid the occurrence of noise, but instead increases the latency of the live broadcast. Summary of the Invention

[0004] To overcome the problems in the related art, the present disclosure provides an audio data processing method, apparatus, storage medium, and electronic device.

[0005] According to a first aspect of an embodiment of the present disclosure, an audio data processing method is provided, including:

[0006] Obtain an audio data sequence generated during a live broadcast, where the audio data sequence includes a plurality of audio data blocks;

[0007] Decode each audio data block in the audio data sequence through a decoding method called by an audio context interface to obtain a PCM audio data sequence;

[0008] Segment the PCM audio data sequence to obtain a plurality of PCM audio data subsequences;

[0009] For each PCM audio data subsequence, determine a processing operation for the PCM audio data subsequence according to the PCM audio data at the head and tail positions of the PCM audio data subsequence, where the processing operation is a fade-in processing operation or a fade-out processing operation;

[0010] For each PCM audio data subsequence, process each PCM audio data in the PCM audio data subsequence according to the corresponding processing operation for the PCM audio data subsequence to obtain a target PCM audio data sequence.

[0011] Optionally, calling the decoding method through the audio context interface to decode each audio data block in the audio data sequence to obtain a PCM audio data sequence, including:

[0012] Calling the decoding method through the audio context interface to decode each audio data block in the audio data sequence to obtain an initial PCM audio data sequence;

[0013] Calling the acquisition method through the audio buffer interface to convert the PCM audio data corresponding to each audio data block in the initial PCM audio data sequence into 32-bit floating-point PCM audio data to obtain the PCM audio data sequence.

[0014] Optionally, splitting the PCM audio data sequence to obtain a plurality of PCM audio data subsequences, including:

[0015] Determining a partitioning rule according to the live scene type, where the partitioning rule is used to characterize the maximum amount of data that can be accommodated in a PCM audio data subsequence;

[0016] Splitting the PCM audio data sequence according to the partitioning rule to obtain a plurality of PCM audio data subsequences.

[0017] Optionally, for each of the PCM audio data subsequences, determining the processing operation of the PCM audio data subsequence according to the PCM audio data at the head and tail positions of the PCM audio data subsequence, including:

[0018] For each of the PCM audio data subsequences, when the volume value corresponding to the first PCM audio data is less than the volume value corresponding to the second PCM audio data, determining that the processing operation of the PCM audio data subsequence is a fade-in processing operation; when the volume value corresponding to the first PCM audio data is greater than the volume value corresponding to the second PCM audio data, determining that the processing operation of the PCM audio data subsequence is a fade-out processing operation;

[0019] Wherein, the first PCM audio data is the PCM audio data at the first position in the PCM audio data subsequence, and the second PCM audio data is the PCM audio data at the last position in the PCM audio data subsequence.

[0020] Optionally, for each of the PCM audio data subsequences, determining the processing operation of the PCM audio data subsequence according to the PCM audio data at the head and tail positions of the PCM audio data subsequence, further including:

[0021] For each of the PCM audio data subsequences, when the volume value corresponding to the first PCM audio data is equal to the volume value corresponding to the second PCM audio data, divide the PCM audio data subsequence, and determine the processing operation of the subsequence based on the PCM audio data at the head and tail positions in the divided subsequence.

[0022] Optionally, for each of the PCM audio data subsequences, processing each PCM audio data in the PCM audio data subsequence according to the processing operation corresponding to the PCM audio data subsequence to obtain a target PCM audio data sequence includes:

[0023] For each of the PCM audio data subsequences, processing each PCM audio data in the PCM audio data subsequence according to a preset curve corresponding to the processing operation corresponding to the PCM audio data subsequence to obtain the target PCM audio data sequence, where the preset curve is used to represent the trend formed by all the PCM audio data in the target PCM audio data sequence.

[0024] Optionally, the method further includes:

[0025] Create an audio cache resource node object by calling a resource creation method through the audio context interface;

[0026] Store the target PCM audio data sequence into the cache attribute of the audio cache resource node object;

[0027] Play the audio corresponding to the target PCM audio data sequence by calling the playback method of the audio cache resource node object.

[0028] Optionally, the method further includes:

[0029] Determine the sum of the noises of the audios corresponding to all the played target PCM audio data sequences;

[0030] When the sum of the noises is greater than a preset noise, update the preset curve, so as to perform corresponding processing operations on the PCM audio data subsequences in the next audio data sequence adjacent to the audio data sequence according to the updated preset curve.

[0031] According to the second aspect of the embodiments of the present disclosure, there is provided an audio data processing device, including:

[0032] An acquisition module, configured to acquire an audio data sequence generated during a live broadcast, where the audio data sequence includes a plurality of audio data blocks;

[0033] A decoding module, configured to call the decoding method of the audio context interface to decode each audio data block in the audio data sequence, so as to obtain a PCM audio data sequence;

[0034] A splitting module, configured to split the PCM audio data sequence to obtain a plurality of PCM audio data subsequences;

[0035] A determining module, configured to, for each of the PCM audio data subsequences, determine a processing operation for the PCM audio data subsequence according to the PCM audio data at the head and tail positions in the PCM audio data subsequence, where the processing operation is a fade-in processing operation or a fade-out processing operation;

[0036] A processing module, configured to, for each of the PCM audio data subsequences, process each PCM audio data in the PCM audio data subsequence according to the corresponding processing operation of the PCM audio data subsequence, so as to obtain a target PCM audio data sequence.

[0037] According to a third aspect of the embodiments of the present disclosure, there is provided a computer-readable storage medium, on which computer program instructions are stored, and when the program instructions are executed by a processor, the steps of the audio data processing method provided in the first aspect of the present disclosure are implemented.

[0038] According to a fourth aspect of the embodiments of the present disclosure, there is provided an electronic device, including:

[0039] A storage device, on which a computer program is stored;

[0040] A processing device, configured to execute the computer program in the storage device to implement the steps of the audio data processing method provided in the first aspect of the present disclosure.

[0041] The technical solutions provided by the embodiments of the present disclosure may include the following beneficial effects: obtaining an audio data sequence generated during a live broadcast, calling a decoding method through an audio context interface to decode each audio data block in the audio data sequence to obtain a PCM audio data sequence; segmenting the PCM audio data sequence to obtain multiple PCM audio data subsequences; for each PCM audio data subsequence, determining a processing operation for the PCM audio data subsequence according to the PCM audio data at the head and tail positions of the PCM audio data subsequence, where the processing operation is a fade-in processing operation or a fade-out processing operation; for each PCM audio data subsequence, processing each PCM audio data in the PCM audio data subsequence according to the corresponding processing operation of the PCM audio data subsequence to obtain a target PCM audio data sequence. Since PCM audio data is a discrete digital signal, a fade-in processing operation or a fade-out processing operation is performed on the PCM audio data to avoid the occurrence of noise when playing audio in a streaming manner using HTML5.

[0042] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and do not limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] The accompanying drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present disclosure and used together with the specification to explain the principles of the present disclosure.

[0044] Figure 1 is a flowchart of an audio data processing method shown according to an exemplary embodiment.

[0045] Figure 2 is a schematic diagram of a preset curve shown according to an exemplary embodiment.

[0046] Figure 3 is another schematic diagram of a preset curve shown according to an exemplary embodiment.

[0047] Figure 4 is a block diagram of an audio data processing device shown according to an exemplary embodiment.

[0048] Figure 5 is a schematic structural diagram of an electronic device shown according to an exemplary embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0049] Exemplary embodiments will be described in detail herein, and examples thereof are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. On the contrary, they are merely examples of apparatuses and methods consistent with some aspects of the present disclosure as detailed in the appended claims.

[0050] Figure 1 is a flowchart of an audio data processing method shown according to an exemplary embodiment, as Figure 1 shown. This audio data processing method can be used in an electronic device, and the electronic device can be a tablet, a smart phone, etc. The audio data processing method includes the following steps.

[0051] Step S101, obtain an audio data sequence generated during a live broadcast, where the audio data sequence includes a plurality of audio data blocks.

[0052] It should be noted that the data in each sequence involved in all embodiments of the present disclosure is sorted in the order of the playback time.

[0053] In some embodiments, the electronic device can set a corresponding cache space, and according to the cache space, the data volume size that the audio data sequence can accommodate can be determined, that is, the number of audio data blocks included in the audio data sequence is determined by the set cache space. Exemplarily, the audio data sequence can include 100 audio data blocks.

[0054] In some embodiments, the format of the audio data block can be MP3 (Moving Picture Experts Group Audio Layer III) format, AAC (Advanced Audio Coding) format, etc.

[0055] Step S102, call a decoding method through an audio context interface to decode each audio data block in the audio data sequence to obtain a PCM audio data sequence.

[0056] It should be noted that the audio context interface is the AudioContext interface in HTML5. The AudioContext interface provides many attributes and methods for creating various audio sources and audio processing modules, etc. The AudioContext interface represents an audio processing graph constructed by linked audio processing modules, and each audio processing module is represented by an AudioNode (audio resource node).

[0057] The decoding methods provided by the AudioContext interface (such as the decodeAudioData method) can be used to asynchronously decode audio data chunks to obtain the PCM audio data corresponding to the audio data chunks. PCM audio data is an uncompressed raw stream of audio sampling data, which is the standard digital audio data converted from an analog signal through sampling, quantization, and encoding. Among them, the parameters used to describe PCM audio data may include sampling frequency, quantization bits, number of channels, a parameter used to identify whether the PCM audio data has a sign bit, byte order, and a parameter used to identify integer or floating point type.

[0058] It should be understood that each PCM audio data in the PCM audio data sequence uniquely corresponds to an audio data chunk in the audio data sequence.

[0059] In some embodiments, Figure 1 The shown step S102 may include: calling a decoding method through the audio context interface to decode each audio data chunk in the audio data sequence to obtain an initial PCM audio data sequence; calling an acquisition method through the audio buffer interface to convert the PCM audio data corresponding to each audio data chunk in the initial PCM audio data sequence into 32-bit floating-point PCM audio data to obtain a PCM audio data sequence.

[0060] It should be noted that the decoding method will obtain a call result, which includes the initial PCM audio data sequence, and this call result will be stored in a cache object. In some embodiments, the audio buffer interface is the AudioBuffer interface in HTML5, and the acquisition method can be the getChannelData() method. The getChannelData() method can be called through the AudioBuffer interface to obtain the data of this cache object, and calling the getChannelData() method can return a PCM audio data sequence with 32-bit floating-point PCM audio data.

[0061] Through the above method, since the data precision of 32-bit floating-point type is higher than that of integer type, therefore, by converting all PCM audio data into 32-bit floating-point type data, the data precision can be improved, and thus the processing result of data processing can be enhanced.

[0062] Step S103, segment the PCM audio data sequence to obtain multiple PCM audio data subsequences.

[0063] It should be noted that the purpose of segmentation is to perform a fade-in operation or a fade-out operation on a continuous segment of PCM audio data. The fade-in operation can gradually increase the audio volume at the start of playback, and the fade-out operation can gradually decrease the audio volume when playback stops, preventing sudden volume changes caused by sound switching. The fade-in operation or the fade-out operation is essentially a linear transformation of the audio volume.

[0064] In some embodiments, step 103 can be implemented in the following manner: Determine a segmentation rule according to the live broadcast scenario type, where the segmentation rule is used to represent the maximum amount of data that a PCM audio data subsequence can accommodate; Segment the PCM audio data sequence according to the segmentation rule to obtain multiple PCM audio data subsequences.

[0065] Among them, the live broadcast scenario type can be determined by identifying the live broadcast screen. For relevant scene recognition technologies, reference can be made to related technologies, which will not be elaborated in this embodiment. The live broadcast scenario type can, for example, include scenarios with high audio quality requirements, such as live concert scenarios, live lecture scenarios, etc.; For another example, the live broadcast scenario type can include scenarios with low audio quality requirements, such as live sales scenarios, etc. For scenarios with high audio quality requirements, obvious noise seriously affects the audio-visual experience. For such scenarios, high-quality audio without noise needs to be played. In this embodiment, the number of segmented PCM audio data subsequences determines the quality of the played audio. The more the number, the lower the noise probability. Therefore, for live broadcast scenarios that require playing high-quality audio without noise, a segmentation rule that can obtain more PCM audio data subsequences can be set. However, at the same time, the increase in the number of PCM audio data subsequences leads to an increase in the number of processing times (such as the processing operations for determining PCM audio data subsequences), and thus leads to a longer processing time and a longer delay. Therefore, for some live broadcast scenarios with high delay requirements and low audio quality requirements, a segmentation rule that can obtain fewer PCM audio data subsequences can be set. The segmentation rule can be used to represent the maximum amount of data that a PCM audio data subsequence can accommodate. Compared with a segmentation rule with a smaller data amount, a segmentation rule with a larger data amount results in fewer PCM audio data subsequences obtained by segmenting the PCM audio data sequence according to this rule, and a segmentation rule with a smaller data amount results in more PCM audio data subsequences obtained by segmenting the PCM audio data sequence according to this rule.

[0066] Through the above method, considering the requirements of two dimensions, namely audio quality and delay degree, from the actual live broadcast scenario type, determine the number of PCM audio data subsequences corresponding to the PCM audio data sequence according to the actual live broadcast scenario type, so as to balance the audio quality and the delay degree.

[0067] Step S104: For each PCM audio data subsequence, determine the processing operation for this PCM audio data subsequence according to the PCM audio data at the head and tail positions in this PCM audio data subsequence. The processing operation is a fade-in processing operation or a fade-out processing operation.

[0068] For example, in the case where the volume value corresponding to the first PCM audio data is less than the volume value corresponding to the second PCM audio data, determine that the processing operation for this PCM audio data subsequence is a fade-in processing operation; in the case where the volume value corresponding to the first PCM audio data is greater than the volume value corresponding to the second PCM audio data, determine that the processing operation for this PCM audio data subsequence is a fade-out processing operation. Herein, the first PCM audio data is the PCM audio data at the first position in this PCM audio data subsequence, and the second PCM audio data is the PCM audio data at the last position in this PCM audio data subsequence.

[0069] In some embodiments, in the case where the volume value corresponding to the first PCM audio data is equal to the volume value corresponding to the second PCM audio data, the PCM audio data subsequence can be divided, and then determine the processing operation for this subsequence according to the PCM audio data at the head and tail positions in the subsequence obtained after division. In the case where it is impossible to determine whether the PCM audio data subsequence is a fade-in processing operation or a fade-out processing operation, further divide this PCM audio data subsequence to improve the accuracy of determining the fade-in and fade-out processing operations. It should be noted that the specific implementation manner of determining the processing operation for this subsequence according to the PCM audio data at the head and tail positions in this subsequence in this embodiment can refer to the examples of the first PCM audio data and the second PCM audio data above, and this embodiment will not be elaborated herein.

[0070] Step S105: For each PCM audio data subsequence, process each PCM audio data in this PCM audio data subsequence according to the corresponding processing operation to obtain a target PCM audio data sequence.

[0071] It should be noted that the fade-in processing operation and the fade-out processing operation are used to transform the PCM audio data to achieve a gradual increase and decrease in volume. Calculate the transformation coefficient of the current PCM data in the PCM audio data subsequence according to the number of data in the PCM audio data subsequence and the preset curve, and then calculate the corresponding data according to the transformation coefficient and the current PCM data. This data is the data after the fade-in processing operation or the fade-out processing operation.

[0072] In some embodiments, the preset curve is a curve representing the relationship between time and amplitude. The amplitude trend formed by all the PCM audio data in the target PCM audio data sequence obtained by processing the PCM audio data subsequence according to the preset curve is consistent with the trend presented by the preset curve.

[0073] In some embodiments, the preset curve can be a linear curve, such as y = kx, where y is the amplitude, x is the time, and k is the transformation coefficient. The following further explains step S105 with the curve of y = kx as the preset curve.

[0074] Refer to Figure 2 , the preset curve includes curve A1 and curve B1. Perform a fade-in processing operation according to curve A1 in Figure 2 , and perform a fade-out processing operation according to curve B1 in Figure 2 . Figure 2 The horizontal axis in Figure 2 represents time, and the vertical axis represents the change in amplitude. The change in amplitude corresponds to the change in volume. In Figure 2 , the change in amplitude is between 0 and 1. If the playback times of the PCM audio data at the beginning and end positions in the PCM audio data subsequence are t1 and t2 respectively, then the PCM audio data subsequence includes the PCM audio data from playback time t1 to t2. According to t1 and t2 and the sampling frequency, the number of samples between t1 and t2 can be calculated, that is, the number of PCM audio data from t1 to t2. Taking the example that the PCM audio data from t1 to t2 includes 100 data and performing a fade-in processing operation on these 100 data, the calculation methods for the 0th, 38th, and 70th PCM audio data can be as follows:

[0075] For the 0th PCM audio data, k 0 = 0 / 100, out 0 = in 0 * k 0 , where k 0 is the transformation coefficient corresponding to the 0th PCM audio data, out 0 is the data obtained by performing a fade-in processing operation on the 0th PCM audio data, and in 0 is the 0th PCM audio data.

[0076] For the 38th PCM audio data, k 38 = 38 / 100, out 38 = in 38 * k 38 , where k 38 is the transformation coefficient corresponding to the 38th PCM audio data, out 38 is the data obtained by performing a fade-in processing operation on the 38th PCM audio data, and in38 is the 38th PCM audio data.

[0077] For the 70th PCM audio data, k 70 = 70 / 100, out 70 = in 70 * k 70 where k 70 is the transformation coefficient corresponding to the 70th PCM audio data, and out 70 is the data obtained by performing a fade-in processing operation on the 70th PCM audio data, and in 70 is the 70th PCM audio data.

[0078] In some embodiments, the preset curve can also be a non-linear curve, such as a sine curve, a non-linear curve having the same trend as an equal-power curve, etc. Referring to Figure 3 For the non-linear curve having the same trend as the equal-power curve shown in Figure 3 as shown, Figure 3 the preset curve shown in Figure 3 includes curve A2 and curve B2. Perform a fade-in processing operation according to curve A2 in Figure 3 and perform a fade-out processing operation according to curve B2 in

[0079] The horizontal axis of this non-linear curve represents time, and the vertical axis represents the change in amplitude. For such non-linear curves, they intersect at a higher amplitude, which can minimize the volume drop between audio regions (between PCM audio data subsequences and PCM audio data subsequences), resulting in a more uniform cross-fade between slightly different audio regions in terms of volume level.

[0080] In some embodiments, the audio data processing method further includes: creating an audio cache resource node object by calling a resource creation method through an audio context interface; storing a target PCM audio data sequence in a cache attribute of the audio cache resource node object; and playing an audio corresponding to the target PCM audio data sequence by calling a play method of the audio cache resource node object.

[0081] Among them, the creation method can be the createBufferSource() method. After creating an audio buffer resource node object (the AudioBufferSourceNode object in HTML5), the target PCM audio data sequence is assigned to the buffer attribute of the audio buffer resource node object. At this time, the audio has been passed to the audio source (which can be understood as the audio rendering device). By calling the play method of the audio buffer resource node object, the audio rendering device can play the audio data pointed to by the buffer attribute.

[0082] In some embodiments, the method further includes: determining the sum of the noises of the audios corresponding to all target PCM audio data sequences to be played; in the case where the sum of the noises is greater than a preset noise, updating the preset curve, so as to perform corresponding processing operations on the PCM audio data subsequence in the next audio data sequence adjacent to the audio data sequence according to the updated preset curve.

[0083] Taking the preset curve as an example of a sine curve or a non-linear curve consistent with the trend of the equal-power curve, the PCM audio data subsequence corresponding to the currently obtained audio data sequence uses a sine curve (for example, a sine curve) to perform a fade-in processing operation or a fade-out operation. After playing the audios corresponding to all target PCM audio data sequences obtained by processing the PCM audio data subsequence of this audio data sequence, determining the sum of the noises of the audios corresponding to all target PCM audio data sequences to be played, comparing the sum of the noises with the preset noise. When the sum of the noises is smaller than the preset noise, the preset curve (i.e., the sine curve) used for processing this audio data sequence may not be updated. In the case where the sum of the noises is greater than the preset noise, update the preset curve used for processing this audio data sequence (i.e., the sine curve can be replaced with a non-linear curve consistent with the trend of the equal-power curve). That is, when a new audio data sequence (for example, the next audio data sequence adjacent to the current audio data sequence) is obtained next time, the fade-in processing operation or the fade-in processing operation can be performed on the next audio data sequence according to the non-linear curve consistent with the trend of the equal-power curve.

[0084] In the above manner, during the actual audio playback process, the preset curve for performing the fade-in processing operation or the fade-in processing operation is adjusted according to the actual noise, so as to ensure that the sum of the noises of the audios played during the entire live broadcast process is minimized.

[0085] The present disclosure also provides an audio data processing device. Figure 4 It is a block diagram of an audio data processing device shown according to an exemplary embodiment. Refer to Figure 4 This audio data processing device 400 includes:

[0086] An acquisition module 401, configured to acquire an audio data sequence generated during a live broadcast, where the audio data sequence includes a plurality of audio data blocks;

[0087] A decoding module 402, configured to call a decoding method of an audio context interface to decode each audio data block in the audio data sequence, so as to obtain a PCM audio data sequence;

[0088] A splitting module 403, configured to split the PCM audio data sequence to obtain a plurality of PCM audio data subsequences;

[0089] A determination module 404, configured to, for each PCM audio data subsequence, determine a processing operation of the PCM audio data subsequence according to the PCM audio data at the head and tail positions in the PCM audio data subsequence, where the processing operation is a fade-in processing operation or a fade-out processing operation;

[0090] A processing module 405, configured to, for each PCM audio data subsequence, process each PCM audio data in the PCM audio data subsequence according to the processing operation corresponding to the PCM audio data subsequence to obtain a target PCM audio data sequence.

[0091] Optionally, the decoding module 402 includes:

[0092] A decoding sub-module, configured to call a decoding method through an audio context interface to decode each audio data block in the audio data sequence to obtain an initial PCM audio data sequence;

[0093] A conversion sub-module, configured to call an acquisition method through an audio buffer interface to convert the PCM audio data corresponding to each audio data block in the initial PCM audio data sequence into 32-bit floating-point PCM audio data, so as to obtain the PCM audio data sequence.

[0094] Optionally, the splitting module 403 includes:

[0095] A splitting rule determination sub-module, configured to determine a splitting rule according to a live broadcast scene type, where the splitting rule is used to represent the maximum data volume that a PCM audio data subsequence can accommodate;

[0096] A splitting sub-module, configured to split the PCM audio data sequence according to the splitting rule to obtain a plurality of PCM audio data subsequences.

[0097] Optionally, the determining module 404 includes a first determining sub-module, configured to, for each of the PCM audio data subsequences, determine that the processing operation for this PCM audio data subsequence is a fade-in processing operation when the volume value corresponding to the first PCM audio data is less than the volume value corresponding to the second PCM audio data; and determine that the processing operation for this PCM audio data subsequence is a fade-out processing operation when the volume value corresponding to the first PCM audio data is greater than the volume value corresponding to the second PCM audio data;

[0098] Wherein, the first PCM audio data is the PCM audio data at the first position in this PCM audio data subsequence, and the second PCM audio data is the PCM audio data at the last position in this PCM audio data subsequence.

[0099] Optionally, the determining module 404 further includes a second determining sub-module, configured to, for each of the PCM audio data subsequences, when the volume value corresponding to the first PCM audio data is equal to the volume value corresponding to the second PCM audio data, divide this PCM audio data subsequence, and determine the processing operation for this subsequence based on the PCM audio data at the head and tail positions in the divided subsequences.

[0100] Optionally, the processing module 405 is specifically configured to, for each of the PCM audio data subsequences, process each PCM audio data in this PCM audio data subsequence according to a preset curve corresponding to the processing operation corresponding to this PCM audio data subsequence, to obtain the target PCM audio data sequence, wherein the preset curve is used to represent the trend formed by all the PCM audio data in the target PCM audio data sequence.

[0101] Optionally, the audio data processing device 400 further includes:

[0102] A creation module, configured to create an audio cache resource node object by invoking a resource creation method through the audio context interface;

[0103] A storage module, configured to store the target PCM audio data sequence into the cache attribute of the audio cache resource node object;

[0104] A playback module, configured to play the audio corresponding to the target PCM audio data sequence by invoking the playback method of the audio cache resource node object.

[0105] Optionally, the device 400 further includes:

[0106] A noise determination module, configured to determine the sum of the noises of the audios corresponding to all the target PCM audio data sequences played;

[0107] An update module, configured to update the preset curve when the sum of the noises is greater than a preset noise, so as to perform corresponding processing operations on the PCM audio data subsequence in the next audio data sequence adjacent to the audio data sequence according to the updated preset curve.

[0108] Regarding the device in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated herein.

[0109] The present disclosure also provides a computer-readable storage medium, on which computer program instructions are stored, and when the program instructions are executed by a processor, the steps of the audio data processing method provided by the present disclosure are implemented.

[0110] The present disclosure also provides an electronic device, including:

[0111] A storage device, on which a computer program is stored;

[0112] A processing device, configured to execute the computer program in the storage device to implement the steps of the audio data processing method provided by the present disclosure.

[0113] Figure 5 is a block diagram of an electronic device 500 shown according to an exemplary embodiment. For example, the electronic device 500 may be a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.

[0114] Referring to Figure 5 , the electronic device 500 may include one or more of the following components: a processing component 502, a memory 504, a power component 506, a multimedia component 508, an audio component 510, an input / output (I / O) interface 512, a sensor component 514, and a communication component 516.

[0115] The processing component 502 generally controls the overall operation of the electronic device 500, such as operations associated with display, telephone calls, data communication, camera operations, and recording operations. The processing component 502 may include one or more processors 520 to execute instructions to complete all or part of the steps of the above audio data processing method. In addition, the processing component 502 may include one or more modules to facilitate the interaction between the processing component 502 and other components. For example, the processing component 502 may include a multimedia module to facilitate the interaction between the multimedia component 508 and the processing component 502.

[0116] The memory 504 is configured to store various types of data to support the operation of the electronic device 500. Examples of such data include instructions for any application or method operating on the electronic device 500, contact data, phone book data, messages, pictures, videos, and the like. The memory 504 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk, or optical disk.

[0117] The power component 506 provides power to the various components of the electronic device 500. The power component 506 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the electronic device 500.

[0118] The multimedia component 508 includes a screen that provides an output interface between the electronic device 500 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors can sense not only the boundaries of the touch or swipe actions but also detect the duration and pressure associated with the touch or swipe operation. In some embodiments, the multimedia component 508 includes a front camera and / or a rear camera. When the electronic device 500 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera can receive external multimedia data. Each of the front camera and the rear camera can be a fixed optical lens system or have a focal length and optical zoom capabilities.

[0119] The audio component 510 is configured to output and / or input audio signals. For example, the audio component 510 includes a microphone (MIC) that is configured to receive external audio signals when the electronic device 500 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signals can be further stored in the memory 504 or transmitted via the communication component 516. In some embodiments, the audio component 510 further includes a speaker for outputting audio signals.

[0120] The I / O interface 512 provides an interface between the processing component 502 and a peripheral interface module, which can be a keyboard, a click wheel, buttons, etc. These buttons can include, but are not limited to: a home button, a volume button, a power button, and a lock button.

[0121] The sensor assembly 514 includes one or more sensors for providing a status assessment of various aspects of the electronic device 500. For example, the sensor assembly 514 can detect the on / off state of the electronic device 500, the relative positioning of components, such as the display and keypad of the electronic device 500. The sensor assembly 514 can also detect a change in the position of the electronic device 500 or a component of the electronic device 500, the presence or absence of user contact with the electronic device 500, the orientation or acceleration / deceleration of the electronic device 500, and a change in the temperature of the electronic device 500. The sensor assembly 514 can include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor assembly 514 can also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor assembly 514 can also include an acceleration sensor, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.

[0122] The communication component 516 is configured to facilitate communication between the electronic device 500 and other devices in a wired or wireless manner. The electronic device 500 can access a wireless network based on communication standards, such as WiFi, 2G, or 3G, or a combination thereof. In an exemplary embodiment, the communication component 516 receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 516 further includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.

[0123] In an exemplary embodiment, the electronic device 500 can be implemented by one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components for performing the above-described audio data processing method.

[0124] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions, such as a memory 504 including instructions, is also provided. The above instructions can be executed by a processor 520 of the electronic device 500 to complete the above-described audio data processing method. For example, the non-transitory computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device, etc.

[0125] In another exemplary embodiment, there is also provided a computer program product, which includes a computer program executable by a programmable electronic device, and the computer program has a code portion for performing the above-described audio data processing method when executed by the programmable electronic device.

[0126] Those skilled in the art will readily conceive of other embodiments of the present disclosure after considering the specification and practicing the present disclosure. This application is intended to cover any variations, uses, or adaptations of the present disclosure, which follow the general principles of the present disclosure and include well-known common general knowledge or conventional technical means in the technical field not disclosed in the present disclosure. The specification and embodiments are only regarded as exemplary, and the true scope and spirit of the present disclosure are pointed out by the following claims.

[0127] It should be understood that the present disclosure is not limited to the exact structures already described and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.

Claims

1. An audio data processing method, characterized in that, it includes: Obtain an audio data sequence generated during the live broadcast, where the audio data sequence includes multiple audio data blocks; Call a decoding method through an audio context interface to decode each audio data block in the audio data sequence to obtain a PCM audio data sequence; Segment the PCM audio data sequence to obtain multiple PCM audio data subsequences; For each of the PCM audio data subsequences, determine the processing operation of the PCM audio data subsequence according to the PCM audio data at the head and tail positions of the PCM audio data subsequence, and the processing operation is a fade-in processing operation or a fade-out processing operation; For each of the PCM audio data subsequences, process each PCM audio data in the PCM audio data subsequence according to the processing operation corresponding to the PCM audio data subsequence to obtain a target PCM audio data sequence; The step of determining the processing operation of each PCM audio data subsequence according to the PCM audio data at the head and tail positions of the PCM audio data subsequence includes: for each of the PCM audio data subsequences, when the volume value corresponding to the first PCM audio data is less than the volume value corresponding to the second PCM audio data, determine that the processing operation of the PCM audio data subsequence is a fade-in processing operation; when the volume value corresponding to the first PCM audio data is greater than the volume value corresponding to the second PCM audio data, determine that the processing operation of the PCM audio data subsequence is a fade-out processing operation; where the first PCM audio data is the PCM audio data at the first position in the PCM audio data subsequence, and the second PCM audio data is the PCM audio data at the last position in the PCM audio data subsequence; The step of processing each PCM audio data in the PCM audio data subsequence according to the processing operation corresponding to the PCM audio data subsequence to obtain a target PCM audio data sequence includes: for each of the PCM audio data subsequences, process each PCM audio data in the PCM audio data subsequence according to the preset curve corresponding to the processing operation corresponding to the PCM audio data subsequence to obtain the target PCM audio data sequence, where the preset curve is used to represent the trend formed by all PCM audio data in the target PCM audio data sequence.

2. The method according to claim 1, characterized in that, The step of calling a decoding method through an audio context interface to decode each audio data block in the audio data sequence to obtain a PCM audio data sequence includes: Call a decoding method through an audio context interface to decode each audio data block in the audio data sequence to obtain an initial PCM audio data sequence; Convert the PCM audio data corresponding to each audio data block in the initial PCM audio data sequence into 32-bit floating-point PCM audio data through the method called by the audio buffer interface to obtain the PCM audio data sequence.

3. The method according to claim 1, wherein, The splitting of the PCM audio data sequence to obtain multiple PCM audio data subsequences includes: Determine a division rule according to the live scene type, where the division rule is used to characterize the maximum amount of data that a PCM audio data subsequence can accommodate; According to the division rule, split the PCM audio data sequence to obtain multiple PCM audio data subsequences.

4. The method according to claim 1, wherein, For each of the PCM audio data subsequences, determining the processing operation of the PCM audio data subsequence according to the PCM audio data at the head and tail positions of the PCM audio data subsequence further includes: For each of the PCM audio data subsequences, in the case where the volume value corresponding to the first PCM audio data is equal to the volume value corresponding to the second PCM audio data, divide the PCM audio data subsequence, and based on the PCM audio data at the head and tail positions of the divided subsequence, determine the processing operation of the subsequence.

5. The method according to claim 1, wherein, The method further includes: Create an audio buffer resource node object through the resource creation method called by the audio context interface; Store the target PCM audio data sequence in the cache attribute of the audio buffer resource node object; Play the audio corresponding to the target PCM audio data sequence by calling the play method of the audio buffer resource node object.

6. The method according to claim 5, wherein, The method further includes: Determine the sum of the noises of the audio corresponding to all the target PCM audio data sequences played; In the case where the sum of the noises is greater than the preset noise, update the preset curve, so as to perform corresponding processing operations on the PCM audio data subsequences in the next audio data sequence adjacent to the audio data sequence obtained according to the updated preset curve.

7. An audio data processing device, wherein, includes: An acquisition module for acquiring an audio data sequence generated during a live broadcast, where the audio data sequence includes a plurality of audio data blocks; A decoding module for calling the decoding method of the audio context interface to decode each audio data block in the audio data sequence to obtain a PCM audio data sequence; A splitting module for splitting the PCM audio data sequence to obtain multiple PCM audio data subsequences; A determination module for, for each of the PCM audio data subsequences, determining the processing operation of the PCM audio data subsequence according to the PCM audio data at the head and tail positions of the PCM audio data subsequence, and the processing operation is a fade-in processing operation or a fade-out processing operation; A processing module, configured to, for each of the PCM audio data subsequences, process each PCM audio data in the PCM audio data subsequence according to a processing operation corresponding to the PCM audio data subsequence, so as to obtain a target PCM audio data sequence; The determining module includes a first determining sub-module, and the first determining sub-module is configured to, for each of the PCM audio data subsequences, when the volume value corresponding to the first PCM audio data is less than the volume value corresponding to the second PCM audio data, determine that the processing operation of the PCM audio data subsequence is a fade-in processing operation; When the volume value corresponding to the first PCM audio data is greater than the volume value corresponding to the second PCM audio data, determine that the processing operation of the PCM audio data subsequence is a fade-out processing operation; wherein, the first PCM audio data is the PCM audio data at the first position in the PCM audio data subsequence, and the second PCM audio data is the PCM audio data at the last position in the PCM audio data subsequence; The processing module is specifically configured to, for each of the PCM audio data subsequences, process each PCM audio data in the PCM audio data subsequence according to a preset curve corresponding to the processing operation corresponding to the PCM audio data subsequence, so as to obtain the target PCM audio data sequence, where the preset curve is used to represent the trend formed by all the PCM audio data in the target PCM audio data sequence.

8. A computer-readable storage medium, on which computer program instructions are stored, Characterized in that, When the program instructions are executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.

9. An electronic device, Characterized in that, Comprising: A storage device, on which a computer program is stored; A processing device, configured to execute the computer program in the storage device to implement the steps of the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Call audio mixing processing method and device, storage medium and computer equipment

    CN111048119A

  • Audio and video playing method and device, electronic equipment and storage medium

    CN113825009A