Audio processing method and related device
By splitting the audio stream into data segments less than or equal to the player's maximum set length, and performing delay processing and forward buffer gain value calculation, the problem of over-sounding sound when playing audio and video is solved, achieving a better user experience.
Patent Information
- Application Number
- CN202510360834.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-06-13
AI Technical Summary
If the player comes with a sound effect when playing audio and video, it may cause excessive sound and reduce the user experience.
By splitting the audio stream to be processed into multiple data segments, the length of each segment is less than or equal to the maximum set length that the player can process, and each data segment is subjected to delay processing and prospective buffer gain value calculation, and the product of the delay data and prospective buffer gain value is finally used as the processing result.
It effectively avoids the situation where the audio signal exceeds the player's processing range due to frame length exceeding the standard, avoids the problem of sound overexposure and improves the user experience.
Smart Images

Figure CN120148529A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of audio playback, and more specifically, to an audio processing method and related device. Background Art
[0002] In application programs APP such as audio and video software, there are usually video and audio sources with different coding formats, and audio and video in various coding formats have different frame lengths.
[0003] When a player plays audio and video, if the player has built-in sound effects, there may be a situation where the audio signal after adding the sound effects exceeds the range that the player can process, which may lead to the problem of over-explosive sound and reduce the user experience. Summary of the Invention
[0004] In view of this, this application provides an audio processing method and related device to solve the problem that when a player plays audio and video, if the player has built-in sound effects, there may be a problem of over-explosive sound and reduce the user experience.
[0005] To solve the above technical problems, this application adopts the following technical solutions:
[0006] An audio processing method includes:
[0007] Splitting the audio stream to be processed into multiple data segments; the length of the data segment is less than or equal to the maximum set length that the player can process;
[0008] Determining the data to be processed that needs audio optimization corresponding to the data segment;
[0009] For each piece of data to be processed, performing a delay process on the data to be processed to obtain the delayed data corresponding to each piece of data to be processed, and determining the look-ahead buffer gain value corresponding to the data to be processed;
[0010] Taking the product of the delayed data corresponding to the data to be processed and the look-ahead buffer gain value as the processing result corresponding to the data to be processed.
[0011] Optionally, splitting the audio stream to be processed into multiple data segments includes:
[0012] Performing a parsing operation on the audio stream to be processed to obtain each audio frame in the audio stream;
[0013] Dividing the audio frames according to the maximum set length to obtain multiple data segments.
[0014] Optionally, determining the data to be processed that needs audio optimization corresponding to the data segment includes:
[0015] When the length of the audio frame is a multiple of the maximum set length, directly use the data segment as the data to be processed that needs audio optimization;
[0016] When the length of the audio frame is not a multiple of the maximum set length, based on the data segment and the adjacent data segments of the data segment, determine the data to be processed that needs audio optimization.
[0017] Optionally, perform a delay process on the data to be processed to obtain delay data corresponding to each data to be processed, including:
[0018] Obtain target data to be processed, and the target data to be processed is initially the first data to be processed in the data to be processed;
[0019] Obtain padding data set in the delay buffer area, and the delay buffer area is initially configured with multiple data whose padding value is a set value;
[0020] Add the padding data to the front of the target data to be processed to obtain combined data;
[0021] In the combined data, screen out the part with the same length as the target data to be processed in the order from front to back, and use it as the delay data corresponding to the data to be processed. Use the remaining part in the combined data as the updated padding data, and update the updated padding data to the delay buffer area;
[0022] Use the next data to be processed after the target data to be processed as the new target data to be processed, return to the step of obtaining the target data to be processed, and execute sequentially until the delay data corresponding to each data to be processed is obtained.
[0023] Optionally, determine the look-ahead buffer gain value corresponding to the data to be processed, including:
[0024] Calculate the loudness value of the data to be processed;
[0025] Perform a correction operation on the loudness value to obtain a corrected value corresponding to the loudness value;
[0026] Use the loudness value and the corrected value to calculate a gain threshold;
[0027] Use the gain smoothing formula and the gain threshold to obtain the gain smoothing matrix corresponding to the data to be processed; the attack time used when determining the gain smoothing formula is a set value. In the gain smoothing formula, the gain value of the first sampling point is the gain value of the last sampling point of the previous data to be processed of the data to be processed; the gain value of the last sampling point is the gain value of the first sampling point of the next data to be processed of the data to be processed;
[0028] Perform look-ahead buffer delay processing on the gain smoothing matrix corresponding to the data to be processed, to obtain the look-ahead buffer gain value corresponding to the data to be processed.
[0029] Optionally, performing look-ahead buffer delay processing on the gain smoothing matrix corresponding to the data to be processed, to obtain the look-ahead buffer gain value corresponding to the data to be processed, includes:
[0030] Write the gain smoothing matrix corresponding to the data to be processed into the look-ahead buffer, to obtain the look-ahead buffer gain, and the write position of the gain smoothing matrix;
[0031] Perform dynamic gain adjustment on the look-ahead buffer gain, to obtain the adjusted look-ahead buffer gain value;
[0032] Read the adjusted look-ahead buffer gain value from the look-ahead buffer, to obtain the look-ahead buffer gain value corresponding to the data to be processed.
[0033] Optionally, performing dynamic gain adjustment on the look-ahead buffer gain, to obtain the adjusted look-ahead buffer gain value, includes:
[0034] Obtain the look-ahead buffer gain corresponding to the index value;
[0035] In the case where the obtained look-ahead buffer gain is greater than the gain attenuation reference value, set the look-ahead buffer gain corresponding to the index value in the look-ahead buffer to the gain attenuation reference value, and use the sum of the step size and the gain attenuation reference value as the new gain attenuation reference value;
[0036] In the case where the obtained look-ahead buffer gain is less than the gain attenuation reference value, calculate a new step size by using the obtained look-ahead buffer gain and the initial length in the delay buffer area, and use the sum of the new step size and the gain attenuation reference value as the new gain attenuation reference value.
[0037] An audio processing device, includes:
[0038] An audio stream processing module, configured to split an audio stream to be processed into a plurality of data segments; the length of the data segment is less than or equal to the maximum set length that the player can process;
[0039] A data determination module, configured to determine the data to be processed that needs to perform audio optimization corresponding to the data segment;
[0040] A data processing module, configured to, for each of the data to be processed, perform delay processing on the data to be processed, to obtain the delay data corresponding to each of the data to be processed, and determine the look-ahead buffer gain value corresponding to the data to be processed;
[0041] An audio optimization module, which is used to take the product of the delay data corresponding to the data to be processed and the look-ahead buffer gain value as the processing result corresponding to the data to be processed.
[0042] Optionally, the data processing module includes:
[0043] A first data acquisition sub-module, which is used to acquire target data to be processed, and the target data to be processed is initially the first data to be processed among the data to be processed;
[0044] A second data acquisition sub-module, which is used to acquire padding data set in a delay buffer area, and the delay buffer area is initially configured with multiple data whose padding values are set values;
[0045] A combination sub-module, which is used to add the padding data to the front of the target data to be processed to obtain combined data;
[0046] A delay sub-module, which is used to screen out, in the combined data, a part with the same length as the target data to be processed in the order from front to back as the delay data corresponding to the data to be processed, take the remaining part in the combined data as the updated padding data, and update the updated padding data to the delay buffer area;
[0047] A judgment sub-module, which is used to judge whether the delay data corresponding to each data to be processed is obtained;
[0048] A data determination sub-module, which is used, after the judgment sub-module judges that the delay data corresponding to each data to be processed is not obtained, to take the next data to be processed after the target data to be processed as the new target data to be processed;
[0049] After the data determination sub-module takes the next data to be processed after the target data to be processed as the new target data to be processed, the first data acquisition sub-module is used to acquire the target data to be processed until the delay data corresponding to each data to be processed is obtained.
[0050] An electronic device, including at least one processor and a memory connected to the processor, wherein:
[0051] The memory is used to store a computer program;
[0052] The processor is used to execute the computer program so that the electronic device can implement the above-mentioned audio processing method.
[0053] The present application provides an audio processing method and related devices. In the present application, if the frame length of the played audio is different from the frame length of the audio processed by the player, the audio stream to be processed will be split into multiple data segments, and the length of the data segments is less than or equal to the maximum set length that the player can process, so that the frame length of the data segments is within the frame length range that the player can process, thereby avoiding the situation where the audio signal after adding sound effects exceeds the range that the player can process due to the excessive frame length, and further avoiding the problem of sound overexposure. In addition, in the present application, the product of the delay data corresponding to the data to be processed and the look-ahead buffer gain value is used as the processing result corresponding to the data to be processed, that is, the data to be processed is gain-adjusted by the look-ahead buffer gain value, avoiding the situation where the audio signal after adding sound effects exceeds the range that the player can process, and further avoiding the problem of sound overexposure. Description of the Drawings
[0054] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following will briefly introduce the drawings required for use in the description of the embodiments or related technologies. Obviously, the drawings in the following description are only the embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.
[0055] Figure 1 It is a flowchart of an audio processing method provided by an embodiment of the present application;
[0056] Figure 2 It is a schematic diagram of time setting provided by an embodiment of the present application;
[0057] Figure 3 It is a logic diagram of an audio processing method provided by an embodiment of the present application;
[0058] Figure 4 It is a logic diagram of a method for determining delay data provided by an embodiment of the present application;
[0059] Figure 5 It is a schematic diagram of data to be processed and delay data provided by an embodiment of the present application;
[0060] Figure 6 It is a flowchart of a method for determining the look-ahead buffer gain value provided by an embodiment of the present application;
[0061] Figure 7 It is a schematic structural diagram of an audio processing device provided by an embodiment of the present application;
[0062] Figure 8 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application. Detailed Embodiments
[0063] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.
[0064] In application programs APP such as audio and video software, there are usually video and audio sources in different coding formats, and the audio and video in various coding formats have different frame lengths.
[0065] When the player plays audio and video, if the player has its own sound effects (sound effects are a way to enhance the user's viewing experience, and audio and video may cover many sound effect options), due to the different loudness of the audio and video, there may be a situation where the audio signal after adding the sound effects exceeds the processing range of the player, which may lead to the problem of over-exploded sound and reduce the user experience.
[0066] Analyzing the reasons for the over-exploded sound problem, it can be known that there are the following two points:
[0067] 1. Since the frame length of the audio is different from the frame length supported by the player, there may be a problem of over-exploded sound caused by the player's inability to process the audio of this frame length.
[0068] 2. The limiter cannot be combined with the player. Among them, the limiter is an audio post-processing algorithm applied to various Digital Audio Workstations (DAWs) to assist users in preventing audio distortion during the post-mixing process. However, during the use of the limiter, the traditional limiter algorithm is based on file processing and requires the entire file stream to be loaded into memory, while the player's processing method is to read and play frame by frame and cannot process the entire file stream (that is, the player logic only reads and processes data through frame addressing, rather than the entire file). Therefore, the limiter cannot be combined with the player to achieve the effect of audio limiting to avoid the problem of over-exploded sound. In addition, as a reference numerical algorithm, the limiter will affect the effect due to the discontinuity of numerical reference, and even introduce noise in severe cases.
[0069] Therefore, in the embodiments of the present application, for the frame length problem, the audio stream to be processed is split into multiple data segments, and the length of the data segments is less than or equal to the maximum set length that the player can process, so that the frame length of the data segments is within the range of the frame length supported by the player, thereby avoiding the situation where the audio signal after adding the sound effects exceeds the processing range of the player due to the excessive frame length, and further avoiding the problem of over-exposed sound.
[0070] Regarding the problem that the limiter cannot be combined with the player, in the embodiments of the present application, in the improved gain smoothing formula, the gain value of the first sampling point is the gain value of the last sampling point of the previous data to be processed of the data to be processed; the gain value of the last sampling point is the gain value of the first sampling point of the next data to be processed of the data to be processed, realizing the continuity of the gain between frames, thus avoiding the problem of affecting the effect due to the discontinuity of the numerical reference, and even introducing noise in severe cases. In addition, the audio is gain-adjusted by the look-ahead buffer gain value to avoid the situation that the audio signal after adding the sound effect exceeds the range that the player can process, and further avoid the problem of sound overexposure.
[0071] Based on the above, an embodiment of the present application provides an audio processing method, and the execution entity can be a player. Referring to Figure 1 , an audio processing method may include:
[0072] S11. Split the audio stream to be processed into multiple data segments.
[0073] Among them, the audio stream to be processed can be only a single audio stream or the audio stream in a video. In this embodiment, the frame length of the audio stream is not limited and can be any frame length such as 1000, 4000, 1024, etc.
[0074] The length of the data segment is less than or equal to the maximum set length that the player can process. In one implementation, the maximum frame length supported by a player is 1024. Therefore, the maximum set length in this embodiment can be 1024.
[0075] In one implementation, splitting the audio stream to be processed into multiple data segments includes:
[0076] Perform a parsing operation on the audio stream to be processed to obtain each audio frame in the audio stream, and perform segmentation processing on the audio frames according to the maximum set length to obtain multiple data segments.
[0077] Specifically, parse the audio stream obtained by the player or the audio stream in the video to obtain the complete audio stream PCM (Pulse Code Modulation) data, determine the position of the first frame in the data, and start processing from the first frame of the audio stream. Process the audio stream to obtain each audio frame. The frame length of the audio frames in this embodiment is not limited and can be any frame length such as 1000, 4000, 1024, etc.
[0078] Then, taking 1024 as an example of the maximum set length, the audio frames are segmented to obtain multiple data segments. For example, if the frame length of an audio frame is 2048, when segmented by 1024, 2048 / 1024 = 2, then the audio frame can be split into two data segments. Another example, if the frame length of an audio frame is 4000, when segmented by 1024, the quotient of 4000 and 1024 is 3 with a remainder, then the audio frame can be split into four data segments. The frame lengths of the first three data segments are all 1024, and the frame length of the last data segment is 928.
[0079] Therefore, for each audio frame in the audio stream, the maximum set length is used to split the audio frame into multiple data segments.
[0080] S12. Determine the data to be processed that needs audio optimization corresponding to the data segment.
[0081] In this embodiment, since the frame lengths of some data segments are 1024 and the frame lengths of some data segments are less than 1024, and the player processes audio according to a frame length of at most 1024, therefore, for a data segment, if its frame length is less than 1024, it needs to be processed separately, or combined with the data of adjacent data segments into data with a frame length of 1024 and then processed together. Which specific method to use is set according to the actual situation.
[0082] In one implementation manner, determining the data to be processed that needs audio optimization corresponding to the data segment includes:
[0083] When the length of the audio frame is a multiple of the maximum set length, directly use the data segment as the data to be processed that needs audio optimization. When the length of the audio frame is not a multiple of the maximum set length, based on the data segment and the adjacent data segments of the data segment, determine the data to be processed that needs audio optimization.
[0084] Specifically, when the length of the audio frame is a multiple of the maximum set length, it means that when splitting the audio frame according to the maximum set length, data segments with a frame length of 1024 can be obtained. Since the frame length of the audio processed by the player is 1024, therefore, the player can directly process this data segment. At this time, the data segment can be directly used as the data to be processed that needs audio optimization. For example, if the audio frame length is 4096, it can be split into four data segments according to 1024, namely segment 1, segment 2, segment 3, and segment 4. Each segment is used as a data to be processed respectively, and then subsequent processing is performed on each data to be processed. At this time, since there are four data to be processed, it needs to be looped four times.
[0085] When the length of the audio frame is not a multiple of the maximum set length, it means that splitting the audio frame according to the maximum set length can obtain multiple data segments with a frame length of 1024 each, and a data segment with a frame length less than 1024. For example, if the frame length of an audio frame is 4000, when splitting according to 1024, the audio frame can be split into four data segments. The frame lengths of the first three data segments are all 1024, and the frame length of the last data segment is 928.
[0086] At this time, since the frame length of the last data segment is 928, if it is combined with the first 96 samples of the first data segment in the next audio frame and processed after audio optimization, it will cause the audio of the current frame to be output until the next frame. If the player is playing a video, there will be an audio-video out-of-sync problem. Therefore, to ensure audio-video synchronization, when encountering a data segment with a frame length less than 1024, the audio is output after being processed. Then, for the first frame, each data segment should be regarded as a data to be processed.
[0087] However, for the first data segment of the second audio frame, since the number of sampling points of the last data to be processed in the previous frame is less than 1024, when performing subsequent audio optimization, the processing methods of the threshold gain data buffer, the look-ahead buffer gain, and the write position of the look-ahead buffer limiter writing the look-ahead buffer gain are not the standard 1024 processing methods. To ensure that the subsequent processing methods of the threshold gain data buffer, the look-ahead buffer gain, and the write position of the look-ahead buffer limiter writing the look-ahead buffer gain are restored to the standard 1024 processing method, the first 96 sampling points in the first data segment of the second frame can be combined with the last data segment of the first frame, that is, 928 sampling points, to obtain a data to be processed including 1024 sampling points, so as to ensure that the subsequent write positions in the threshold gain data buffer, the look-ahead buffer limiter, and the look-ahead buffer limiter are restored to the standard 1024 processing method.
[0088] Therefore, after combining the last (including 928 sampling points) data segment of the first frame with the first 96 sampling points in the first data segment of the second frame to form a data to be processed, there are still 928 sampling points in the first data segment of the second frame that have not been processed. At this time, the remaining 928 sampling points in the first data segment of the second frame can be combined with the first 96 sampling points in the second data segment of the second frame to form the next data to be processed, and so on... After combining the remaining 928 sampling points in the third data segment of the second frame with the first 96 sampling points in the fourth data segment of the second frame to form the next data to be processed, there are still 832 sampling points in the fourth data segment of the second frame that have not been processed. At this time, 832 is directly used as a data to be processed.
[0089] The processing of the third frame, the fourth frame, … up to the last frame is the same as that of the second frame. Through the above method, multiple pieces of data to be processed arranged in sequence can be obtained.
[0090] S13. For each piece of the data to be processed, perform delay processing on the data to be processed to obtain delay data corresponding to each piece of the data to be processed, and determine a look-ahead buffer gain value corresponding to the data to be processed.
[0091] In this embodiment, after obtaining multiple pieces of data to be processed arranged in sequence, since the maximum number of data sampling points of the data to be processed is 1024 and can be processed by the player, the data to be processed can be input into the player for processing, and a limiter operation can be performed after selecting a suitable sound effect. The limiter can be a file-based limiter. The process of the file-based limiter includes loudness conversion, threshold detection, gain calculation, gain smoothing, and gain application, etc. If the process of the file-based limiter is applied in the player, problems will occur in the gain smoothing calculation step during the use of the player.
[0092] Specifically, the original formula for file-based limiter gain smoothing is:
[0093]
[0094] Wherein, denotes the attack time, which is determined by the sampling rate and the time . denotes the release time, which is determined by the sampling rate and the time . is the gain smoothing matrix, which can also be called the gain smoothing array. The initial value of this array is all 0, the maximum value of is the total number of audio data sampling points + 1, and the minimum value is 1, is the gain threshold of the audio after threshold calculation, denotes the current sampling point of the gain smoothing array, denotes the previous sampling point of the gain smoothing array.
[0095] The result of combining the smoothed gain of the attack time and the release time and the audio data (Compressed Signal) is as Figure 2 shown.
[0096] Figure 2Indicates the effect of achieving a smoothing effect by changing the attack time, where ReleaseTime = 0s. The role of the attack time is to restore the signal exceeding the threshold to the threshold. It can be seen from Figure 2 that different Attack Times will produce different curves for smoothing. The longer the time in the attack time, the smoother the curve, that is, the less sensitive to the change in decibels (dB), and the longer the time to reach the target dB. Without the attack time, the audio data can reach the predetermined amplitude value immediately. Specifically, refer to the blue line. The longer the attack time, the longer the time for the audio data to reach the predetermined amplitude value. Specifically, refer to the red line and the yellow line.
[0097] Figure 2 In Figure 2 , for the attack time in the file-type limiter, if a time is set in the gain smoothing, the threshold of the sampling points exceeding the threshold loudness will not reach the threshold immediately, but requires a period of buffering. However, if the attack time is set to 0, it can be seen that although all the data will be forced to be set to the threshold, it will actually seriously affect the listening experience after the limiter. For example, if the original data has exceeded the loudness limit, it will directly clip the audio and generate a hissing noise. Therefore, in the embodiments of the present application, the file-type limiter is first improved by keeping the attack time always 0 to ensure true loudness limitation, and at the same time, the logic of look-ahead buffering is used to ensure the problem of poor limiter effect caused by canceling the attack time. The look-ahead buffering can identify the inflection points or mutations of the signal in advance, so as to achieve a smoother transition. The actual principle of the look-ahead buffering is to use data delay in the original audio data and gain smoothing for side-chain limiting (side-chain limiting is a reference limiter that can make the audio produce the effect of reference data, so as to achieve the predetermined effect) to imitate the effect of the attack time. The optimized process is as Figure 3 shown.
[0098] Figure 3 In Figure 3 , after the input audio, for each piece of data to be processed in the audio, the data to be processed is delayed to obtain the delayed data corresponding to each piece of the data to be processed. This delayed data can be called the delayed input signal. In addition, the gain of the data to be processed is calculated to obtain the smoothing gain, and the smoothing gain is delayed to obtain the delayed smoothing gain. The delayed smoothing gain in this embodiment refers to the look-ahead buffering gain value , so as to use the delayed data and the look-ahead buffering gain value to optimize the audio according to different frame length logics, and then output the audio.
[0099] S14. Multiply the delay data corresponding to the data to be processed by the look-ahead buffer gain value, and use the product as the processing result corresponding to the data to be processed.
[0100] In this embodiment, for each piece of data to be processed, the delay data with a maximum sampling point of 1024 corresponding to the data to be processed is multiplied by the look-ahead buffer gain value to obtain the final audio signal, and the player terminal is used for audio output operation.
[0101] In this embodiment, if the frame length of the played audio is different from the frame length of the audio processed by the player, the audio stream to be processed will be split into multiple data segments, and the length of the data segments is less than or equal to the maximum set length that the player can process, so that the frame length of the data segments is within the frame length range that the player can process, thereby avoiding the situation that the audio signal after adding sound effects exceeds the processable range of the player due to the excessive frame length, and further avoiding the problem of sound overexposure. In addition, in this application, the product of the delay data corresponding to the data to be processed and the look-ahead buffer gain value is used as the processing result corresponding to the data to be processed, that is, the data to be processed is gain-adjusted through the look-ahead buffer gain value, avoiding the situation that the audio signal after adding sound effects exceeds the processable range of the player, and further avoiding the problem of sound overexposure.
[0102] In another implementation manner of this application, the specific implementation of "performing delay processing on the data to be processed to obtain the delay data corresponding to each piece of data to be processed" is given. Refer to Figure 4 and it may include:
[0103] S21. Obtain the target data to be processed.
[0104] Among them, the target data to be processed is initially the first piece of data to be processed among the data to be processed, that is, in the arrangement order of the data to be processed. Therefore, delay operations are performed on each piece of data to be processed.
[0105] S22. Obtain the padding data set in the delay buffer area.
[0106] In this embodiment, the delay buffer area is initially configured with multiple data with padding values being set values, where the set value can be 0, that is the initial value is [0, 0, 0..., 0].
[0107] Set the initial length of the delay to be , which represents the audio frame length. Usually, the length of one frame of audio is set There are 1024 sampling points. 1024 sampling points are the number of sampling points in one frame of the most common AAC (Advanced Audio Coding, a lossy audio compression format), and also the number of sampling points in one frame of the 3D audio technology Audio vivid coding format, and also the common multiple of 4096, which is the number of sampling points in one frame of the audio lossless format FLAC (Free Lossless Audio Codec, lossless audio compression coding). Therefore, set the length of one frame of audio to be 1024 sampling points, which basically covers more than 95% of the playback service requirements.
[0108] In this embodiment, set the delay time to be 0.005 seconds, that is, 5 milliseconds. 5 milliseconds is a relatively small time value that can have a differential effect on the signal. A short delay will not cause the problem of audio-visual asynchronization, and at the same time, it can effectively improve the file limiter.
[0109] In addition, the number of delay sampling points is determined by multiplying the delay time (5 milliseconds) by the sampling rate. Here, the common audio sampling rate is taken as 44100 hz, so set the number of delay sampling points to be 5 milliseconds multiplied by 44100 = 220, that is the number of 0s in the initial value [0, 0, 0……, 0] is 220.
[0110] S23. Add the padding data to the front of the target data to be processed to obtain combined data.
[0111] In this embodiment, since the number of 0s in the initial value [0, 0, 0……, 0] is 220. If it is illustrated with 220 0s as an example, it will lead to a large amount of content. Therefore, in the embodiments of this application, it is illustrated with 4 0s included as an example, and the processing logic for the rest
[0112] including 220 0s is the same. Similarly, since the frame length of a data to be processed is at most 1024, if it is illustrated with 1024 as an example, it will also lead to a large amount of content for the example. Therefore, in this embodiment, the frame length of the data to be processed is taken as 7 for illustration to simplify the content of the illustration.
[0113] When combining the data to be processed with for illustration, if the frame length requirement is 7, for example, the first-frame data to be processed is [1, 2, 3, 4, 5, 6, 7], and it is [0, 0, 0, 0], then The data is placed before the data to be processed to combine the two. After combination, it is [0,0,0,0,1,2,3,4,5,6,7].
[0114] S24. In the combined data, the part with the same length as the target data to be processed is filtered out in order from front to back, and used as the delayed data corresponding to the data to be processed. The remaining part of the combined data is used as the updated filling data, and the updated filling data is updated into the delay buffer area.
[0115] In this embodiment, the delayed data of the data to be processed output after the delay is [0,0,0,0,1,2,3], and the remaining data is saved in ,Right now After updating, it becomes [4,5,6,7].
[0116] S25, determining whether the delayed data corresponding to each of the to-be-processed data is obtained; if so, the process ends; if not, executing step S26.
[0117] S26: taking the next data to be processed after the target data to be processed as new target data to be processed.
[0118] After executing step S26, the process returns to executing step S21, and the process is executed sequentially until the delayed data corresponding to each of the data to be processed is obtained.
[0119] Specifically, if the delayed data corresponding to each data to be processed is obtained, it means that the delayed data is determined and the process can be ended. If the delayed data corresponding to each data to be processed is not obtained, the delayed data corresponding to the next data to be processed should be determined. Specifically, the second frame of data to be processed is [8,9,10,11,12,13,14], [4,5,6,7], the delayed data output after the delay corresponding to the second frame of data to be processed is [4,5,6,7,8,9,10], Update to [11,12,13,14]. The third frame is processed in the same way. and , Output , Update and retain When processing the last frame of data to be processed, the last digit of the data to be processed is Data with the same number of bits will not be used directly to ensure that the total number of sampling points of input and output is consistent.
[0120] like Figure 5 As shown, in the input of the first frame and the second frame, 220 represents , the subsequent numbers represent the data to be processed. The numbers before 220 in the outputs of the first and second frames are the latency data corresponding to the data to be processed in the input, or when 220 does not exist, this output is the latency data corresponding to the data to be processed.
[0121] In this embodiment, by performing latency processing on the audio, the inflection points or mutations of the signal can be identified in advance through look-ahead buffering, so as to achieve a smoother transition.
[0122] In another implementation manner of this application, the implementation process of "determining the look-ahead buffer gain value corresponding to the data to be processed" is given. Refer to Figure 6 , and it may include:
[0123] S31. Calculate the loudness value of the data to be processed.
[0124] In this embodiment, a frame of data to be processed is obtained, and its loudness value is calculated , and the calculation formula of the loudness value is , where is the set of all sampling points of a frame of data to be processed.
[0125] S32. Perform a correction operation on the loudness value to obtain the corrected value corresponding to the loudness value.
[0126] In this embodiment, for the loudness value of the audio sampling points, in order to improve the user's auditory experience, generally the loudness value should not be too large, such as it should not exceed the threshold . Among them, generally 0 dB is default as the maximum loudness, then the threshold can be set to -0.1 dB, which can not only ensure that the audio value does not exceed the boundary, but also does not affect the actual effect of the player's sound effect.
[0127] Therefore, after the loudness value exceeds the threshold , the loudness value of the audio sampling point will be assigned to , that is, the corrected value corresponding to the loudness value of this sampling point is , if it does not exceed the threshold , then the loudness of the audio sampling point remains unchanged, and the corrected value corresponding to the loudness value is still this actual loudness value.
[0128] S33. Use the loudness value and the corrected value to calculate the gain threshold.
[0129] In this embodiment, among them, the gain threshold .
[0130] S34. Use the gain smoothing formula and the gain threshold to obtain the gain smoothing matrix corresponding to the data to be processed.
[0131] Among them, according to the above discussion, the attack time needs to be set to 0. Then, the attack time used when determining the gain smoothing formula is the set value 0. At this time = . When setting the attack time to 0 seconds, the in the release time can be set to 0.2 seconds based on requirements.
[0132] In addition, since the input of the original file-based limiter is the entire file and the sampling points in the file are continuous, but this method cannot meet the requirements of the player for reading and playing frame by frame. Therefore, in the embodiments of the present application, the audio stream is processed frame by frame, but this will result in no continuity between adjacent frames, thus not meeting the requirement of the file-based limiter for continuous sampling points. Therefore, in this embodiment, continuity between the previous data to be processed and the next data to be processed will be constructed to make the sampling points continuous. In practical applications, the threshold gain data buffer will be set to , initially [0, 0]. The last two values of the gain threshold of each data to be processed are saved in the threshold gain data buffer for use as the initial gain threshold reference for the next data to be processed. Subsequently it will be continuously updated according to the processing of the data to be processed. Initially, in the gain smoothing formula, the gain value of the first sampling point is the gain value of the last sampling point of the previous data to be processed of the data to be processed; the gain value of the last sampling point is the gain value of the first sampling point of the next data to be processed of the data to be processed.
[0133] Therefore, the improved limiter frame-level gain smoothing formula in this embodiment is:
[0134]
[0135] Among them, is the gain smoothing matrix of the left channel, is the gain smoothing matrix of the right channel. When it is the first frame of data to be processed, since there is no previous frame reference, so and the first sampling points of both are assigned the gain threshold reference values of the left and right channels of . After the gain smoothing calculation of the first data to be processed is completed, and are the last values of the left and right channels, which are assigned to to ensure the continuity of the gain threshold reference value of the first sampling point of the next frame of data to be processed.
[0136] The principle of processing according to the above logic is as follows: If a frame of data to be processed is processed in a file-based manner, there is a set of true gain thresholds for the left and right channels at the end of the actual frame of data to be processed. However, due to the logic of the player, if these true gain thresholds for the left and right channels are not reserved in the buffer, then when creating the second frame of data to be processed, they will all be assigned a value of 0 again, that is, it cannot correspond to the file-based result, that is, it cannot ensure the correctness of the effect of the player's frame-level limiter. Therefore, the frame-level gain smoothing of the limiter is improved as follows.
[0137] S35. Perform look-ahead buffer delay processing on the gain smoothing matrix corresponding to the data to be processed to obtain the look-ahead buffer gain value corresponding to the data to be processed.
[0138] In this embodiment, after calculating the gain smoothing matrix for the left and right channels, look-ahead buffer delay processing is started on the gain smoothing matrix. The overall process of look-ahead buffer delay processing is divided into writing the gain smoothing matrix into the look-ahead buffer, using the look-ahead buffer to delay and process the gain smoothing matrix, and reading the processed gain smoothing matrix from the buffer.
[0139] In one implementation, step S35 may include:
[0140] 1) Write the gain smoothing matrix corresponding to the data to be processed into the look-ahead buffer to obtain the look-ahead buffer gain and the write position of the gain smoothing matrix.
[0141] In this embodiment, set the look-ahead buffer gain The initial length of the array is + , the initial filling value is 0, and set the initial write position of the look-ahead buffer limiter write to 0.
[0142] In one implementation, perform look-ahead buffer write operations on , , , and . The steps are as follows:
[0143] Calculate , and these three values respectively represent the write position, the length of the first block, and the length of the second segment. The first block refers to the block where the current data to be processed is located in the look-ahead buffer, and the second block refers to the block where the previous data to be processed is located in the look-ahead buffer.
[0144] These three values are jointly determined by , and , and the formula is:
[0145]
[0146] Among them, 、 are intermediate parameter values, takes a value of , takes a value of , is the length of + . is a calculation method to find the smaller value of two values. is a ternary operator, which is represented here as , takes a value of 0, otherwise takes a value of .
[0147] After obtaining , and can be written into , and are combined into , and the formula is as follows:
[0148]
[0149] At the same time, update , and the formula is as follows:
[0150]
[0151] The writing logic is the logic of a standard circular buffer queue. By calculating the writing position and the size of the writing block, an efficient writing operation of the circular buffer is realized. This design makes full use of the circular characteristics of the circular buffer and avoids the problem of buffer overflow.
[0152] 2) Dynamically adjust the look-ahead buffer gain to obtain an adjusted look-ahead buffer gain value.
[0153] In this embodiment, the overall logic of the dynamic gain adjustment is to obtain the sample value at the current index, which specifically refers to the look-ahead buffer gain. If the look-ahead buffer gain is greater than the current gain attenuation reference value, limit the look-ahead buffer gain to the gain attenuation reference value to avoid excessive audio volume, and then update the gain attenuation reference value. If the look-ahead buffer gain is less than or equal to the gain attenuation reference value, calculate a new step size and update the gain attenuation reference value.
[0154] In one implementation, the look-ahead buffer gain corresponding to the index value can be obtained. When the obtained look-ahead buffer gain is greater than the gain attenuation reference value, the look-ahead buffer gain corresponding to the index value in the look-ahead buffer is set to the gain attenuation reference value, and the sum of the step size and the gain attenuation reference value is used as the new gain attenuation reference value. When the obtained look-ahead buffer gain is less than the gain attenuation reference value, a new step size is calculated using the obtained look-ahead buffer gain and the initial length in the delay buffer region, and the sum of the new step size and the gain attenuation reference value is used as the new gain attenuation reference value.
[0155] Specifically, when implementing, set the gain attenuation reference value , the step size and the index with initial values of 0, 0, and -1 respectively.
[0156] Then, calculate Partition 1 ( and ) and Partition 2 ( ) according to ), and The calculation formulas are as follows:
[0157]
[0158] wherein, and serve to divide the processing tasks of the look-ahead buffer into two parts: the partition where the current data to be processed is located in the look-ahead buffer and the partition where the previous data to be processed is located in the look-ahead buffer, so as to process the samples crossing the buffer boundary separately. This block processing method ensures the continuity of the circular buffer, avoids index out-of-bounds problems, and can efficiently process all samples that need to be processed.
[0159] Subsequently, start the segmented processing of for and . First, process the part of , and circularly fetch data, specifically referring to the look-ahead buffer gain corresponding to the index value . Compare with in terms of size. If the current data is larger than , perform numerical update using the following formula
[0160]
[0161] If it is less, perform numerical update using the following formula:
[0162]
[0163] It should be noted that for the meanings of the letters in this embodiment, please refer to the corresponding descriptions above.
[0164] After processing the index is incremented by one, the index is decremented by one. After the processing ends, start processing At this time the initial value of is the length of minus 1. The logical judgment and processing method are the same as the
[0165] In this embodiment, through dynamic gain adjustment, the sample values in the look-ahead buffer are restricted to prevent the peak value of the audio signal from being too large, and the step size is dynamically adjusted according to the sample values and the delay time to adapt to different audio signal characteristics, so as to achieve smoother gain adjustment.
[0166] 3) Read the adjusted look-ahead buffer gain value from the look-ahead buffer to obtain the look-ahead buffer gain value corresponding to the data to be processed.
[0167] In this embodiment, the process of reading the processed gain data from the buffer is similar to the process of writing the gain data into the look-ahead buffer. First, the initialization assignment is modified, and the formula is as follows:
[0168]
[0169] It should be noted that for the meanings of the letters in this embodiment, please refer to the corresponding descriptions above. The purpose of modifying the initialization assignment of is to calculate the start position and the block size of the read operation. It needs to consider the delay of the look-ahead buffer ( ), and the number of samples pushed last time ( ), so
[0170] is the processed look-ahead buffer gain value read, and the formula is:
[0171]
[0172] Thus, the gain smoothing data processed by the look-ahead buffer is obtained, so that it has the data smoothing ability when the attack time is greater than 0 and can ensure that the loudness can be limited within the set maximum threshold.
[0173] To enable those skilled in the art to more clearly understand this application, an example of buffer processing for audio frames of different lengths in this application will now be given.
[0174] Due to the existence of audio with different frame lengths, after the original signal and the gain signal are delayed, the buffer update method itself should be divided into three types. Several situations are exemplified as follows:
[0175] Embodiment 1: The length of the audio frame is an integer multiple of 1024
[0176] For example, if the audio frame length is 4096, it can be processed four times in a logical loop of 1024, which are called Segment 1, Segment 2, Segment 3, and Segment 4. Each segment is processed according to the above logic, and the delayed data is multiplied by the look-ahead buffer gain value to obtain the final audio signal for output at the player terminal. Each time, it is necessary to update and retain , , and , and output and of each segment. This process continues until the entire audio ends. If the last frame is less than 1024, the processing method refers to Embodiment 2.
[0177] Embodiment 2: The length of the audio frame is not an integer multiple of 1024 and it is not the last frame
[0178] For example, if the audio frame length is 4000, then the first three segments are processed according to the logic of 1024, and , , and are updated and retained. The actual audio length entering the fourth time is only 928. It is still calculated according to the above logic. At this time, , and of Segment 4 are not retained, and the actual audio length entering the fourth time is all put into . At this time, should be the normally retained in the third time plus all the audio sampling points of the fourth frame, and and are output. To maintain audio-video synchronization, the result of the fourth output signal is directly output at the terminal. When the next frame of audio is obtained and transmitted, the first 96 sampling points of Segment 1 are combined with of the first frame to form the first data to be processed in the real second frame. At this time, , , and , at this time, the output signal of the buffer with 96 points will not be output by the player and will be retained in the playback buffer. It will be played after the entire first segment is processed. The second input is Adding the last 928 sampling points of segment 1 and the first 96 sampling points of segment 2 to form a 220 + 1024 pattern. Among the 1024 output, the first 928 sampling points are combined with the output signal of the 96 points in the player buffer and output to the terminal. The last 96 sampling point signals continue to enter the player buffer. The third input-output operation is the same as the second until the last time. The length of the last segment 4 is 220 + 832, and actually 96 + 832 points are output.
[0179] The first four times will all be updated , , and , and the fifth time will only update . The processing output result of segment 1 is output starting from the second time, and the audio processing result of the second frame is completely processed and output by the fifth time. The processing method of the third frame is the same as that of the second frame, and so on for the subsequent frames' operation logic.
[0180] In this embodiment, since the player needs to maintain audio-visual synchronization, when encountering different audio frame lengths, the audio is output after being processed. If the audio waits until the next frame to be output in this frame, there will be a problem of audio-visual asynchronization.
[0181] In addition, for the last time or the last segment processing of each frame, only is updated. The reason is: , and are not the processing methods of the standard 1024. If the buffer data is retained, then for example , the first 928 points will be the processing result of segment 4, and the last 96 points will be the processing result of the previous segment 3. When the next frame of audio comes in, obviously the continuity of this data buffer is broken, resulting in incorrect subsequent references and causing the limiter's processing to fail. And only retaining can meet the standard 1024 processing logic, enabling , and to be accurately updated. The specific implementation process is referred to Figure 5 as shown.
[0182] Embodiment 3: If the length of the audio frame is not an integer multiple of 1024 and it is the last frame situation.
[0183] For example, if the length of the last audio frame is 1000, since there are no subsequent frames for reference, it can be calculated directly according to the normal logic; for example, if the length of the last audio frame is 4000, it is calculated according to the above logic and the output ends.
[0184] Through the above improvements, the second improvement of the original limiter is achieved to support audio and video data with different frame lengths. Thus, the improvement of the file-based limiter is completed, which can apply to the problem of audio over-explosion that may be caused by different audio and video frame lengths and sound effects applications of the player, control the audio loudness within the standard range, and further ensure the user's audio-visual experience.
[0185] After performing the look-ahead buffer delay processing, the delay data with a length of 1024 obtained from each normal processing is multiplied by the look-ahead buffer gain value to obtain the final audio signal for output at the player terminal.
[0186] In this embodiment, to ensure that the sound effects are played normally in videos and audios with different frame lengths and adapt to their characteristics, this embodiment proposes a look-ahead buffer limiter solution based on real-time processing of player sound effects. By adopting a real-time limiter algorithm with arbitrary-sized frame input to meet the video and audio inputs and sound effects applications with different frame lengths, and in addition, adopting a look-ahead buffer logic to optimize the effect of the limiter algorithm, enhance the audio listening experience after the limiter's action, and thus be able to adapt to the video and audio inputs and sound effects applications with different frame lengths, and optimize the limiter algorithm logic through the look-ahead buffer technology to ensure that the audio listening experience after the limiter's processing is not lost and prevent audio distortion.
[0187] Based on the above embodiment of the audio processing method, another embodiment of the present application provides an audio processing device, referring to Figure 7 , which may include:
[0188] An audio stream processing module 11, configured to split the audio stream to be processed into multiple data segments; the length of the data segment is less than or equal to the maximum set length that the player can process;
[0189] A data determination module 12, configured to determine the data to be processed that needs to be audio-optimized corresponding to the data segment;
[0190] A data processing module 13, configured to perform delay processing on each piece of the data to be processed to obtain the delay data corresponding to each piece of the data to be processed, and determine the look-ahead buffer gain value corresponding to the data to be processed;
[0191] An audio optimization module 14, configured to use the product of the delay data corresponding to the data to be processed and the look-ahead buffer gain value as the processing result corresponding to the data to be processed.
[0192] In one implementation, when the audio stream processing module 11 is used to split the audio stream to be processed into multiple data segments, it is specifically used for:
[0193] Performing a parsing operation on the audio stream to be processed to obtain each audio frame in the audio stream, and splitting the audio frames according to the maximum set length to obtain multiple data segments.
[0194] In one implementation, the data determination module 12 includes:
[0195] A first determination sub-module, configured to directly use the data segment as the data to be processed that needs to be audio-optimized when the length of the audio frame is a multiple of the maximum set length;
[0196] A second determination sub-module, configured to, when the length of the audio frame is not a multiple of the maximum set length, determine the data to be processed that needs to be audio-optimized based on the data segment and the adjacent data segment of the data segment.
[0197] In one implementation, the data processing module 13 includes:
[0198] A first data acquisition sub-module, configured to acquire target data to be processed, where the target data to be processed is initially the first data to be processed in the data to be processed;
[0199] A second data acquisition sub-module, configured to acquire padding data set in a delay buffer area, where the delay buffer area is initially configured with multiple data with a padding value of a set value;
[0200] A combination sub-module, configured to add the padding data to the front of the target data to be processed to obtain combined data;
[0201] A delay sub-module, configured to, in the combined data, screen out the part with the same length as the target data to be processed in the order from front to back as the delay data corresponding to the data to be processed, use the remaining part in the combined data as the updated padding data, and update the updated padding data to the delay buffer area;
[0202] A judgment sub-module, configured to judge whether the delay data corresponding to each data to be processed is obtained;
[0203] A data determination sub-module, configured to, after the judgment sub-module determines that the delay data corresponding to each data to be processed is not obtained, use the next data to be processed after the target data to be processed as the new target data to be processed;
[0204] The first data acquisition sub-module is used to acquire the target data to be processed after the data determination sub-module uses the next data to be processed after the target data to be processed as the new target data to be processed, until the delay data corresponding to each piece of data to be processed is obtained.
[0205] In one implementation, the data processing module 13 includes:
[0206] A loudness calculation sub-module, which is used to calculate the loudness value of the data to be processed;
[0207] A correction sub-module, which is used to perform a correction operation on the loudness value to obtain a corrected value corresponding to the loudness value;
[0208] A threshold determination sub-module, which is used to calculate a gain threshold by using the loudness value and the corrected value;
[0209] A matrix determination sub-module, which is used to obtain a gain smoothing matrix corresponding to the data to be processed by using a gain smoothing formula and the gain threshold; the attack time used when determining the gain smoothing formula is a set value. In the gain smoothing formula, the gain value of the first sampling point is the gain value of the last sampling point of the previous data to be processed of the data to be processed; the gain value of the last sampling point is the gain value of the first sampling point of the next data to be processed of the data to be processed;
[0210] A delay processing sub-module, which is used to perform a look-ahead buffer delay processing on the gain smoothing matrix corresponding to the data to be processed to obtain a look-ahead buffer gain value corresponding to the data to be processed.
[0211] In one implementation, the delay processing sub-module includes:
[0212] A writing unit, which is used to write the gain smoothing matrix corresponding to the data to be processed into a look-ahead buffer to obtain a look-ahead buffer gain and the writing position of the gain smoothing matrix;
[0213] An adjustment unit, which is used to perform a dynamic gain adjustment on the look-ahead buffer gain to obtain an adjusted look-ahead buffer gain value;
[0214] A reading unit, which is used to read the adjusted look-ahead buffer gain value from the look-ahead buffer to obtain a look-ahead buffer gain value corresponding to the data to be processed.
[0215] In one implementation, the adjustment unit includes:
[0216] A gain acquisition sub-unit, which is used to acquire the look-ahead buffer gain corresponding to an index value;
[0217] The first adjustment subunit is configured to, when the obtained look-ahead buffer gain is greater than the gain attenuation reference value, set the look-ahead buffer gain corresponding to the index value in the look-ahead buffer to the gain attenuation reference value, and use the sum of the step size and the gain attenuation reference value as the new gain attenuation reference value;
[0218] The second adjustment subunit is configured to, when the obtained look-ahead buffer gain is less than the gain attenuation reference value, calculate a new step size by using the obtained look-ahead buffer gain and the initial length in the delay buffer area, and use the sum of the new step size and the gain attenuation reference value as the new gain attenuation reference value.
[0219] In this embodiment, if the frame length of the played audio is different from the frame length of the audio processed by the player, the audio stream to be processed will be split into multiple data segments, and the length of the data segment is less than or equal to the maximum set length that the player can process, so that the frame length of the data segment is within the frame length range that the player can process, thereby avoiding the situation that the audio signal after adding the sound effect exceeds the range that the player can process due to the excessive frame length, and further avoiding the problem of sound overexposure. In addition, in this application, the product of the delay data corresponding to the data to be processed and the look-ahead buffer gain value is used as the processing result corresponding to the data to be processed, that is, the look-ahead buffer gain value is used to perform gain adjustment on the data to be processed, avoiding the situation that the audio signal after adding the sound effect exceeds the range that the player can process, and further avoiding the problem of sound overexposure.
[0220] It should be noted that for the working processes of the various modules, sub-modules, units, and sub-units in this embodiment, please refer to the corresponding descriptions in the above embodiments, and will not be elaborated here.
[0221] An electronic device is further provided in an embodiment of the present application, including at least one processor and a memory connected to the processor, where:
[0222] The memory is used to store a computer program;
[0223] The processor is configured to execute the computer program so that the electronic device can implement the above audio processing method.
[0224] Reference Figure 8 As shown, it shows a schematic structural diagram of an electronic device suitable for implementing the electronic device in an embodiment of the present application. The electronic device in an embodiment of the present application may include, but is not limited to, fixed terminals such as mobile phones, laptop computers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), desktop computers, and the like. Figure 8 The shown electronic device is only an example and should not bring any limitation to the functions and usage scopes of the embodiments of the present application.
[0225] As Figure 8As shown, the electronic device may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 601, which may perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 602 or a program loaded from a storage device 608 into a random access memory (RAM) 603. When the electronic device is powered on, various programs and data required for the operation of the electronic device are also stored in the RAM 603. The processing device 601, the ROM 602, and the RAM 603 are connected to each other through a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.
[0226] Generally, the following devices may be connected to the I / O interface 605: an input device 606 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 607 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 608 including, for example, a memory card, a hard disk, etc.; and a communication device 609. The communication device 609 may allow the electronic device to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 8 an electronic device with various devices is shown, it should be understood that it is not required to implement or have all the shown devices. More or fewer devices may be implemented or had alternatively.
[0227] An embodiment of the present application also provides a computer program product including computer-readable instructions, which, when running on an electronic device, enable the electronic device to implement any one of the audio processing methods provided by the embodiments of the present application.
[0228] An embodiment of the present application also provides a computer-readable storage medium carrying one or more computer programs, which, when executed by an electronic device, can enable the electronic device to implement any one of the audio processing methods provided by the embodiments of the present application.
[0229] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present application. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to these embodiments shown herein, but rather to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. An audio processing method, characterized in that: include: Splitting the audio stream to be processed into multiple data segments; the length of the data segments is less than or equal to the maximum set length that the player can process; Determining the data to be processed corresponding to the data segment that needs to be audio optimized; For each of the data to be processed, delay processing is performed on the data to be processed to obtain delayed data corresponding to each of the data to be processed, and a forward-looking buffer gain value corresponding to the data to be processed is determined; The product of the delayed data corresponding to the data to be processed and the forward-looking buffer gain value is taken as the processing result corresponding to the data to be processed.
2. The audio processing method according to claim 1, characterized in that: Split the audio stream to be processed into multiple data segments, including: Performing a parsing operation on the audio stream to be processed to obtain each audio frame in the audio stream; The audio frame is segmented according to a maximum set length to obtain multiple data segments.
3. The audio processing method according to claim 2, characterized in that: Determining the data to be processed corresponding to the data segment and requiring audio optimization, including: When the length of the audio frame is a multiple of the maximum set length, the data segment is directly used as the data to be processed that needs to be optimized for audio; When the length of the audio frame is not a multiple of the maximum set length, the to-be-processed data that needs to be audio optimized is determined based on the data segment and adjacent data segments of the data segment.
4. The audio processing method according to claim 1, characterized in that: Delay processing is performed on the data to be processed to obtain delayed data corresponding to each of the data to be processed, including: Acquire target data to be processed, where the target data to be processed is initially the first data to be processed among the data to be processed; Acquire filling data set in the delay buffer area, wherein the delay buffer area is initially configured with a plurality of data having filling values of set values; Add the padding data to the front of the target data to be processed to obtain combined data; In the combined data, the portion having the same length as the target data to be processed is selected in a forward-to-backward order and used as the delayed data corresponding to the data to be processed, and the remaining portion in the combined data is used as the updated padding data, and the updated padding data is updated into the delay buffer area; The next data to be processed after the target data to be processed is taken as the new target data to be processed, and the step of obtaining the target data to be processed is returned to and executed sequentially until the delayed data corresponding to each of the data to be processed is obtained.
5. The audio processing method according to claim 1, characterized in that: Determining a forward buffer gain value corresponding to the data to be processed includes: Calculating the loudness value of the data to be processed; Performing a correction operation on the loudness value to obtain a correction value corresponding to the loudness value; Calculating a gain threshold using the loudness value and the correction value; The gain smoothing matrix corresponding to the data to be processed is obtained by using the gain smoothing formula and the gain threshold; the attack time used when determining the gain smoothing formula is the set value, and in the gain smoothing formula, the gain value of the first sampling point is the gain value of the last sampling point of the previous data to be processed of the data to be processed; the gain value of the last sampling point is the gain value of the first sampling point of the next data to be processed of the data to be processed; Performing forward buffer delay processing on the gain smoothing matrix corresponding to the data to be processed to obtain a forward buffer gain value corresponding to the data to be processed.
6. The audio processing method according to claim 5, characterized in that: Performing forward buffer delay processing on the gain smoothing matrix corresponding to the data to be processed to obtain a forward buffer gain value corresponding to the data to be processed, including: Writing the gain smoothing matrix corresponding to the data to be processed into the forward buffer to obtain the forward buffer gain and the writing position of the gain smoothing matrix; Performing dynamic gain adjustment on the forward-looking buffer gain to obtain an adjusted forward-looking buffer gain value; The adjusted forward buffer gain value is read from the forward buffer to obtain the forward buffer gain value corresponding to the data to be processed.
7. The audio processing method according to claim 6, characterized in that: Dynamically adjusting the forward-looking buffer gain to obtain an adjusted forward-looking buffer gain value includes: Get the forward buffer gain corresponding to the index value; In the case where the acquired forward buffer gain is greater than the gain attenuation reference value, setting the forward buffer gain corresponding to the index value in the forward buffer to the gain attenuation reference value, and taking the sum of the step size and the gain attenuation reference value as a new gain attenuation reference value; When the acquired look-ahead buffer gain is less than the gain attenuation reference value, a new step length is calculated using the acquired look-ahead buffer gain and the initial length in the delay buffer area, and the sum of the new step length and the gain attenuation reference value is used as a new gain attenuation reference value.
8. An audio processing device, characterized in that: include: An audio stream processing module, used for splitting the audio stream to be processed into multiple data segments; the length of the data segment is less than or equal to the maximum set length that the player can process; A data determination module, used to determine the to-be-processed data that needs to be audio optimized corresponding to the data segment; A data processing module, configured to perform delay processing on each of the data to be processed, obtain delayed data corresponding to each of the data to be processed, and determine a forward buffer gain value corresponding to the data to be processed; The audio optimization module is used to take the product of the delay data corresponding to the data to be processed and the forward-looking buffer gain value as the processing result corresponding to the data to be processed.
9. The audio processing device according to claim 8, characterized in that: The data processing module comprises: A first data acquisition submodule is used to acquire target data to be processed, wherein the target data to be processed is initially the first data to be processed among the data to be processed; A second data acquisition submodule is used to acquire filling data set in the delay buffer area, wherein the delay buffer area is initially configured with a plurality of data having filling values of set values; A combining submodule, used for adding the filling data to the front of the target data to be processed to obtain combined data; a delay submodule, configured to filter out, from the combined data, a portion having the same length as the target data to be processed in a forward-to-backward order and use the portion as delayed data corresponding to the data to be processed, use the remaining portion of the combined data as updated padding data, and update the updated padding data into the delay buffer area; A judging submodule, used for judging whether delayed data corresponding to each of the data to be processed is obtained; A data determination submodule, configured to use the next data to be processed after the target data to be processed as a new target data to be processed after the determination submodule determines that the delayed data corresponding to each of the data to be processed is not obtained; The first data acquisition submodule is used for the data determination submodule to take the next to-be-processed data after the target to-be-processed data as the new target to-be-processed data, and then acquire the target to-be-processed data until the delayed data corresponding to each of the to-be-processed data is obtained.
10. An electronic device, characterized in that: The method comprises at least one processor and a memory connected to the processor, wherein: The memory is used to store computer programs; The processor is used to execute the computer program so that the electronic device can implement the audio processing method as described in any one of claims 1 to 7.