Audio data encoding method, decoding method, and apparatus
The method addresses audio discontinuity and noise issues in VOIP calls by processing low-frequency data from consecutive audio frames based on their encoding delays when switching between MDC and SDC modes, enhancing the quality of the audio signal.
Patent Information
- Application Number
- JP2024572349
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-11-07
- Filing Date
- 2023-11-03
- Publication Date
- 2025-06-26
- Estimated Expiration
- 2043-11-03
AI Technical Summary
Existing audio encoding methods for VOIP calls face issues with audio discontinuity and noise when switching between Multiple Description Coding (MDC) and Single Description Coding (SDC) modes due to mismatched parameters such as delay and sampling rate.
The proposed method determines the encoding mode of consecutive audio frames and generates target data by processing low-frequency data from both frames based on their encoding delays when a mode switch occurs, ensuring seamless transition and minimizing discontinuity.
This approach effectively reduces audio discontinuity and noise when switching encoding modes, thereby improving the overall quality of the audio signal in VOIP calls.
Smart Images

Figure 2025519557000001_ABST
Abstract
Description
Technical Field
[0001] This application claims priority based on a Chinese application with an application number of 202211387602.8 and a filing date of November 7, 2022, and the disclosure content of this Chinese application is incorporated herein by reference in its entirety.
[0002] This disclosure relates to the field of data processing technology, and particularly to an encoding method, a decoding method, and an apparatus for audio data.
Background Art
[0003] In a VOIP (Voice over Internet Protocol) call, in order to improve the quality of an audio signal, an encoding end adjusts an encoding mode according to a real-time network situation, for example, switches between a Multiple Description Coding (MDC) mode and a Single Description Coding (SDC) mode.
[0004] Since the MDC mode and the SDC mode use different encoding algorithms, parameters such as delay and sampling rate may not match, and problems such as audio discontinuity and / or noise may occur when decoding audio data when switching the encoding mode.
Summary of the Invention
Problems to be Solved by the Invention
[0005] In view of this, embodiments of the present disclosure provide an encoding method, a decoding method, a processing method, and an apparatus for audio data to improve the quality of an audio signal when switching an encoding mode.
Means for Solving the Problems
[0006] To achieve the above object, embodiments of the present disclosure provide the following technical solutions.
[0007] According to a first aspect, embodiments of the present disclosure provide a method for encoding audio data, the method comprising: determining an encoding mode of a first audio frame; determining whether the encoding mode of the first audio frame is the same as that of a second audio frame, where the second audio frame is the audio frame preceding the first audio frame; when they are not the same and the encoding mode of the first audio frame is multi-description coding, generating third data based on first data, second data, and a first delay, where the first data is low-frequency data obtained by downsampling the original audio data of the first audio frame, the second data is low-frequency data obtained by downsampling the original audio data of the second audio frame, and the first delay is the encoding delay of the multi-description coding; multiplying the third data by the multi-description coding to obtain encoded data of the first audio frame.
[0008] In an alternative embodiment of the embodiments of the present application, the method further comprises: when the encoding mode of the first audio frame is not the same as that of the second audio frame and the encoding mode of the first audio frame is single-description coding, generating sixth data based on fourth data, fifth data, and a second delay, where the fourth data is the original audio data of the first audio frame, the fifth data is the original audio data of the second audio frame, and the second delay is the encoding delay of the single-description coding; encoding the sixth data by the single-description coding to obtain encoded data of the first audio frame.
[0009] As an alternative embodiment of the example of the present application, generating the third data based on the first data, the second data, and the first delay includes: obtaining seventh data by cutting off sample points having a length of the first delay from the end of the second data; obtaining eighth data by connecting the seventh data to the front end of the first data; and obtaining the third data by deleting sample points having a length of the first delay from the end of the eighth data.
[0010] As an alternative embodiment of the example of the present application, generating the sixth data based on the fourth data, the fifth data, and the second delay includes: obtaining ninth data by cutting off sample points having a length of the second delay from the end of the fifth data; obtaining tenth data by connecting the ninth data to the front end of the fourth data; and obtaining the sixth data by deleting sample points having a length of the second delay from the end of the tenth data.
[0011] As an alternative embodiment of the example of the present application, determining the encoding mode of the first audio frame includes: determining whether an encoding mode switching condition is satisfied based on an encoding mode duration length and a signal type of the first audio frame, where the encoding mode duration length is a playback time length of audio frames continuously encoded in a current encoding mode; if NO, determining the encoding mode of the second audio frame as the encoding mode of the first audio frame; if YES, determining the encoding mode of the first audio frame based on network parameters of an audio encoding data transmission network.
[0012] As an alternative embodiment of the example of the present application, determining whether to satisfy the encoding mode switching condition based on the encoding mode duration and the signal type of the first audio frame includes: determining whether the encoding mode duration is greater than a threshold duration; determining whether the probability that the first audio frame is a voice audio frame is less than a threshold probability; when the encoding mode duration is greater than the threshold duration and the probability that the first audio frame is a voice audio frame is less than the threshold probability, determining that the encoding mode switching condition is satisfied; when the encoding mode duration is less than or equal to the threshold duration and / or the probability that the first audio frame is a voice audio frame is greater than or equal to the threshold probability, determining that the encoding mode switching condition is not satisfied.
[0013] As an alternative embodiment of the example of the present application, determining the encoding mode of the first audio frame based on the network parameters of the audio encoding data transmission network includes: determining the packet loss rate of the audio encoding data transmission network based on the network parameters; determining whether the packet loss rate is greater than or equal to a threshold packet loss rate; if YES, determining that the encoding mode of the first audio frame is multi-description encoding; if NO, determining that the encoding mode of the first audio frame is single-description encoding.
[0014] According to a second aspect, an example of the present disclosure provides a method for decoding audio data, the method including: Determining an encoding mode of a first audio frame based on the encoded data of the first audio frame; Decoding the encoded data of the first audio frame based on the encoding mode to obtain decoded data; Determining whether the encoding mode of the first audio frame is the same as the encoding mode of a second audio frame, where the second audio frame is the audio frame before the first audio frame; When they are not the same and the encoding mode of the first audio frame is multi-description encoding, generating packet loss compensation data based on the second audio frame; Smoothing the decoded data based on the delay data of the second audio frame and the packet loss compensation data to obtain playback data of the first audio frame.
[0015] As an alternative embodiment of the embodiment of the present application, the method includes: When the encoding mode of the first audio frame is not the same as the encoding mode of the second audio frame and the encoding mode of the first audio frame is single-description encoding, generating packet loss compensation data based on the second audio frame; Smoothing the decoded data based on the packet loss compensation data to obtain a smoothing result corresponding to the decoded data; Further including delaying the smoothing result based on the packet loss compensation data and the number of delay sample points to obtain playback data of the first audio frame, where the number of delay sample points is the number of delay sample points of multi-description encoding.
[0016] As an alternative embodiment of the embodiment of the present application, the step of smoothing the decoded data based on the packet loss compensation data to obtain a smoothing result corresponding to the decoded data is: Replacing a first sample point sequence in the decrypted data with a second sample point sequence in the packet loss compensation data to obtain a first replacement result, where the first sample point sequence is a sample point sequence composed of the first first several sample points of the decrypted data, the first number is a difference value between a first preset number and the number of delay sample points, and the second sample point sequence is a sample point sequence composed of sample points in the packet loss compensation data where the index value ranges from the number of delay sample points to the first preset number. Windowing and superimposing a third sample point sequence in the first replacement result and a fourth sample point sequence in the packet loss compensation data based on a first window function to obtain a smoothing result corresponding to the decrypted data, where the third sample point sequence is a sample point sequence composed of sample points in the first replacement result where the index value ranges from the first number to the sum of the first number and a second preset number, and the fourth sample point sequence is a sample point sequence composed of sample points in the packet loss compensation data where the index value ranges from the first preset number to the sum of the first preset number and a second preset number.
[0017] As an alternative embodiment of the examples of this application, the delaying the smoothing result based on the packet loss compensation data and the number of delay sample points to obtain the playback data of the first audio frame is Obtaining a fifth sample point sequence, where the fifth sample point sequence is a sample point sequence composed of the first number of delay sample points of the packet loss compensation data. Connect the fifth sample point sequence in front of the smoothing result to obtain a first connection result; Delete the sixth sample point sequence in the first connection result to obtain the playback data of the first audio frame, where the sixth sample point sequence is a sample point sequence composed of sample points of the last delay sample point number in the first connection result.
[0018] As an alternative embodiment of the embodiment of the present application, smoothing the decoded data based on the delay data of the second audio frame and the packet loss compensation data to obtain the playback data of the first audio frame includes: Replacing the seventh sample point sequence in the decoded data with the delay data to obtain a second replacement result, where the seventh sample point sequence is a sample point sequence composed of the first delay sample point number of sample points in the decoded data; Windowing and superimposing the eighth sample point sequence in the second replacement result and the ninth sample point sequence in the packet loss compensation data based on a second window function to obtain the playback data of the first audio frame, where the eighth sample point sequence is a sample point sequence composed of sample points where the index value in the second replacement result ranges from the delay sample point number to the sum of the delay sample point number and a third preset number, and the ninth sample point sequence is a sample point sequence composed of the first third preset number of sample points in the packet loss compensation data.
[0019] As an alternative embodiment of the embodiment of the present application, the method includes: When the encoding mode of the first audio frame is the same as the encoding mode of the second audio frame, and the encoding mode of the first audio frame is single description encoding, the method further includes delaying the decoded data based on the delay data of the second audio frame and the number of delay sample points to obtain playback data of the first audio frame.
[0020] In an alternative embodiment of the example of this application, the delaying the decoded data based on the delay data of the second audio frame and the number of delay sample points to obtain playback data of the first audio frame includes: concatenating the delay data in front of the decoded data to obtain a second concatenation result; deleting a tenth sample point sequence in the second concatenation result to obtain playback data of the first audio frame, where the tenth sample point sequence is a sample point sequence composed of the last number of delay sample points of the second concatenation result.
[0021] According to a third aspect, an embodiment of the present disclosure provides an audio data encoding apparatus, the apparatus includes: a determination unit for determining an encoding mode of a first audio frame; a determination unit for determining whether the encoding mode of the first audio frame is the same as the encoding mode of the second audio frame, where the second audio frame is an audio frame before the first audio frame. When the encoding mode of the first audio frame is not the same as the encoding mode of the second audio frame, and the encoding mode of the first audio frame is multi-description encoding, a generation unit for generating third data based on first data, second data, and a first delay, where the first data is low-frequency data obtained by decimating the original audio data of the first audio frame, the second data is low-frequency data obtained by decimating the original audio data of the second audio frame, and the first delay is the encoding delay of the multi-description encoding. An encoding unit that multi-description encodes the third data to obtain the encoded data of the first audio frame.
[0022] In an alternative embodiment of the example of the present application, when the encoding mode of the first audio frame is not the same as the encoding mode of the second audio frame, and the encoding mode of the first audio frame is single-description encoding, the generation unit is further used to generate sixth data based on fourth data, fifth data, and a second delay, where the fourth data is the original audio data of the first audio frame, the fifth data is the original audio data of the second audio frame, and the second delay is the encoding delay of the single-description encoding. The encoding unit is further used to single-description encode the sixth data to obtain the encoded data of the first audio frame.
[0023] In an alternative embodiment of the example of the present application, specifically, the generation unit is used to cut off sample points with a length equal to the first delay from the end of the second data to obtain seventh data, connect the seventh data to the front end of the first data to obtain eighth data, and delete sample points with a length equal to the first delay from the end of the eighth data to obtain the third data.
[0024] As an alternative embodiment of the embodiment of the present application, specifically, the generating unit cuts off a sample point having a length of the second delay from the end of the fifth data to obtain ninth data, connects the ninth data to the front end of the fourth data to obtain tenth data, and deletes a sample point having a length of the second delay from the end of the tenth data to obtain the sixth data.
[0025] As an alternative embodiment of the embodiment of the present application, specifically, the determining unit determines whether the encoding mode switching condition is satisfied based on the encoding mode duration length and the signal type of the first audio frame, where the encoding mode duration length is the playback duration length of the audio frames continuously encoded in the current encoding mode, and in the case of NO, determines the encoding mode of the second audio frame as the encoding mode of the first audio frame, and in the case of YES, determines the encoding mode of the first audio frame based on the network parameters of the audio encoding data transmission network.
[0026] As an alternative embodiment of the embodiment of the present application, specifically, the determining unit determines whether the encoding mode duration length is greater than a threshold duration length, determines whether the probability that the first audio frame is a voice audio frame is less than a threshold probability, and when the encoding mode duration length is greater than the threshold duration length and the probability that the first audio frame is a voice audio frame is less than the threshold probability, determines that the encoding mode switching condition is satisfied, and when the encoding mode duration length is less than or equal to the threshold duration length and / or the probability that the first audio frame is a voice audio frame is greater than or equal to the threshold probability, determines that the encoding mode switching condition is not satisfied.
[0027] As an alternative embodiment of the embodiments of the present application, specifically, the determination unit is used to determine the packet loss rate of the audio encoding data transmission network based on the network parameters, determine whether the packet loss rate is greater than or equal to the threshold packet loss rate, and if YES, determine that the encoding mode of the first audio frame is multi-description encoding, and if NO, determine that the encoding mode of the first audio frame is single-description encoding.
[0028] According to a fourth aspect, an embodiment of the present disclosure provides a decoding apparatus for audio data, the apparatus comprising: a determination unit for determining an encoding mode of a first audio frame based on the encoded data of the first audio frame; a decoding unit for decoding the encoded data of the first audio frame based on the encoding mode to obtain decoded data; a determination unit for determining whether the encoding mode of the first audio frame is the same as the encoding mode of a second audio frame, where the second audio frame is the audio frame before the first audio frame; and a processing unit for generating packet loss compensation data based on the second audio frame and smoothing the decoded data based on the delay data of the second audio frame and the packet loss compensation data to obtain playback data of the first audio frame when the encoding mode of the first audio frame is not the same as the encoding mode of the second audio frame and the encoding mode of the first audio frame is multi-description encoding.
[0029] As an alternative embodiment of the embodiments of the present disclosure, when the encoding mode of the first audio frame is different from the encoding mode of the second audio frame and the encoding mode of the first audio frame is single-description encoding, the processing unit further generates packet loss compensation data based on the second audio frame, smoothes the decoded data based on the packet loss compensation data to obtain a smoothed result corresponding to the decoded data, and delays the smoothed result based on the packet loss compensation data and the number of delay sample points to obtain playback data of the first audio frame, where the number of delay sample points is the number of delay sample points of multi-description encoding.
[0030] As an alternative embodiment of the example of the present application, specifically, the processing unit replaces a first sample point sequence in the decoded data with a second sample point sequence in the packet loss compensation data to obtain a first replacement result, where the first sample point sequence is a sample point sequence composed of the first first several sample points of the decoded data, and the first number is a difference value between a first preset number and the number of delay sample points, and the second sample point sequence is a sample point sequence composed of sample points in the packet loss compensation data whose index value ranges from the number of delay sample points to the first preset number, and is used to window and superimpose a third sample point sequence in the first replacement result and a fourth sample point sequence in the packet loss compensation data based on a first window function to obtain a smoothing result corresponding to the decoded data, where the third sample point sequence is a sample point sequence composed of sample points in the first replacement result whose index value ranges from the first number to the sum of the first number and a second preset number, and the fourth sample point sequence is a sample point sequence composed of sample points in the packet loss compensation data whose index value ranges from the first preset number to the sum of the first preset number and a second preset number.
[0031] As an alternative embodiment of the embodiment of the present application, specifically, the processing unit is configured to obtain a fifth sample point sequence, where the fifth sample point sequence is composed of sample points corresponding to the first delay sample point number of the packet loss compensation data, concatenate the fifth sample point sequence in front of the smoothing result to obtain a first concatenation result, and delete a sixth sample point sequence in the first concatenation result to obtain the playback data of the first audio frame, where the sixth sample point sequence is composed of sample points corresponding to the last delay sample point number of the first concatenation result.
[0032] As an alternative embodiment of the embodiment of the present application, specifically, the processing unit is configured to replace a seventh sample point sequence in the decoded data with the delay data to obtain a second replacement result, where the seventh sample point sequence is composed of sample points corresponding to the first delay sample point number of the decoded data, window and superimpose an eighth sample point sequence in the second replacement result and a ninth sample point sequence in the packet loss compensation data based on a second window function to obtain the playback data of the first audio frame, where the eighth sample point sequence is composed of sample points with an index value in the second replacement result ranging from the delay sample point number to the sum of the delay sample point number and a third preset number, and the ninth sample point sequence is composed of sample points corresponding to the first third preset number of the packet loss compensation data.
[0033] As an alternative embodiment of the embodiment of the present application, when the encoding mode of the first audio frame is the same as the encoding mode of the second audio frame and the encoding mode of the first audio frame is single-description encoding, the processing unit is further used to delay-process the decoded data based on the delay data of the second audio frame and the number of delay sample points to obtain the playback data of the first audio frame.
[0034] As an alternative embodiment of the embodiment of the present application, specifically, the processing unit is used to splice the delay data in front of the decoded data to obtain a second splicing result, delete the tenth sample point sequence in the second splicing result, and obtain the playback data of the first audio frame. The tenth sample point sequence is a sample point sequence composed of the last number of delay sample points of the second splicing result.
[0035] According to a fifth aspect, an embodiment of the present disclosure provides an electronic device, which includes a memory and a processor. The memory is used to store a computer program, and the processor is used to implement the audio data encoding method or the audio data decoding method described in any one of the above embodiments when executing the computer program.
[0036] According to a sixth aspect, an embodiment of the present disclosure provides a computer-readable storage medium, which realizes the audio data encoding method or the audio data decoding method described in any one of the above embodiments when the computer program is executed by a computer device.
[0037] According to a seventh aspect, an embodiment of the present disclosure provides a computer program product that, when operating on a computer, causes the computer to implement the audio data encoding method or the audio data decoding method described in any one of the above embodiments.
Advantages of the Invention
[0038] The audio data encoding method and decoding method according to the embodiments of the present disclosure generate target data by determining an encoding mode of a first audio frame, determining whether the encoding mode of the first audio frame is the same as an encoding mode of a second audio frame, and generating target data based on first data, second data, and a first delay when they are not the same and the encoding mode of the first audio frame is multi-description coding. The audio data encoding method according to the embodiments of the present disclosure processes low-frequency data obtained by downsampling the original audio data of the first audio frame based on low-frequency data obtained by downsampling the original audio data of the second audio frame and an encoding delay of multi-description coding when the encoding mode of the first audio frame is not the same as the encoding mode of the second audio frame and the encoding mode of the first audio frame is multi-description coding, and encodes the third data obtained by the processing. Therefore, the embodiments of the present application can avoid problems of audio discontinuity and noise and further improve the quality of the audio signal when switching the change mode from single-description coding to multi-description coding.
Brief Description of the Drawings
[0039] The drawings herein are incorporated into the specification and form a part of the specification, showing embodiments consistent with the present disclosure, and are used together with the specification to interpret the principles of the present disclosure.
[0040] To more clearly explain the technical solutions in the embodiments of the present disclosure or the prior art, the following briefly introduces the drawings that need to be referred to in the description of the embodiments or the prior art. Obviously, for those skilled in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0041]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17
Embodiments for Carrying Out the Invention
[0042] In order to more clearly understand the above objects, features and advantages of the present disclosure, the following further describes the solutions of the present disclosure. It should be noted that, unless there is a conflict, the embodiments of the present disclosure and the features in the embodiments can be combined with each other.
[0043] The following description includes many specific details for facilitating a full understanding of the present disclosure. However, the present disclosure may be implemented in a manner different from that described herein. Obviously, the embodiments in the specification are only some embodiments of the present disclosure, not all embodiments.
[0044] In the embodiments of the present disclosure, terms such as "exemplary" or "for example" are used to represent by way of example, illustration or explanation. Any embodiment or design described as "exemplary" or "for example" in the embodiments of the present disclosure should not be construed as being more preferred or having more advantages than other embodiments or designs. Exactly speaking, invoking terms such as "exemplary" or "for example" is intended to present related concepts in a specific manner. In addition, in the description of the embodiments of the present disclosure, unless otherwise specified, "a plurality" means two or more.
[0045] Embodiments of the present disclosure provide a method for encoding audio data. Referring to FIG. 1, this method for encoding audio data includes the following steps.
[0046] S101. Determine the encoding mode of the first audio frame.
[0047] In the embodiments of the present disclosure, the encoding methods for audio frames are Single Description Coding (SDC) and Multiple Description Coding (MDC).
[0048] S102. Determine whether the encoding mode of the first audio frame is the same as that of the second audio frame.
[0049] Here, the second audio frame is the audio frame before the first audio frame.
[0050] In step S102 above, when the encoding mode of the first audio frame is different from that of the second audio frame, and the encoding mode of the first audio frame is Multiple Description Coding, steps S103 and S104 below are executed.
[0051] S103. Generate third data based on first data, second data, and a first delay.
[0052] Here, the first data is low-frequency data obtained by downsampling the original audio data of the first audio frame, the second data is low-frequency data obtained by downsampling the original audio data of the second audio frame, and the first delay is the encoding delay of the Multiple Description Coding.
[0053] In some embodiments, when the encoding mode of the current audio frame is multi-description encoding, the original data of the current audio frame is written into a delay buffer. However, when the encoding mode of the current audio frame is single-description encoding, in order to obtain a second data, when it is necessary to obtain the low-frequency data obtained by downsampling the original audio data of the previous audio frame, the low-frequency data obtained by downsampling the original audio data of the previous audio frame is directly read from the delay buffer by writing the low-frequency data obtained by downsampling the original data of the current audio frame into a specified buffer.
[0054] S104. Multiply-description encode the third data to obtain the encoded data of the first audio frame.
[0055] In step S102 above, when the encoding mode of the first audio frame is different from the encoding mode of the second audio frame, and the encoding mode of the first audio frame is single-description encoding, the following steps S105 and S106 are executed.
[0056] S105. Generate a sixth data based on a fourth data, a fifth data, and a second delay.
[0057] Here, the fourth data is the original audio data of the first audio frame, and the fifth data is the original audio data of the second audio frame. The second delay is the encoding delay of the single-description encoding.
[0058] S106. Single-description encode the sixth data to obtain the encoded data of the first audio frame.
[0059] According to an embodiment of the present disclosure, an audio data encoding method and a decoding method generate target data through steps of determining an encoding mode of a first audio frame, determining whether the encoding mode of the first audio frame is the same as an encoding mode of a second audio frame, and generating target data based on first data, second data, and a first delay when they are not the same and the encoding mode of the first audio frame is multi-description encoding. According to an embodiment of the present disclosure, an audio data encoding method divides low-frequency data obtained by downsampling original audio data of a first audio frame based on low-frequency data obtained by downsampling original audio data of a second audio frame and an encoding delay of multi-description encoding when the encoding mode of the first audio frame is not the same as the encoding mode of the second audio frame and the encoding mode of the first audio frame is multi-description encoding, processes the low-frequency data obtained by downsampling the original audio data of the first audio frame, and encodes third data obtained by the processing. Therefore, an embodiment of the present application can avoid problems of audio discontinuity and noise and further improve the quality of an audio signal when switching a change mode from single-description encoding to multi-description encoding.
[0060] As a refinement and extension of the above embodiment, an embodiment of the present disclosure provides an audio data encoding method. Referring to FIG. 2, this audio data encoding method includes the following steps.
[0061] S201. Determine an encoding mode of a first audio frame.
[0062] That is, determine the encoding mode of the current audio frame.
[0063] S202. Determine whether the encoding mode of the first audio frame is the same as an encoding mode of a second audio frame.
[0064] Here, the second audio frame is the audio frame before the first audio frame.
[0065] That is, it is determined whether the encoding mode of the current audio frame is the same as the encoding mode of the previous audio frame.
[0066] In S202 above, when the encoding mode of the first audio frame is not the same as the encoding mode of the second audio frame and the encoding mode of the first audio frame is multi-description encoding, the following S203 to S206 are executed.
[0067] S203: Cut off sample points with a length equal to the first delay from the end of the second data to obtain seventh data.
[0068] S204: Connect the fifth data to the front end of the first data to obtain eighth data.
[0069] S205: Delete sample points with a length equal to the first delay from the end of the eighth data to obtain the third data.
[0070] S206: Multiply-descriptively encode the third data to obtain the encoded data of the first audio frame.
[0071] When the encoding mode of the first audio frame is not the same as that of the second audio frame and the encoding mode of the first audio frame is multi-description encoding, that is, when the encoding mode of the current audio frame is multi-description encoding and the encoding mode of the previous frame audio frame is single-description encoding, referring to Figure 4 as shown in Figure 3, the first delay extension in Figure 3 is delay_8kHZ, and the data buffered in the delay buffer is the low-frequency data (second data 31) obtained by decimating the original data of the second audio frame. The input to the encoder during multi-description encoding is the low-frequency data (first data 32) obtained by decimating the first audio frame. The data processing process of steps S203 to S205 above first cuts off sample points with a length of the delay_8kHZ from the end of the second data 31 to obtain the seventh data 311, then concatenates the seventh data 311 to the front end of the first data 32 to obtain the eighth data 33, and finally deletes sample points with a length of delay_8kHZ from the end of the eighth data 33 to obtain the third data 34. As shown in Figure 3, the third data 34 is composed of two parts. One part is the seventh data 311, and the other part is the remaining data of the first data 32 after deleting sample points with a length of delay_8kHZ from the end of the first data 32.
[0072] In S202 above, when the encoding mode of the first audio frame is not the same as that of the second audio frame and the encoding mode of the first audio frame is multi-description encoding, the following S207 to S210 are executed.
[0073] S207: Cut off sample points with a length of the second delay from the end of the fifth data to obtain the ninth data.
[0074] S208. Connect the ninth data to the front end of the fourth data to obtain tenth data.
[0075] S209. Delete sample points with a length equal to the second delay from the end of the tenth data to obtain the sixth data.
[0076] S210. Singly describe and encode the sixth data to obtain the encoded data of the first audio frame.
[0077] When the encoding mode of the first audio frame is not the same as that of the second audio frame, and the encoding mode of the first audio frame is single-description encoding, that is, when the encoding mode of the current audio frame is single-description encoding and the encoding mode of the previous frame of the audio frame is multi-description encoding, referring to FIG. 4 as shown, the first delay length in FIG. 4 is delay_16kHZ, the data buffered in the delay buffer is the original audio data (the fifth data 41) of the second audio frame, and the input of the output encoder during single-description encoding is the original audio data (the fourth data 42) of the first audio frame. The data processing process of the above steps S207 to S209 includes first cutting off sample points with a length equal to the delay_16kHZ from the end of the fifth data 41 to obtain the ninth data 411, then connecting the ninth data 411 to the front end of the fourth data 42 to obtain the tenth data 43, and finally deleting sample points with a length equal to delay_16kHZ from the end of the tenth data 43 to obtain the sixth data 44. As shown in FIG. 4, the sixth data 44 is composed of two parts. One part is the ninth data 411, and the other part is the remaining data of the fourth data 42 after deleting sample points with a length equal to delay_16kHZ from the end of the fourth data 42.
[0078] As a refinement and extension of the above embodiments, embodiments of the present disclosure provide a method for processing audio data. Referring to FIG. 5, this method for processing audio data includes the following.
[0079] S501. Determine whether the encoding mode switching condition is satisfied based on the encoding mode duration and the signal type of the first audio frame.
[0080] Here, the encoding mode duration is the playback duration of the audio frames continuously encoded in the current encoding mode.
[0081] In some embodiments, the implementation manner of determining whether the encoding mode switching condition is satisfied based on the encoding mode duration and the signal type of the first audio frame may include the following steps a to d.
[0082] Step a. Determine whether the encoding mode duration is greater than a threshold duration.
[0083] Embodiments of the present application do not limit the threshold duration. Exemplarily, the threshold duration may be 2 s.
[0084] In the above step a, when the encoding mode duration is less than or equal to the threshold duration, execute the following step b.
[0085] Step b. Determine that the encoding mode switching condition is not satisfied.
[0086] In the above step a, when the encoding mode duration is greater than the threshold duration, execute the following steps c to e.
[0087] Step c. Determine whether the probability that the first audio frame is a voice audio frame is less than a threshold probability.
[0088] In the above step c, if the probability that the first audio frame is a voice audio frame is less than the threshold probability, the following step d is executed.
[0089] Step d: Determine that the encoding mode switching condition is satisfied.
[0090] In the above step c, if the probability that the first audio frame is a voice audio frame is greater than or equal to the threshold probability, the following step e is executed.
[0091] Step e: Determine that the encoding mode switching condition is not satisfied.
[0092] That is, when the encoding mode duration is less than or equal to the threshold duration and / or the probability that the first audio frame is a voice audio frame is greater than or equal to the threshold probability, it is determined that the encoding mode switching condition is not satisfied.
[0093] In the above step S501, if the encoding mode switching condition is not satisfied, the following step S502 is executed.
[0094] S502: Determine the encoding mode of the second audio frame as the encoding mode of the first audio frame.
[0095] That is, reuse the encoding mode of the previous audio frame for encoding.
[0096] In the above step S501, if the encoding mode switching condition is satisfied, the following step S503 is executed.
[0097] S503: Determine the encoding mode of the first audio frame based on the network parameters of the audio encoding data transmission network.
[0098] In some embodiments, the steps of the implementation method of step S503 (determining the encoding mode of the first audio frame based on the network parameters of the audio encoded data transmission network) are steps 1 to 3 below.
[0099] Step 1, determine the packet loss rate of the audio encoded data transmission network based on the network parameters.
[0100] The packet loss rate in the embodiments of the present application refers to the ratio of the number of data packets lost during the data packet transmission process to all the transmitted data packets.
[0101] Step 2, determine whether the packet loss rate is greater than or equal to the threshold packet loss rate.
[0102] The embodiments of the present application do not limit the threshold packet loss rate. Exemplarily, the threshold packet loss rate may be 5%.
[0103] In step 2 above, if the packet loss rate is greater than or equal to the threshold packet loss rate, execute step 3 below. If the packet loss rate is less than the threshold packet loss rate, execute step 4 below.
[0104] Step 3, determine that the encoding mode of the first audio frame is multi-description coding.
[0105] Step 4, determine that the encoding mode of the first audio frame is single-description coding.
[0106] S504, determine whether the encoding mode of the first audio frame is the same as the encoding mode of the second audio frame.
[0107] In S504 above, when the encoding mode of the first audio frame is not the same as that of the second audio frame and the encoding mode of the first audio frame is multi-description encoding, the following S505 to S508 are executed.
[0108] S505. Cut off sample points with a length equal to the first delay from the end of the second data to obtain seventh data.
[0109] Here, the second data is low-frequency data obtained by decimating the original audio data of the second audio frame, and the first delay is the encoding delay of the multi-description encoding.
[0110] S506. Connect the seventh data to the front end of the first data to obtain eighth data.
[0111] Here, the first data is low-frequency data obtained by decimating the original audio data of the first audio frame.
[0112] S507. Delete sample points with a length equal to the first delay from the end of the eighth data to obtain the third data.
[0113] S508. Perform multi-description encoding on the third data to obtain the encoded data of the first audio frame.
[0114] In S504 above, when the encoding mode of the first audio frame is not the same as that of the second audio frame and the encoding mode of the first audio frame is single-description encoding, the following S509 to S512 are executed.
[0115] S509. Cut off sample points with a length equal to the second delay from the end of the fifth data to obtain ninth data.
[0116] S510. Connect the ninth data to the tip of the fourth data to obtain tenth data.
[0117] S511. Delete sample points with a length equal to the second delay from the end of the tenth data to obtain the fourth data.
[0118] S512. Singly describe and encode the sixth data to obtain the encoded data of the first audio frame.
[0119] In S504, when the encoding mode of the first audio frame is the same as the encoding mode of the second audio frame and the encoding mode of the first audio frame is multi-description encoding, multi-description encode the low-frequency data obtained by decimating the original audio data of the first audio frame to obtain the encoded data of the first audio frame. In S504, when the encoding mode of the first audio frame is the same as the encoding mode of the second audio frame and the encoding mode of the first audio frame is single-description encoding, singly describe and encode the original audio data of the first audio frame to obtain the encoded data of the first audio frame.
[0120] Embodiments of the present disclosure provide a method for decoding audio data. Referring to FIG. 6, this method for decoding audio data includes the following.
[0121] S601. Determine the encoding mode of the first audio frame based on the encoded data of the first audio frame.
[0122] S602. Decode the encoded data of the first audio frame based on the encoding mode to obtain decoded data.
[0123] S603. Determine whether the encoding mode of the first audio frame is the same as that of the second audio frame.
[0124] Here, the second audio frame is the audio frame before the first audio frame.
[0125] In S603 above, when the encoding mode of the first audio frame is not the same as that of the second audio frame, and the encoding mode of the first audio frame is single description coding, execute the following S604 - S606.
[0126] S604. Generate packet loss compensation data based on the second audio frame.
[0127] The packet loss compensation data is data obtained based on a Packet Loss Concealment (PLC) mechanism. The packet loss compensation mechanism is for the media engine to solve network packet loss problems. When the media engine receives a series of media stream data packets, it cannot guarantee that all packets are received. When a data packet is lost and the Forward Error Correction (FEC) mechanism is not used at this time, the packet loss compensation mechanism functions. The packet loss compensation mechanism is not a standard match and allows the media engine and codec to be realized and extended according to their own situations.
[0128] The packet loss compensation data in the embodiments of this application may be data with a length of 10 ms.
[0129] S605. Smooth the decoded data based on the packet loss compensation data to obtain a smoothed result corresponding to the decoded data.
[0130] S606. Delay-process the smoothing result based on the packet loss compensation data and the number of delay sample points to obtain the playback data of the first audio frame.
[0131] Here, the number of delay sample points is the number of delay sample points of multi-description coding.
[0132] In the embodiments of the present disclosure, since there is a delay of qmf_order - 1 sample points in the MDC algorithm itself, when the coding method of the first audio frame is MDC, the output audio delay after decoding may be set to 0. However, when the coding method of the first audio frame is SDC, in order to match the delay of the MDC algorithm, it is necessary to set the output audio delay after decoding to qmf_order - 1. Matching the delays of the two algorithms may be realized by the following formula. Number of delay sample points = qmf_order - 1, coding mode = SDC Number of delay sample points = 0, coding mode = MDC
[0133] In S603 above, when the coding mode of the first audio frame and the coding mode of the second audio frame are not the same, and the coding mode of the first audio frame is multi-description coding, the following S607 and S608 are executed.
[0134] S607. Generate packet loss compensation data based on the second audio frame.
[0135] S608. Smooth the decoded data based on the delay data of the second audio frame and the packet loss compensation data to obtain the playback data of the first audio frame.
[0136] When decrypting the data packet of the first audio data in the above embodiment, first, determine the encoding mode of the first audio frame based on the encoded data of the first audio frame, and decrypt the encoded data of the first audio frame based on the encoding mode to obtain decrypted data. Further, determine whether the encoding mode of the first audio frame is the same as the encoding mode of the second audio frame. If the encoding modes are not the same and the encoding mode of the first audio frame is single-description encoding, generate packet loss compensation data based on the second audio frame, and further smooth the decrypted data based on the packet loss compensation data to obtain a smoothing result corresponding to the decrypted data. Delay-process the smoothing result based on the packet loss compensation data and the number of delay sample points to obtain the playback data of the first audio frame. If the encoding modes are not the same and the encoding mode of the first audio frame is multi-description encoding, generate packet loss compensation data based on the second audio frame, and further smooth the decrypted data based on the delay data of the second audio frame and the packet loss compensation data to obtain the playback data of the first audio frame.According to an embodiment of the present disclosure, a method for decrypting audio data is as follows: when the encoding mode of the first audio frame is different from the encoding mode of the second audio frame, and the encoding mode of the first audio frame is single-description encoding, packet loss compensation data is generated based on the second audio frame, and further, the decrypted data is smoothed to obtain the playback data of the first audio frame. When the encoding mode of the first audio frame is different from the encoding mode of the second audio frame, and the encoding mode of the first audio frame is multi-description encoding, packet loss compensation data is generated based on the second audio frame, and further, the playback data of the first audio frame is obtained by combining with the delay data of the second audio frame. Therefore, when the encoding mode of the first audio frame is different from the encoding mode of the second audio frame, according to an embodiment of the present application, the encoded data is processed based on the current audio frame encoding mode type, further avoiding the problems of audio discontinuity and noise, and further improving the quality of the audio signal.
[0137] As a refinement and extension of the above embodiment, an embodiment of the present disclosure provides a method for decrypting audio data. Referring to FIG. 7, this method for decrypting audio data includes the following steps.
[0138] S701. Determine the encoding mode of the first audio frame based on the encoded data of the first audio frame.
[0139] S702. Decrypt the encoded data of the first audio frame based on the encoding mode to obtain decrypted data.
[0140] S703. Determine whether the encoding mode of the first audio frame is the same as the encoding mode of the second audio frame.
[0141] Here, the second audio frame is the audio frame before the first audio frame.
[0142] In S703 above, when the encoding mode of the first audio frame and the encoding mode of the second audio frame are the same, and the encoding mode of the first audio frame is single-description coding, the following S704 to S706 are executed.
[0143] S704. Replace the first sample point sequence in the decoded data with the second sample point sequence in the packet loss compensation data to obtain a first replacement result.
[0144] Here, the first sample point sequence is a sample point sequence composed of the first first several sample points of the decoded data, the first number is a difference value between a first preset number and the number of delay sample points, and the second sample point sequence is a sample point sequence composed of sample points in the packet loss compensation data whose index value ranges from the number of delay sample points to the first preset number.
[0145] In some embodiments, when the encoding mode of the current audio frame is multi-description encoding, the original data of the current audio frame is written into the transition buffer. However, when the encoding mode of the current audio frame is single-description encoding, in order to obtain the second data when it is necessary to obtain the low-frequency data obtained by downsampling the original audio data of the previous audio frame by directly reading it from the delay buffer. The decoded data is stored in the pulse code modulation buffer (pcm_buffer), the first replacement result is written to the storage location of the original decoded data in the pulse code modulation buffer, and the second replacement result obtained in S707 below is also written to the storage location of the original decoded data in the pulse code modulation buffer.
[0146] In this embodiment, the decoded data is stored in the pulse code modulation buffer, and the first sample point sequence in the decoded data is the first F5 - Fd sample point sequence of the pulse code modulation buffer. The second sample point sequence in the packet loss compensation data is the sample point sequence where the index value in the packet loss compensation data is Fd~F5, and obtaining the first replacement result may be realized by the following formula. pcm_buffe(i - Fd)=transition_buffer(i) i = Fd,……,F5 - 1
[0147] In this embodiment, referring to FIG. 8, the number of delay sample points is Fd, and the first preset number is F5. In FIG. 8, what is stored in the transition buffer is packet loss compensation data (packet loss compensation data 81) generated based on the second audio frame, and the sample point sequence composed of sample points where the index value in the transition buffer ranges from the number of delay sample points to the first preset number is the second sample point sequence 811. What is stored in the pulse code modulation buffer is decoded data (decoded data 82) obtained by decoding the encoded data of the first audio frame based on the encoding mode, and it is the sample point sequence (first sample point sequence 821) composed of the first few sample points in the pulse code modulation buffer. Then, step S704 is to replace the first sample point sequence 821 in the decoded data 82 with the second sample point sequence 811 in the packet loss compensation data 81 to obtain a first replacement result 83.
[0148] S705. Based on the first window function, window and superimpose the third sample point sequence in the first replacement result and the fourth sample point sequence in the packet loss compensation data to obtain a smoothing result corresponding to the decoded data.
[0149] Here, the third sample point sequence is a sample point sequence composed of sample points where the index value in the first replacement result ranges from the first number to the sum of the first number and the second preset number, and the fourth sample point sequence is a sample point sequence composed of sample points where the index value in the packet loss compensation data ranges from the first preset number to the sum of the first preset number and the second preset number.
[0150] Window function: Since the Fourier transform can only transform time-domain data of finite length, it is necessary to perform signal truncation on the time-domain signal. Even for periodic signals, if the truncation time length is not an integer multiple of the period (periodic truncation), there will be leakage in the signal after truncation. In order to minimize this leakage error, it is necessary to use a weighting function, also called a window function. Windowing is mainly to make the time-domain signal appear to better meet the periodic requirements of Fourier processing and reduce leakage. In this embodiment, smoothing is performed according to the switching type, and transition smoothing is performed using the windowing smoothing method.
[0151] In this embodiment, the third sample point sequence is a sample point sequence in which the index value in the pulse-coded modulation buffer is F5 - Fd to F5 - Fd + F2.5. The fourth sample point sequence is a sample point sequence in which the index value in the transition buffer is F5 - Fd to F5 - Fd + F2.5. Obtaining the smoothing result may be realized by the following formula. pcm_buffe(i + F5 - Fd) = w(i) * pcm(i + F5 - Fd)+(1 - w(i)) * transition_buffer(i + F5) i = 0, 1, …… F2.5 - 1
[0152] Here, w(i) is the expression of the window function. The smoothing method is to window and superimpose the corresponding part and the sample point whose index in the transition buffer is F5 to F5 + F2.5 to achieve the purpose of smooth transition.
[0153] Based on the embodiment shown in FIG. 8 above, referring to FIG. 9, the second preset number is F2.5. Based on S704 above, the sample point sequence composed of sample points where the index value in the transition buffer ranges from the first preset number to the sum of the first preset number and the second preset number is the fourth sample point sequence 812, and the sample point sequence composed of sample points where the index value in the first replacement result 83 in the pulse coding modulation buffer ranges from the first number to the sum of the first number and the second preset number is the third sample point sequence 831. Window and superimpose the third sample point sequence 831 in the first replacement result 83 and the fourth sample point sequence 812 in the packet loss compensation data 81 to obtain the smoothing result 91 corresponding to the decoded data.
[0154] S706. Obtain a fifth sample point sequence.
[0155] Here, the fifth sample point sequence is a sample point sequence composed of the first delay sample point number of sample points in the packet loss compensation data.
[0156] In this embodiment, the fifth sample point sequence is the first Fd sample point sequences in the transition buffer. Obtaining the fifth sample point sequence may be realized by the following formula. delay_buffer(i)=transition_buffer(i) i = 0, 1, …… Fd - 1
[0157] S707. Connect the fifth sample point sequence in front of the smoothing result to obtain a first connection result.
[0158] Delete the sixth sample point sequence in the first splicing result to obtain the playback data of the first audio frame.
[0159] Here, the sixth sample point sequence is a sample point sequence composed of sample points for the last number of delayed sample points in the first splicing result.
[0160] Based on the embodiment shown in FIG. 9 above, referring to FIG. 10 as shown, the sample point sequence composed of the sample points for the first number of delayed sample points of the packet loss compensation data is the fifth sample point sequence 101. First, splice the fifth sample point sequence 101 in front of the smoothing result 91 to obtain the first splicing result 102. The sample point sequence composed of the sample points for the last number of delayed sample points in the first splicing result 102 is the sixth sample point sequence 103. Then, delete the sixth sample point sequence 103 in the first splicing result 102 to obtain the playback data 104 of the first audio frame. The playback data 104 of the first audio frame consists of the fifth sample point sequence 101, the second sample point sequence 811, and the remaining part after deleting the sixth sample point sequence from the end of the first splicing result 102.
[0161] In S703 above, when the encoding mode of the first audio frame is the same as the encoding mode of the second audio frame and the encoding mode of the first audio frame is multi-description coding, execute the following S709 and S710.
[0162] S709: Replace the seventh sample point sequence in the decoded data with the delay data to obtain a second replacement result.
[0163] Here, the seventh sample point sequence is a sample point sequence composed of the first delay sample points in the decoded data, and obtaining the second replacement result may be realized by the following formula. pcm_buffer(i)=delay_buffer(i) i = 0, 1, …… qmf_order - 2
[0164] In this embodiment, referring to FIG. 11 as shown, in the figure, what is in the pulse code modulation buffer is the decoded data (decoded data 112) obtained by decoding the encoded data of the first audio frame based on the encoding mode. The first qmf_order - 1 sample point sequence in the pulse code modulation buffer is the seventh sample point sequence 1121. The first qmf_order - 1 sample point sequence in the delay buffer is the delay data 111. Replace the seventh sample point sequence 1121 in the decoded data 112 with the delay data 111 to obtain a second replacement result 113. The second replacement result 113 consists of the remaining parts at the ends in the delay data 111 and the decoded data 112.
[0165] S710. Window and superimpose the eighth sample point sequence in the second replacement result and the ninth sample point sequence in the packet loss compensation data based on the second window function.
[0166] Here, the eighth sample point sequence is a sample point sequence composed of sample points where the index value in the second replacement result ranges from the number of delay sample points to the sum of the number of delay sample points and a third preset number. The ninth sample point sequence is a sample point sequence composed of the first third preset number of sample points of the packet loss compensation data. Based on the second window function of the above step, windowing and superimposing the eighth sample point sequence in the second replacement result and the ninth sample point sequence in the packet loss compensation data may be realized by the following formula. pcm_buffe(i+qmf_order-1) =(i)*pcm(i+qmf_order-1)+(1-(i)) *transition_buffer(i), i = 0, 1, …… F2.5-1
[0167] In this embodiment, based on the embodiment shown in FIG. 11 above, referring to FIG. 12, the sample point sequence composed of the first third preset number of sample points in the transition buffer is the ninth sample point sequence 1211. The sample point sequence composed of sample points where the index value in the second replacement result 113 in the pulse code modulation buffer ranges from the number of delay sample points to the sum of the number of delay sample points and a third preset number is the eighth sample point sequence 1031. The ninth sample point sequence 1211 and the eighth sample point sequence 1031 in the packet loss compensation data 121 are windowed and superimposed to obtain a result 122. The result 122 obtained by windowing and superimposing is composed of the delay data 102, the smoothing result 1221 obtained by windowing and superimposing the ninth sample point sequence 1211 and the eighth sample point sequence 1031, and the remaining part at the end in the second replacement result 104.
[0168] As a refinement and extension of the above embodiments, embodiments of the present disclosure provide a method for processing audio data. Referring to FIG. 13, this method for processing audio data includes the following steps.
[0169] S1301. Determine the encoding mode of the first audio frame based on the encoded data of the first audio frame.
[0170] S1302. Decode the encoded data of the first audio frame based on the encoding mode to obtain decoded data.
[0171] S1303. Determine whether the encoding mode of the first audio frame is the same as the encoding mode of the second audio frame.
[0172] In S1303 above, if the encoding mode of the first audio frame is the same as the encoding mode of the second audio frame, perform the following steps a and b.
[0173] Step a. Concatenate the delay data in front of the decoded data to obtain a second concatenation result.
[0174] Step b. Delete the tenth sample point sequence in the second concatenation result to obtain the playback data of the first audio frame.
[0175] Here, the tenth sample point sequence is a sample point sequence composed of sample points for the last number of delay sample points in the second concatenation result.
[0176] In some embodiments, steps a and b above may refer to FIG. 14. The delay data 1411 is the first qmf_order-1 sample point sequence in the delay buffer. The delay data 1411 is concatenated in front of the decoded data 142 to obtain a second concatenation result 143. The sample point sequence composed of the sample points of the last delay sample point number of the second concatenation result 143 is the tenth sample point sequence 1431. Then, the tenth sample point sequence 1431 in the second concatenation result 143 is deleted to obtain the playback data 144 of the first audio frame. The playback data 144 of the first audio frame consists of the delay data 1411 and the remaining part at the end of the decoded data 142.
[0177] In S1303 above, when the encoding mode of the first audio frame is not the same as the encoding mode of the second audio frame, and the encoding mode of the first audio frame is single-description coding, the following S1304 to S1306 are executed.
[0178] S1304, Generate packet loss compensation data based on the second audio frame.
[0179] S1305, Replace the first sample point sequence in the decoded data with the second sample point sequence in the packet loss compensation data to obtain a first replacement result.
[0180] Here, the first sample point sequence is a sample point sequence composed of the first first number of sample points of the decoded data. The first number is the difference value between the first preset number and the delay sample point number. The second sample point sequence is a sample point sequence composed of the sample points in the packet loss compensation data where the index value ranges from the delay sample point number to the first preset number.
[0181] S1306. Based on the first window function, window and superimpose the third sample point sequence in the first replacement result and the fourth sample point sequence in the packet loss compensation data to obtain a smoothing result corresponding to the decoded data.
[0182] Here, the third sample point sequence is a sample point sequence composed of sample points where the index value in the first replacement result ranges from the first number to the sum of the first number and a second preset number, and the fourth sample point sequence is a sample point sequence composed of sample points where the index value in the packet loss compensation data ranges from the first preset number to the sum of the first preset number and the second preset number.
[0183] In S1303 above, when the encoding mode of the first audio frame is not the same as the encoding mode of the second audio frame and the encoding mode of the first audio frame is multi-description encoding, the following S1307 to S1313 are executed.
[0184] S1307. Obtain a fifth sample point sequence.
[0185] Here, the fifth sample point sequence is a sample point sequence composed of the first delay sample point number of sample points in the packet loss compensation data.
[0186] S1308. Concatenate the fifth sample point sequence in front of the smoothing result to obtain a first concatenation result.
[0187] S1309. Delete the sixth sample point sequence in the first concatenation result to obtain the playback data of the first audio frame.
[0188] Here, the sixth sample point sequence is a sample point sequence composed of sample points for the last number of delayed sample points in the first splicing result.
[0189] S1310. When the encoding mode of the first audio frame is multi-description coding, generate packet loss compensation data based on the second audio frame.
[0190] S1311. Replace the seventh sample point sequence in the decoded data with the delay data to obtain a second replacement result.
[0191] S1312. Window and superimpose the eighth sample point sequence in the second replacement result and the ninth sample point sequence in the packet loss compensation data based on a second window function to obtain playback data for the first audio frame.
[0192] Here, the eighth sample point sequence is a sample point sequence composed of sample points where the index value in the second replacement result ranges from the number of delayed sample points to the sum of the number of delayed sample points and a third preset number, and the ninth sample point sequence is a sample point sequence composed of the first third preset number of sample points in the packet loss compensation data.
[0193] S1313. Delay-process the smoothing result based on the packet loss compensation data and the number of delayed sample points to obtain playback data for the first audio frame.
[0194] Based on the same inventive concept, as an implementation of the above method, embodiments of the present disclosure further provide an audio data encoding device and a decoding device. This embodiment corresponds to the embodiment of the above method. For the sake of readability, the detailed content in the embodiment of the above method is omitted in this embodiment. However, it should be clearly stated that the audio data processing device in this embodiment can be realized corresponding to all the content in the embodiment of the above method.
[0195] Embodiments of the present disclosure provide an audio data encoding device. FIG. 15 is a schematic structural diagram of this audio data processing device. Referring to FIG. 15 as shown, this audio data processing device 1500 includes the following: The determination unit 1501 is used to determine the encoding mode of the first audio frame. The judgment unit 1502 is used to judge whether the encoding mode of the first audio frame is the same as the encoding mode of the second audio frame. The second audio frame is the audio frame before the first audio frame. The generation unit 1503 is used to generate target data based on the first data, the second data, and the first delay when the encoding mode of the first audio frame is not the same as the encoding mode of the second audio frame and the encoding mode of the first audio frame is multi-description encoding. The first data is the low-frequency data obtained by decimating the original audio data of the first audio frame. The second data is the low-frequency data obtained by decimating the original audio data of the second audio frame. The first delay is the encoding delay of the multi-description encoding. The generating unit 1503 is further used to generate sixth data based on fourth data, fifth data, and a second delay when the encoding mode of the first audio frame is not the same as the encoding mode of the second audio frame and the encoding mode of the first audio frame is single-description encoding. The fourth data is the original audio data of the first audio frame, the fifth data is the original audio data of the second audio frame, and the second delay is the encoding delay of the single-description encoding. The encoding unit 1504 is used to encode the target data based on the encoding mode of the first audio frame to obtain the encoded data of the first audio frame.
[0196] As an alternative embodiment of the embodiment of the present disclosure, specifically, the generating unit 1503 is used to cut off sample points with a length equal to the first delay from the end of the second data to obtain fifth data, splice the fifth data to the front end of the first data to obtain sixth data, and delete sample points with a length equal to the first delay from the end of the sixth data to obtain the target data.
[0197] As an alternative embodiment of the embodiment of the present disclosure, specifically, the generating unit 1503 is used to cut off sample points with a length equal to the second delay from the end of the fifth data to obtain seventh data, splice the seventh data to the front end of the fourth data to obtain eighth data, and delete sample points with a length equal to the second delay from the end of the eighth data to obtain the target data.
[0198] As an alternative embodiment of the embodiments of the present disclosure, specifically, the determination unit 1501 determines whether to satisfy the encoding mode switching condition based on the encoding mode duration and the signal type of the first audio frame, where the encoding mode duration is the playback duration of the audio frames encoded continuously in the current encoding mode, and in the case of NO, determines the encoding mode of the second audio frame as the encoding mode of the first audio frame, and in the case of YES, is used to determine the encoding mode of the first audio frame based on the network parameters of the audio encoding data transmission network.
[0199] As an alternative embodiment of the embodiments of the present disclosure, specifically, the determination unit 1501 determines whether the encoding mode duration is greater than the threshold duration, determines whether the probability that the first audio frame is a voice audio frame is less than the threshold probability, and determines that the encoding mode switching condition is satisfied when the encoding mode duration is greater than the threshold duration and the probability that the first audio frame is a voice audio frame is less than the threshold probability, and is used to determine that the encoding mode switching condition is not satisfied when the encoding mode duration is less than or equal to the threshold duration and / or the probability that the first audio frame is a voice audio frame is greater than or equal to the threshold probability.
[0200] As an alternative embodiment of the embodiments of the present disclosure, specifically, the determination unit 1501 determines the packet loss rate of the audio encoding data transmission network based on the network parameters, determines whether the packet loss rate is greater than or equal to the threshold packet loss rate, and in the case of YES, determines that the encoding mode of the first audio frame is multi-description encoding, and in the case of NO, is used to determine that the encoding mode of the first audio frame is single-description encoding.
[0201] Embodiments of the present disclosure provide a decoder for audio data. FIG. 16 is a schematic structural diagram of the decoder for audio data. Referring to FIG. 16, the audio data decoder 1600 includes: a determination unit 1601 for determining an encoding mode of a first audio frame based on the encoded data of the first audio frame; a decoding unit 1602 for decoding the encoded data of the first audio frame based on the encoding mode to obtain decoded data; a determination unit 1603 for determining whether the encoding mode of the first audio frame is the same as the encoding mode of a second audio frame, where the second audio frame is the audio frame before the first audio frame; a processing unit 1604 configured to, when the encoding mode of the first audio frame is not the same as the encoding mode of the second audio frame and the encoding mode of the first audio frame is single-description coding, generate packet loss compensation data based on the second audio frame, smooth the decoded data based on the packet loss compensation data to obtain a smoothed result corresponding to the decoded data, and delay the smoothed result based on the packet loss compensation data and the number of delay sample points to obtain playback data of the first audio frame, where the number of delay sample points is the number of delay sample points of multi-description coding; The processing unit 1604 is further configured to, when the encoding mode of the first audio frame is not the same as the encoding mode of the second audio frame and the encoding mode of the first audio frame is multi-description coding, generate packet loss compensation data based on the second audio frame, and use the delayed data of the second audio frame and the packet loss compensation data to smooth the decoded data to obtain playback data of the first audio frame.
[0202] As an alternative embodiment of the embodiments of the present disclosure, specifically, the processing unit 1604 replaces the first sample point sequence in the decoded data with the second sample point sequence in the packet loss compensation data to obtain a first replacement result, where the first sample point sequence is a sample point sequence composed of the first first few sample points of the decoded data, the first number is a difference value between a first preset number and the number of delay sample points, the second sample point sequence is a sample point sequence composed of sample points in the packet loss compensation data where the index value ranges from the number of delay sample points to the first preset number, and is used to window and superimpose the third sample point sequence in the first replacement result and the fourth sample point sequence in the packet loss compensation data based on a first window function to obtain a smoothing result corresponding to the decoded data. The third sample point sequence is a sample point sequence composed of sample points in the first replacement result where the index value ranges from the first number to the sum of the first number and a second preset number, and the fourth sample point sequence is a sample point sequence composed of sample points in the packet loss compensation data where the index value ranges from the first preset number to the sum of the first preset number and the second preset number.
[0203] As an alternative embodiment of an example of the present disclosure, specifically, the processing unit 1604 is configured to obtain a fifth sample point sequence, where the fifth sample point sequence is composed of sample points corresponding to the first number of delayed sample points of the packet loss compensation data, connect the fifth sample point sequence in front of the smoothing result to obtain a first concatenation result, and delete a sixth sample point sequence in the first concatenation result to obtain playback data of the first audio frame, where the sixth sample point sequence is composed of sample points corresponding to the last number of delayed sample points of the first concatenation result.
[0204] As an alternative embodiment of an example of the present disclosure, specifically, the processing unit 1604 is configured to replace a seventh sample point sequence in the decoded data with the delay data to obtain a second replacement result, where the seventh sample point sequence is composed of sample points corresponding to the first number of delayed sample points in the decoded data, window and superimpose an eighth sample point sequence in the second replacement result and a ninth sample point sequence in the packet loss compensation data based on a second window function to obtain playback data of the first audio frame. The eighth sample point sequence is composed of sample points where the index value in the second replacement result ranges from the number of delayed sample points to the sum of the number of delayed sample points and a third preset number, and the ninth sample point sequence is composed of sample points corresponding to the first third preset number of sample points in the packet loss compensation data.
[0205] As an alternative embodiment of the embodiments of the present disclosure, specifically, when the encoding mode of the first audio frame is the same as the encoding mode of the second audio frame, and the encoding mode of the first audio frame is single description coding, the processing unit 1604 is used to delay-process the decoded data based on the delay data of the second audio frame and the number of delay sample points to obtain the playback data of the first audio frame.
[0206] As an alternative embodiment of the embodiments of the present disclosure, specifically, the processing unit 1604 is used to splice the delay data in front of the decoded data to obtain a second splicing result, and delete the tenth sample point sequence in the second splicing result to obtain the playback data of the first audio frame. The tenth sample point sequence is a sample point sequence composed of the last number of delay sample points of the second splicing result.
[0207] The audio data processing device according to this embodiment can execute the audio data processing method according to the embodiment of the above method. Its realization principle and technical effect are similar and will not be described further here.
[0208] Based on the same inventive concept, the embodiments of the present disclosure further provide an electronic device. FIG. 17 is a schematic structural diagram of an electronic device according to an embodiment of the present disclosure. As shown in FIG. 17, the electronic device according to this embodiment includes a memory 1701 and a processor 1702. The memory 1701 is used to store a computer program, and the processor 1702 is used to execute the audio data processing method according to the above embodiment when executing the computer program.
[0209] Based on the same inventive concept, embodiments of the present disclosure further provide a computer-readable storage medium, in which a computer program is stored, and when the computer program is executed by a processor, the computer device is caused to implement the method for processing audio data according to the above embodiments.
[0210] Based on the same inventive concept, embodiments of the present disclosure further provide a computer program product, and when the computer program product operates on a computer, the computer device is caused to implement the method for processing audio data according to the above embodiments.
[0211] Those skilled in the art should understand that the embodiments of the present disclosure can be provided as a method, a system, or a computer program product. Therefore, the present disclosure may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. And the present disclosure may also adopt the form of a computer program product implemented on one or more computer-readable storage media including computer-usable program code.
[0212] The processor may be a central rendering unit 103 (Central Processing Unit, CPU), or may be other general-purpose processors, digital signal processors (Digital Signal Processor, DSP), application specific integrated circuits (Application Specific Integrated Circuit, ASIC), field-programmable gate arrays (Field-Programmable Gate Array, FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware assemblies, etc. The general-purpose processor may be a microprocessor, or this processor may be any general processor, etc.
[0213] The memory may include non-persistent memory in a computer-readable medium, forms such as random access memory (RAM) and / or non-volatile memory, for example read-only memory (ROM) or flash memory. The memory is an example of a computer-readable medium.
[0214] Computer-readable media include persistent and non-persistent storage media, removable and non-removable storage media. The storage media can realize information storage by any method or technology, and the information may be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic disk storage or other magnetic storage devices or any other non-transmission media, and may be used to store information accessible by a computing device. According to the definitions herein, computer-readable media do not include transitory computer-readable media, such as modulated data signals and carriers.
[0215] Finally, it should be noted that each of the above embodiments is only used to illustrate the technical solutions of the present disclosure and does not limit them. Although the present disclosure has been described in detail with reference to the foregoing embodiments, those skilled in the art may still modify the technical solutions described in the foregoing embodiments, or equivalently replace some or all of their technical features, and it should be understood that these modifications or replacements do not deviate from the essence of the corresponding technical solutions from the scope of the technical solutions of each embodiment of the present disclosure.
Claims
1. Determining an encoding mode of a first audio frame; Determining whether the encoding mode of the first audio frame is the same as that of a second audio frame, where the second audio frame is the audio frame preceding the first audio frame; When they are not the same and the encoding mode of the first audio frame is multi-description encoding, generating third data based on first data, second data, and a first delay, where the first data is low-frequency data obtained by decimating the original audio data of the first audio frame, the second data is low-frequency data obtained by decimating the original audio data of the second audio frame, and the first delay is the encoding delay of the multi-description encoding; Encoding the third data by multi-description encoding to obtain the encoded data of the first audio frame. An audio data encoding method comprising the above steps.
2. The method further comprises: When the encoding mode of the first audio frame is not the same as that of the second audio frame and the encoding mode of the first audio frame is single-description encoding, generating sixth data based on fourth data, fifth data, and a second delay, where the fourth data is the original audio data of the first audio frame, the fifth data is the original audio data of the second audio frame, and the second delay is the encoding delay of the single-description encoding; Encoding the sixth data by single-description encoding to obtain the encoded data of the first audio frame. The method according to Claim 1 further comprises the above steps.
3. Generating third data based on first data, second data, and a first delay comprises: Cutting off sample points with a length equal to the first delay from the end of the second data to obtain seventh data; Connecting the seventh data to the front end of the first data to obtain eighth data; Deleting sample points with a length equal to the first delay from the end of the eighth data to obtain the third data. The method according to Claim 1 comprises the above steps.
4. Generating the sixth data based on the fourth data, the fifth data, and the second delay includes: Obtaining ninth data by cutting off sample points having a length of the second delay from the end of the fifth data; Obtaining tenth data by splicing the ninth data to the front end of the fourth data; The method according to claim 2, further comprising obtaining the sixth data by deleting sample points having a length of the second delay from the end of the tenth data.
5. Determining the encoding mode of the first audio frame includes: Determining whether an encoding mode switching condition is satisfied based on an encoding mode duration length and a signal type of the first audio frame, where the encoding mode duration length is a playback duration length of audio frames continuously encoded in a current encoding mode; In the case of NO, determining the encoding mode of the second audio frame as the encoding mode of the first audio frame; The method according to claim 1, further comprising determining the encoding mode of the first audio frame based on network parameters of an audio encoding data transmission network in the case of YES.
6. Determining whether an encoding mode switching condition is satisfied based on an encoding mode duration length and a signal type of the first audio frame includes: Determining whether the encoding mode duration length is greater than a threshold duration length; Determining whether a probability that the first audio frame is a voice audio frame is less than a threshold probability; Determining that the encoding mode switching condition is satisfied when the encoding mode duration length is greater than the threshold duration length and the probability that the first audio frame is a voice audio frame is less than the threshold probability; The method according to claim 5, further comprising determining that the encoding mode switching condition is not satisfied when the encoding mode duration length is less than or equal to the threshold duration length and / or the probability that the first audio frame is a voice audio frame is greater than or equal to the threshold probability.
7. Determining the encoding mode of the first audio frame based on network parameters of an audio encoding data transmission network includes: Determining a packet loss rate of the audio encoding data transmission network based on the network parameters; Determining whether the packet loss rate is equal to or higher than a threshold packet loss rate; If YES, determining that the encoding mode of the first audio frame is multi-description encoding; If NO, determining that the encoding mode of the first audio frame is single-description encoding, the method according to claim 5.
8. Determining an encoding mode of a first audio frame based on encoded data of the first audio frame; Decoding the encoded data of the first audio frame based on the encoding mode to obtain decoded data; Determining whether an encoding mode of the first audio frame is the same as an encoding mode of a second audio frame, the second audio frame being an audio frame before the first audio frame; If not the same and the encoding mode of the first audio frame is multi-description encoding, generating packet loss compensation data based on the second audio frame; Smoothing the decoded data based on delay data of the second audio frame and the packet loss compensation data to obtain playback data of the first audio frame, a method for decoding audio data.
9. The method includes: If the encoding mode of the first audio frame is not the same as the encoding mode of the second audio frame and the encoding mode of the first audio frame is single-description encoding, generating packet loss compensation data based on the second audio frame; Smoothing the decoded data based on the packet loss compensation data to obtain a smoothed result corresponding to the decoded data; Further including delaying the smoothed result based on the packet loss compensation data and a number of delay sample points to obtain playback data of the first audio frame, the number of delay sample points being a number of delay sample points of multi-description encoding, the method according to claim 8.
10. Smoothing the decoded data based on the packet loss compensation data to obtain a smoothed result corresponding to the decoded data, replacing a first sample point sequence in the decoded data with a second sample point sequence in the packet loss compensation data to obtain a first replacement result, where the first sample point sequence is a sample point sequence composed of the first several sample points of the decoded data, the first number is a difference value between a first preset number and the number of delay sample points, and the second sample point sequence is a sample point sequence composed of sample points in the packet loss compensation data whose index value ranges from the number of delay sample points to the first preset number, windowing and superimposing a third sample point sequence in the first replacement result and a fourth sample point sequence in the packet loss compensation data based on a first window function to obtain a smoothed result corresponding to the decoded data, where the third sample point sequence is a sample point sequence composed of sample points in the first replacement result whose index value ranges from the first number to the sum of the first number and a second preset number, and the fourth sample point sequence is a sample point sequence composed of sample points in the packet loss compensation data whose index value ranges from the first preset number to the sum of the first preset number and a second preset number. The method according to claim 9.
11. Delaying the smoothed result based on the packet loss compensation data and the number of delay sample points to obtain playback data for the first audio frame, obtaining a fifth sample point sequence, where the fifth sample point sequence is a sample point sequence composed of the first number of delay sample points of the packet loss compensation data, concatenating the fifth sample point sequence in front of the smoothed result to obtain a first concatenation result, Deleting the sixth sample point sequence in the first splicing result to obtain the playback data of the first audio frame, wherein the sixth sample point sequence is composed of sample points corresponding to the last delay sample point number in the first splicing result. The method according to claim 9 includes this step.
12. Smoothing the decoded data based on the delay data of the second audio frame and the packet loss compensation data to obtain the playback data of the first audio frame, which includes: Replacing a seventh sample point sequence in the decoded data with the delay data to obtain a second replacement result, wherein the seventh sample point sequence is composed of sample points corresponding to the first delay sample point number in the decoded data. Windowing and superimposing an eighth sample point sequence in the second replacement result and a ninth sample point sequence in the packet loss compensation data based on a second window function to obtain the playback data of the first audio frame. The eighth sample point sequence is composed of sample points with index values in the second replacement result ranging from the delay sample point number to the sum of the delay sample point number and a third preset number. The ninth sample point sequence is composed of sample points corresponding to the first third preset number of the packet loss compensation data. The method according to claim 9 includes these steps.
13. The method further includes: When the encoding mode of the first audio frame is the same as that of the second audio frame and the encoding mode of the first audio frame is single-description encoding, further delaying the decoded data based on the delay data of the second audio frame and the number of delay sample points to obtain the playback data of the first audio frame. The method according to claim 8 includes this step.
14. Based on the delay data of the second audio frame and the number of delay sample points, performing delay processing on the decoded data to obtain the playback data of the first audio frame, which is concatenating the delay data in front of the decoded data to obtain a second concatenation result, and deleting the tenth sample point sequence in the second concatenation result to obtain the playback data of the first audio frame, where the tenth sample point sequence is a sample point sequence composed of the last number of delay sample points of the second concatenation result. The method according to claim 13 **Claim 15** A determination unit for determining the encoding mode of a first audio frame, A determination unit for determining whether the encoding mode of the first audio frame is the same as the encoding mode of a second audio frame, where the second audio frame is determined to be the audio frame before the first audio frame, A generation unit for generating third data based on first data, second data, and a first delay when the encoding mode of the first audio frame is not the same as the encoding mode of the second audio frame and the encoding mode of the first audio frame is multi-description encoding. The first data is low-frequency data obtained by dividing the original audio data of the first audio frame, the second data is low-frequency data obtained by dividing the original audio data of the second audio frame, and the first delay is the encoding delay of the multi-description encoding. The generation unit An encoding apparatus for audio data, comprising an encoding unit for multi-description encoding the third data to obtain the encoded data of the first audio frame **Claim 16** A determination unit for determining the encoding mode of a first audio frame based on the encoded data of the first audio frame, A decoding unit for decoding the encoded data of the first audio frame based on the encoding mode to obtain decoded data A determination unit for determining whether the encoding mode of the first audio frame is the same as the encoding mode of the second audio frame, the second audio frame being the audio frame preceding the first audio frame, A decoding apparatus for audio data, comprising: a processing unit that, when the encoding mode of the first audio frame is not the same as the encoding mode of the second audio frame and the encoding mode of the first audio frame is multi-description encoding, generates packet loss compensation data based on the second audio frame, and smoothes the decoded data based on the delay data of the second audio frame and the packet loss compensation data to obtain playback data of the first audio frame.
17. An electronic device including a memory and a processor, the memory being used for storing a computer program, and the processor being used for realizing the audio data encoding method according to any one of claims 1 to 7 or the audio data decoding method according to any one of claims 8 to 14 when executing the computer program.
18. A computer-readable storage medium storing a computer program that, when executed by a computer device, realizes the audio data encoding method according to any one of claims 1 to 7 or the audio data decoding method according to any one of claims 8 to 14.
19. A computer program product including a computer program that, when executed by a processor, realizes the audio data encoding method according to any one of claims 1 to 7 or the audio data decoding method according to any one of claims 8 to 14.
Citation Information
Patent Citations
Low-latency acoustic coding that repeatedly performs predictive and transformative coding.
JP2014505272A
audio processing system
JP2016514858A
Method for encoding and decoding audio content using encoder, decoder and parameters for enhancing concealment
JP2017529565A
Method and program product for organizing data into packets
US20030099236A1