Audio signal processing method and apparatus, storage medium, and computer program

The audio signal processing method optimizes subband partitioning and spectral envelope shaping to enhance coding effectiveness and compression efficiency, addressing the challenge of maintaining audio quality in Bluetooth interconnection scenarios.

JP7870399B2Active Publication Date: 2026-06-04HUAWEI TECH CO LTD

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
HUAWEI TECH CO LTD
Filing Date
2023-05-04
Publication Date
2026-06-04

Smart Images

  • Figure 0007870399000021
    Figure 0007870399000021
  • Figure 0007870399000022
    Figure 0007870399000022
  • Figure 0007870399000023
    Figure 0007870399000023
Patent Text Reader

Abstract

An audio signal processing method and apparatus (1100), a storage medium, and a computer program product are provided, belonging to the field of audio encoding / decoding. An optimal subband decomposition scheme is selected from multiple subband decomposition schemes based on audio signal characteristics. In other words, the subband decomposition scheme has signal adaptive characteristics and can adapt to the coding bit rate of the audio signal to improve interference resistance. Specifically, the audio signal is separately divided based on multiple subband decomposition schemes, and a total scale factor corresponding to each subband decomposition scheme is determined based on the spectral values ​​of the audio signal in the subbands obtained through the division, the bandwidth of each subband, and the coding bit rate of the audio signal. An optimal target subband decomposition scheme is selected based on the total scale factor to obtain an optimal subband set. Then, spectral envelope shaping is performed based on the scale factor of each subband in the optimal subband set, thereby improving coding efficiency and compression efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application claims priority to Chinese Patent Application No. 202210894324.9, filed on July 27, 2022, and Chinese Patent Application No. 202211139940.X, filed on September 19, 2022, both entitled "Audio Signal Processing Method and Apparatus, Storage Medium, and Computer Program Product", and incorporates both of them herein by reference in their entirety.

[0002] This application relates to the field of audio encoding / decoding, and in particular, to an audio signal processing method and apparatus, a storage medium, and a computer program product.

Background Art

[0003] As the quality of life improves, people's demands for high-quality audio are increasing. In order to better transmit audio signals with limited bandwidth, it is usually necessary to perform data compression on the audio signals on the encoder side to obtain a bitstream. Then, the bitstream is transmitted to the decoder side. The decoder side decodes the received bitstream to reconstruct the audio signal. The reconstructed audio signal is used for playback. However, in the process of compressing the audio signal, the audio quality of the audio signal may be affected. Therefore, how to improve the compression efficiency of the audio signal while ensuring the audio quality of the audio signal has become a technical problem that needs to be urgently solved.

Summary of the Invention

[0004] This application provides an audio signal processing method and apparatus, a storage medium, and a computer program product for improving the coding effect and compression efficiency. The technical solution is as follows.

[0005] According to the first aspect, an audio signal processing method is provided. This method is: The method includes performing subband partitioning on an audio signal separately based on multiple subband partitioning schemes and cutoff subbands corresponding to the multiple subband partitioning schemes to obtain multiple candidate subband sets, wherein the multiple candidate subband sets correspond one-to-one with the multiple subband partitioning schemes, each candidate subband set contains multiple subbands, and the total scale value of each candidate subband set is determined based on the spectral values ​​of the audio signal in the subbands included in the candidate subband set, the encoded bitrate of the audio signal, and the subband bandwidth of the subbands included in the candidate subband set; selecting one candidate subband set from the multiple candidate subband sets as a target subband set based on the total scale value of each candidate subband set, wherein each subband included in the target subband set has a scale factor used to shape the spectral envelope of the audio signal.

[0006] In this application, the optimal subband division method is selected from multiple subband division methods based on the characteristics of the audio signal. In other words, the subband division method has signal adaptive characteristics and can improve interference immunity by adapting to the encoding bitrate of the audio signal. Specifically, the audio signal is divided separately based on multiple subband division methods, and a total scale value corresponding to each subband division method is determined based on the spectral values ​​of the audio signal in the subbands obtained through division, the bandwidth of each subband, and the encoding bitrate of the audio signal. An optimal subband set is obtained by selecting the optimal target subband division method based on this total scale value. Subsequently, spectral envelope shaping is performed based on the scale factor of each subband in the optimal subband set, thereby improving coding effectiveness and compression efficiency.

[0007] Optionally, you can select one candidate subbandset as the target subbandset from among multiple candidate subbandsets based on the total scale value of each candidate subbandset. Among multiple candidate subband sets, the candidate subband set with the smallest total scale value is selected as the target subband set. This includes the following.

[0008] Optionally, the total scale value of each candidate subband set may be determined based on the spectral values ​​of the audio signals within the subbands included in the candidate subband set, the encoded bitrate of the audio signals, and the subband bandwidth of the subbands included in the candidate subband set. For a first candidate subband set among multiple candidate subband sets, the scale factor of each subband included in the first candidate subband set is determined based on the spectral values ​​of the audio signals within the subbands included in the first candidate subband set, and the first candidate subband set is one of the multiple candidate subband sets. The total scale value of the first candidate subband set is determined based on the encoding bitrate of the audio signal and the scale factor and subband bandwidth of each subband included in the first candidate subband set. This includes the following.

[0009] Optionally, the scale factor of each subband included in the first candidate subband set may be determined based on the spectral values ​​of the audio signals within the subbands included in the first candidate subband set. For a first subband included in the first candidate subband set, obtain the maximum value among the absolute values ​​of all spectral values ​​of the audio signal within the first subband, and determine that the first subband is any of the subbands in the first candidate subband set. The scale factor of the first subband is determined based on the maximum value. This includes the following.

[0010] Optionally, the encoding bitrate of the audio signal is not lower than a first bitrate threshold, and / or the energy concentration of the audio signal is greater than a concentration threshold.

[0011] Determining the total scale value of the first candidate subband set based on the encoded bitrate of the audio signal and the scale factor and subband bandwidth of each subband included in the first candidate subband set is: An energy smoothing reference value is determined based on the encoding bitrate of the audio signal and a second bitrate threshold. Based on the energy smoothing reference value and the scale factor and subband bandwidth of each subband included in the first candidate subband set, the total energy value of each subband included in the first candidate subband set is determined. The total energy values ​​of the subbands included in the first candidate subband set are summed to obtain the total scale value of the first candidate subband set. This includes the following.

[0012] Optionally, the total energy value of each subband in the first candidate subband set can be determined based on an energy smoothing reference value and the scale factor and subband bandwidth of each subband in the first candidate subband set. For the first subband included in the first candidate subband set, the larger of the scale factor and the energy smoothing reference value of the first subband is determined as the reference scale value of the first subband, and the first subband is one of the subbands in the first candidate subband set. The product of the reference scale value of the first subband and the subband bandwidth of the first subband is determined as the total energy value of the first subband. This includes the following.

[0013] Optionally, the encoding bitrate of the audio signal is lower than the first bitrate threshold, and the energy concentration of the audio signal is not greater than the concentration threshold.

[0014] Determining the total scale value of the first candidate subband set based on the encoded bitrate of the audio signal and the scale factor and subband bandwidth of each subband included in the first candidate subband set is: An energy smoothing reference value is determined based on the encoding bitrate of the audio signal and a second bitrate threshold. Based on the energy smoothing reference value and the scale factor of each subband included in the first candidate subband set, the scale difference value of the subbands included in the first candidate subband set is determined, and this scale difference value represents the difference between the scale factor of the corresponding subband and the scale factor of the adjacent subband of the corresponding subband. The total scale value of the first candidate subband set is determined based on the scale difference values ​​of the subbands included in the first candidate subband set and the subband bandwidth of each subband. This includes the following.

[0015] Optionally, the scale difference values ​​of the subbands included in the first candidate subband set can be determined based on the energy smoothing reference value and the scale factor of each subband included in the first candidate subband set. For a first subband included in the first candidate subband set, the first smoothing value, second smoothing value, and third smoothing value of the first subband are determined based on the energy smoothing reference value, the scale factor of the first subband, and the scale factor of the adjacent subbands of the first subband, and the first subband is one of the subbands in the first candidate subband set. Based on the first smoothing value, second smoothing value, and third smoothing value of the first subband, the scale difference value of the first subband is determined. This includes the following.

[0016] Optionally, determining the first smoothing value, the second smoothing value, and the third smoothing value of the first sub-band based on an energy smoothing reference value, a scale factor of the first sub-band, and a scale factor of an adjacent sub-band of the first sub-band is when the first sub-band is the first sub-band within the first candidate sub-band set, determining the larger value between the scale factor of the first sub-band and the energy smoothing reference value as the first smoothing value of the first sub-band; when the first sub-band is not the first sub-band within the first candidate sub-band set, determining the larger value between the scale factor of the previous sub-band adjacent to the first sub-band and the energy smoothing reference value as the first smoothing value of the first sub-band, determining the larger value between the scale factor of the first sub-band and the energy smoothing reference value as the second smoothing value of the first sub-band, when the first sub-band is the last sub-band within the first candidate sub-band set, determining the larger value between the scale factor of the first sub-band and the energy smoothing reference value as the third smoothing value of the first sub-band; when the first sub-band is not the last sub-band within the first candidate sub-band set, determining the larger value between the scale factor of the next sub-band adjacent to the first sub-band and the energy smoothing reference value as the third smoothing value of the first sub-band, including.

[0017] Optionally, determining a scale difference value of the first sub-band based on the first smoothing value, the second smoothing value, and the third smoothing value of the first sub-band is for the first sub-band included in the first candidate sub-band set, determining a first difference value and a second difference value of the first sub-band, where the first difference value is the absolute value of the difference between the first smoothing value and the second smoothing value of the first sub-band, and the second difference value is the absolute value of the difference between the second smoothing value and the third smoothing value of the first sub-band, and the first sub-band is any sub-band within the first candidate sub-band set, Determining a scale difference value of a first sub-band based on a first difference value and a second difference value of the first sub-band including this

[0018] Optionally, determining a total scale value of a first candidate sub-band set based on scale difference values of sub-bands included in the first candidate sub-band set and sub-band bandwidths of each sub-band Determining a smoothing weighting coefficient for each sub-band included in the first candidate sub-band set based on the number of sub-bands included in the first candidate sub-band set and sub-band bandwidths of each sub-band Adding up the smoothing weighting coefficients of the sub-bands included in the first candidate sub-band set to obtain a total smoothing weighting coefficient of the first candidate sub-band set Multiplying the scale difference values of the sub-bands included in the first candidate sub-band set by the smoothing weighting coefficients to obtain weighted scale difference values of the sub-bands included in the first candidate sub-band set Adding up the weighted scale difference values of the sub-bands included in the first candidate sub-band set to obtain a total scale value of the first candidate sub-band set Dividing the total scale value by the total smoothing weighting coefficient of the first candidate sub-band set to obtain a total scale value of the first candidate sub-band set including this

[0019] Optionally, the method further includes When the coding bit rate of the audio signal is lower than a first bit rate threshold, performing bandwidth detection on the spectrum of the audio signal to obtain a cut-off frequency of the audio signal Determining a cut-off sub-band corresponding to each of a plurality of sub-band division methods based on the cut-off frequency including this

[0020] Optionally, the method further includes If the encoding bitrate of the audio signal is not lower than the first bitrate threshold, the last subband indicated by each of the multiple subband division schemes is determined as the cutoff subband corresponding to each subband division scheme. This includes the following: Optionally, the method further, Perform feature analysis on the spectrum of the audio signal to obtain the feature analysis results. Based on the feature analysis results and the encoding bitrate of the audio signal, multiple subband partitioning schemes are determined from a list of candidate subband partitioning schemes. This includes the following.

[0021] Optionally, the feature analysis results include a subjective signal flag or an objective signal flag, where the subjective signal flag indicates that the energy concentration of the audio signal is not greater than the concentration threshold, and the objective signal flag indicates that the energy concentration of the audio signal is greater than the concentration threshold.

[0022] Optionally, the audio signal frame length may be 10 milliseconds and the sampling rate 88.2 kHz or 96 kHz, or the audio signal frame length may be 5 milliseconds and the sampling rate 88.2 kHz or 96 kHz, or the audio signal frame length may be 10 milliseconds and the sampling rate 44.1 kHz or 48 kHz.

[0023] Determining multiple subband partitioning schemes from a pool of candidate schemes based on feature analysis results and the encoding bitrate of the audio signal is possible. When the encoded bitrate of the audio signal is lower than the first bitrate threshold, and the feature analysis result includes a subjective signal flag, the first group of subband partitioning schemes among the multiple candidate subband partitioning schemes is determined to be one of the multiple subband partitioning schemes. This includes the following.

[0024] The subband division scheme for the first group is as follows: { {0,1,2,3,4,6,8,10,13,16,20,24,28,33,38,45,52,61,70,79,88,100,112,127,142,160,178,196,217,238,259,280,480}, {0,1,2,3,5,7,9,12,15,18,22,26,30,35,41,48,56,65,74,84,94,106,118,134,150,166,184,202,220,240,260,280,480}, {0,1,2,3,4,5,7,9,11,14,17,21,25,29,34,40,46,52,60,68,76,86,98,110,126,144,162,180,200,224,250,280,480}, {0,2,4,6,8,12,16,21,26,31,36,41,46,51,56,61,66,71,77,83,89,95,103,111,121,131,147,163,179,203,240,280,480}, {0,1,2,3,5,7,9,12,15,19,23,27,32,37,43,49,57,66,76,86,98,110,125,140,158,176,194,216,238,264,290,320,480}, {0,1,2,3,5,7,10,13,17,21,25,30,35,41,47,54,62,70,80,90,102,114,130,146,162,180,198,218,240,264,290,320,480}, {0,1,2,4,6,8,11,14,18,22,26,30,36,42,50,58,66,76,88,100,112,128,144,160,182,204,226,256,286,316,352,400,480}, {0,1,2,4,6,8,11,14,18,22,26,30,36,42,50,58,68,78,90,102,116,132,148,166,186,208,234,262,292,324,360,400,480} }。

[0025] Optionally, the audio signal frame length may be 10 milliseconds and the sampling rate 88.2 kHz or 96 kHz, or the audio signal frame length may be 5 milliseconds and the sampling rate 88.2 kHz or 96 kHz, or the audio signal frame length may be 10 milliseconds and the sampling rate 44.1 kHz or 48 kHz.

[0026] Determining multiple subband partitioning schemes from a pool of candidate schemes based on feature analysis results and the encoding bitrate of the audio signal is possible. If the encoded bitrate of the audio signal is not lower than the first bitrate threshold, and / or the feature analysis result includes an objective signal flag, then the second group of subband partitioning schemes among the multiple candidate subband partitioning schemes is determined to be one of the multiple subband partitioning schemes. This includes the following.

[0027] The subband division scheme for the second group is as follows: { {0,1,2,3,4,5,6,7,8,10,12,14,16,19,22,26,30,35,40,45,50,57,64,73,82,92,102,112,124,136,148,160,480}, {0,1,2,3,4,5,7,9,11,13,15,18,21,24,28,33,38,44,50,57,64,73,82,93,104,116,128,140,155,170,185,200,480}, {0,1,2,3,4,6,8,10,13,16,20,24,28,33,38,45,52,61,70,79,88,100,112,127,142,160,178,196,217,238,259,280,480}, {0,1,2,4,6,10,14,18,22,26,30,34,42,50,58,66,74,84,96,108,120,136,152,168,192,216,240,272,304,336,376,424,480}, {0,1,2,4,6,10,14,18,26,34,42,50,62,74,86,98,112,128,144,160,176,196,216,236,256,280,304,328,352,384,416,448,480}, {0,80,92,104,112,120,128,136,144,148,152,156,160,164,168,172,176,180,184,188,192,196,200,208,216,224,232,240,248,256,268,280,480}, {0,200,212,224,232,240,248,256,264,268,272,276,280,284,288,292,296,300,304,308,312,316,320,328,336,344,352,360,368,376,388,400,480}, {0,320,332,344,356,364,372,380,384,388,392,396,400,404,408,412,416,420,424,428,432,436,440,444,448,452,456,460,464,468,472,476,480} }.

[0028] Optionally, the audio signal frame length is 5 milliseconds, and the sampling rate is 44.1 kHz or 48 kHz.

[0029] Determining multiple subband partitioning schemes from a pool of candidate schemes based on feature analysis results and the encoding bitrate of the audio signal is possible. If the encoding bitrate of the audio signal is lower than the first bitrate threshold, and the feature analysis results include a subjective signal flag, then the third group of subband partitioning schemes among the multiple candidate subband partitioning schemes is selected as one of the multiple subband partitioning schemes. This includes the following.

[0030] The subband division scheme for the third group is as follows: { {0,1,2,3,4,5,6,7,8,9,10,12,14,16,19,22,26,30,35,39,44,50,56,63,71,80,89,98,108,119,129,140,240}, {0,1,2,3,4,5,6,7,8,9,11,13,15,17,20,24,28,32,37,42,47,53,59,67,75,83,92,101,110,120,130,140,240}, {0,1,2,3,4,5,6,7,8,9,10,11,12,14,17,20,23,26,30,34,38,43,49,55,63,72,81,90,100,112,125,140,240}, {0,1,2,3,4,6,8,10,13,15,18,20,23,25,28,30,33,35,38,41,44,47,51,55,60,65,73,81,89,101,120,140,240}, {0,1,2,3,4,5,6,7,9,11,13,14,16,18,21,24,28,33,38,43,49,55,62,70,79,88,97,108,119,132,145,160,240}, {0,1,2,3,4,5,6,7,8,10,12,14,17,20,23,27,31,35,40,45,51,57,65,73,81,90,99,109,120,132,145,160,240}, {0,1,2,3,4,5,6,7,9,11,13,15,18,21,25,29,33,38,44,50,56,64,72,80,91,102,113,128,143,158,176,200,240}, {0,1,2,3,4,5,6,7,9,11,13,15,18,21,25,29,34,39,45,51,58,66,74,83,93,104,117,131,146,162,180,200,240} }.

[0031] Optionally, the audio signal frame length is 5 milliseconds, and the sampling rate is 44.1 kHz or 48 kHz.

[0032] Determining multiple subband partitioning schemes from a pool of candidate schemes based on feature analysis results and the encoding bitrate of the audio signal is possible. If the encoding bitrate of the audio signal is not lower than the first bitrate threshold, and / or the feature analysis result includes an objective signal flag, then the fourth group of subband partitioning schemes among the multiple candidate subband partitioning schemes is determined to be one of the multiple subband partitioning schemes. This includes the following.

[0033] The subband division scheme for the fourth group is as follows: { {0,1,2,3,4,5,6,7,8,9,10,11,12,13,14,15,16,17,18,19,20,21,22,23,24,26,28,30,32,34,37,40,120}, {0,1,2,3,4,5,6,7,8,9,10,11,12,13,14,15,16,18,20,22,24,26,28,30,32,34,36,38,41,44,47,50,120}, {0,1,2,3,4,5,6,7,8,9,10,11,12,14,16,18,20,22,24,26,28,31,34,37,40,44,48,52,56,60,65,70,120}, {0,1,2,3,4,5,6,7,8,9,10,11,12,13,15,17,19,21,24,27,30,34,38,42,48,54,60,68,76,84,94,106,120}, {0,1,2,3,4,5,6,7,8,10,12,14,16,19,22,25,28,32,36,40,44,49,54,59,64,70,76,82,88,96,104,112,120}, {0,20,23,26,28,30,32,34,36,37,38,39,40,41,42,43,44,45,46,47,48,49,50,52,54,56,58,60,62,64,67,70,120}, {0,50,53,56,58,60,62,64,66,67,68,69,70,71,72,73,74,75,76,77,78,79,80,82,84,86,88,90,92,94,97,100,120}, {0,80,83,86,89,91,93,95,96,97,98,99,100,101,102,103,104,105,106,107,108,109,110,111,112,113,114,115,116,117,118,119,120} }.

[0034] Optionally, the audio signal is a dual-channel signal.

[0035] This method further, Based on the scale factor and subband bandwidth of each subband included in the target subband set, a first total scale value is determined. Perform mid / side stereo conversion coding on the spectrum of a dual-channel signal to obtain the converted spectrum of the dual-channel signal. Based on the transformed spectral values ​​of the dual-channel signals within each subband included in the target subband set, the transformed scale factor of each subband in the target subband set is determined. Based on the transformed scale factor and subband bandwidth of each subband included in the target subband set, a second total scale value is determined. If the first total scale value is not greater than the second total scale value, the dual-channel signal is determined to be the signal to be encoded. This includes the following.

[0036] Optionally, this method further, The converted dual-channel signal is determined to be the signal to be encoded if the first total scale value is greater than the second total scale value, and the encoding bitrate of the audio signal is not lower than the first bitrate threshold, and / or the energy concentration of the audio signal is greater than the concentration threshold. This includes the following.

[0037] Optionally, the scale factor includes the left channel scale factor and the right channel scale factor.

[0038] This method further, If the first total scale value is greater than the second total scale value, the encoded bitrate of the audio signal is lower than the first bitrate threshold, and the energy concentration of the audio signal is not greater than the concentration threshold, then, based on the left channel scale factor and right channel scale factor of each subband included in the target subband set, the left channel of each subband included in the target subband set channel Scale factor and right channel Determine the difference value between the scale factor and the other factor. Based on the initial frequency and cutoff frequency of each subband included in the target subband set, the subband center frequency of each subband included in the target subband set is determined. Within the target subband set, left-hand values ​​greater than the difference threshold. channel Scale factor and right channel A dual-channel signal is determined to be the signal to be encoded if there is at least one subband that has a difference value between it and the scale factor and has a subband center frequency within a first range. This includes the following.

[0039] Optionally, this method further, If at least one subband is not present in the target subband set, the converted dual-channel signal is determined to be the signal to be encoded. This includes the following.

[0040] According to a second embodiment, an audio signal processing device is provided. The audio signal processing device has the function of implementing the audio signal processing method of the first embodiment. The audio signal processing device includes one or more modules, the one or more modules being configured to implement the audio signal processing method of the first embodiment.

[0041] According to a third aspect, an audio signal processing device is provided. The audio signal processing device includes a processor and memory. The memory is configured to store a program used to perform the audio signal processing method of the first aspect, and to store data used to perform the audio signal processing method of the first aspect. The processor is configured to execute the program stored in memory. The audio signal processing device may further include a communication bus, which is configured to establish a connection between the processor and the memory.

[0042] According to a fourth aspect, a computer-readable storage medium is provided. The computer-readable storage medium stores instructions, and when the instructions are executed on a computer, the computer is enabled to execute the audio signal processing method according to the first aspect.

[0043] According to the fifth aspect, a computer program product containing instructions is provided. When the computer program product runs on a computer, the computer is enabled to perform the audio signal processing method according to the first aspect.

[0044] The technical effects achieved in the second, third, fourth, and fifth embodiments are the same as those achieved by the corresponding technical means in the first embodiment. Further details will not be described here. [Brief explanation of the drawing]

[0045] [Figure 1] This is a diagram of a Bluetooth interconnection scenario according to one embodiment of this application. [Figure 2] This is a diagram of a system architecture related to an audio signal processing method according to one embodiment of this application. [Figure 3A] Figures 3A and 3B illustrate an overall audio coding / decoding framework according to one embodiment of this application. [Figure 3B] Figures 3A and 3B illustrate an overall audio coding / decoding framework according to one embodiment of this application. [Figure 4] This is a diagram showing the configuration of an electronic device according to one embodiment of this application. [Figure 5] This is a flowchart of an audio signal processing method according to one embodiment of this application. [Figure 6] This is a diagram showing the relationship between subbands and initial frequencies in a first group subband division scheme according to one embodiment of this application. [Figure 7] This is a diagram showing the relationship between subbands and initial frequencies in a second group subband division scheme according to one embodiment of this application. [Figure 8] This is a diagram showing the relationship between subbands and initial frequencies in a third group subband division scheme according to one embodiment of this application. [Figure 9] This is a diagram showing the relationship between subbands and initial frequencies in a subband division scheme of the fourth group according to one embodiment of this application. [Figure 10] This is a flowchart of a method for determining whether an MS conversion has gain, according to one embodiment of this application. [Figure 11] This is a diagram showing the configuration of an audio signal processing device according to one embodiment of this application. [Modes for carrying out the invention]

[0046] To further clarify the purpose, technical solutions, and advantages of the embodiments of this application, the implementation of this application will be described in detail below with reference to the accompanying drawings.

[0047] First, we will describe the implementation environment and background knowledge related to the embodiments of this application.

[0048] As wireless Bluetooth® devices such as true wireless stereo (TWS) headsets, smart speakers, and smartwatches become widespread and used in people's daily lives, the demand for high-quality audio playback experiences in various scenarios is becoming increasingly urgent, especially in environments where Bluetooth signals are vulnerable to interference, such as subways, airports, and train stations. This is because, in Bluetooth interconnection scenarios, the Bluetooth channel connecting the audio transmitting and receiving devices limits the size of the transmitted data. The audio encoder in the audio transmitting device performs data compression on the audio signal before sending it to the audio receiving device. The compressed audio signal can only be played back after it has been decoded by the audio decoder in the audio receiving device. It is also clear that the proliferation of wireless Bluetooth devices has facilitated the flourishing of various audio codecs.

[0049] Currently, Bluetooth audio codecs include sub-band coding (SBC), advanced audio coding (AAC), aptX series encoders, low-latency high-definition audio codec (LHDC), low-energy low-latency LC3 audio codec, and LC3plus.

[0050] It should be understood that the audio provided in the embodiments of this application Signal processing The method can be applied to audio transmitting devices (i.e., encoders) and audio receiving devices (i.e., decoders) in Bluetooth interconnection scenarios.

[0051] Figure 1 is a diagram of a Bluetooth interconnection scenario according to one embodiment of this application. Please refer to Figure 1. The audio transmitting device in the Bluetooth interconnection scenario may be a mobile phone, computer, tablet computer, or similar. The computer may be a notebook computer or a desktop computer, and the tablet computer may be a handheld tablet computer or an in-car tablet computer, etc. The audio receiving device in the Bluetooth interconnection scenario may be a TWS headset, smart speaker, wireless headphones, wireless neckband headset, smartwatch, smart glasses, smart in-car device, or similar. In some other embodiments, the audio receiving device in the Bluetooth interconnection scenario may instead be a mobile phone, computer, tablet computer, or similar.

[0052] In addition to the Bluetooth interconnection scenario, the audio provided in the embodiments of this application is also available. Signal processing The method may also be applicable to other device interconnection scenarios. In other words, the system architectures and service scenarios described in embodiments of this application are intended to further illustrate the technical solutions in embodiments of this application and do not constitute a limitation on the technical solutions provided in embodiments of this application. As will be apparent to those skilled in the art, the technical solutions provided in embodiments of this application are also applicable to similar technical problems as system architectures evolve and new service scenarios emerge.

[0053] Figure 2 is a diagram of a system architecture relating to an audio signal processing method according to one embodiment of this application. Please refer to Figure 2. The system includes an encoder side and a decoder side. The encoder side includes an input module, an encoding module, and a transmission module. The decoder side includes a receiving module, an input module, a decoding module, and a playback module.

[0054] On the encoder side, the user selects one of two encoding modes, a low-latency encoding mode and a high-quality encoding mode, based on the usage scenario. The encoding frame lengths for these two encoding modes are 5ms and 10ms, respectively. For example, if the usage scenario is playing games, live broadcasting, or making phone calls, the user can select the low-latency encoding mode, or if the usage scenario is enjoying music through a headset or speaker, the user can select the high-quality encoding mode. The user also needs to provide the encoder with the audio signal to be encoded (pulse code modulation (PCM) data shown in Figure 2). In addition, the user needs to set the target bitrate of the bitstream obtained through encoding, i.e., the encoding bitrate of the audio signal. A higher target bitrate results in better sound quality but poorer interference immunity of the bitstream in the short-distance transmission process. A lower target bitrate results in poorer sound quality but better interference immunity of the bitstream in the short-distance transmission process. Simply put, the encoder's input module receives the encoded frame length, encoded bitrate, and the audio signal to be encoded, which are submitted by the user.

[0055] The input module on the encoder side inputs the data submitted by the user into the frequency domain encoder of the encoding module.

[0056] The frequency domain encoder in the encoding module performs encoding based on the received data to obtain a bitstream. The frequency domain encoder analyzes the audio signal to be encoded to obtain signal characteristics (including mono / dual channel signals, stable / unstable signals, full-bandwidth / narrow-bandwidth signals, and subjective / objective signals). Based on the signal characteristics and bitrate level (i.e., encoding bitrate), the audio signal enters the corresponding encoding processing submodule. This encoding processing submodule encodes the audio signal and packages the bitstream packet header (including sampling rate, channel number, encoding mode, and frame length) to finally obtain the bitstream.

[0057] The encoder-side transmitting module transmits the bitstream to the decoder. Optionally, the transmitting module may be the short-range transmitting module shown in Figure 2, or another type of transmitting module. This is not limited to the embodiments of this application.

[0058] On the decoder side, after receiving the bitstream, the decoder's receiving module notifies the decoder's input module to send the bitstream to the decoding module's frequency domain decoder and to obtain the configured bit depth, configured channel decoding mode, or similar. Optionally, the receiving module may be the short-range receiving module shown in Figure 2, or another type of receiving module. This is not limited to the embodiments of this application.

[0059] The decoder's input module inputs acquired information, such as bit depth and audio channel decoding mode, to the decoding module's frequency domain decoder.

[0060] The frequency domain decoder in the decoding module decodes the bitstream based on the bit depth and channel decoding mode to obtain the necessary audio data (PCM data shown in Figure 2), and sends the obtained audio data to the playback module. The playback module then plays the audio. The audio channel decoding mode indicates the channels that need to be decoded.

[0061] Figures 3A and 3B illustrate an overall audio encoding / decoding framework according to one embodiment of this application. Please refer to Figures 3A and 3B. The encoding procedure on the encoder side includes the following steps:

[0062] (1) PCM input module PCM data is input. The PCM data can be mono-channel or dual-channel data, and the bit depth can be 16 bits, 24 bits, 32 bits floating-point, or 32 bits fixed-point. Optionally, the PCM input module converts the input PCM data to the same bit depth, for example, 24 bits, performs deinterleaving on the PCM data, and then places the deinterleaved PCM data on the left and right channels.

[0063] (2) Low-latency analysis window addition module and modified discrete cosine transform (MDCT) transformation module A low-latency analysis window is added to the PCM data processed in step (1), and an MDCT transformation is performed to obtain spectral data in the MDCT domain. The window is added in a way that prevents spectral leakage.

[0064] (3) MDCT Domain Signal Analysis Module and Adaptive Bandwidth Detection Module The MDCT domain signal analysis module is effective in full bitrate scenarios, while the adaptive bandwidth detection module is activated at low bitrates (e.g., bitrate <150kbps / channel). First, bandwidth detection is performed on the spectral data in the MDCT domain obtained in step (2) to acquire the cutoff frequency or effective bandwidth. Next, signal analysis is performed on the spectral data within the effective bandwidth, i.e., whether the frequency distribution is concentrated or flat is analyzed to obtain energy concentration, and based on this energy concentration, a flag is obtained indicating whether the audio signal to be encoded is an objective or subjective signal (flag 1 for objective signals and flag 0 for subjective signals). If the audio signal is an objective signal, spectral noise shaping (SNS) and MDCT spectral smoothing are not performed on the scale factor at low bitrates because this reduces the encoding effect of the objective signal. Then, based on the bandwidth detection result and the subjective signal flag and objective signal flag, it is determined whether a subband cutoff operation should be performed in the MDCT domain. If the audio signal is an objective signal, the subband cutoff operation is not performed. If the audio signal is a subjective signal and the bandwidth detection result is identified as 0 (within the entire bandwidth), the subband cutoff operation is determined based on the bitrate. If the audio signal is a subjective signal and the bandwidth detection result is not identified as 0 (i.e., the bandwidth is less than half of the sampling rate limiting bandwidth), the subband cutoff operation is determined based on the bandwidth detection result.

[0065] (4) Subband division selection and scale factor calculation module Based on the bitrate level, the subjective and objective signal flags obtained in step (3), and the cutoff frequency, the optimal subband division scheme is selected from several subband division schemes, and the total number of subbands for encoding the audio signal is obtained. In addition, the spectral envelope is obtained through calculation, i.e., the scale factor corresponding to the selected subband division scheme is calculated.

[0066] (5) MS Channel conversion module For dual-channel PCM data, a joint coding determination is made based on the scale factor calculated in step (4), that is, it is determined whether MS channel conversion should be performed on the left channel data and the right channel data.

[0067] (6) Spectral smoothing module and scale factor-based spectral noise shaping module The spectral smoothing module performs MDCT spectral smoothing based on a low bitrate setting (e.g., bitrate < 150kbps / channel), and the spectral noise shaping module performs spectral noise shaping on the data to which spectral smoothing is performed, based on a scale factor, to obtain an adjustment factor used to quantize the spectral values ​​of the audio signal. The low bitrate setting is controlled by the low bitrate determination module. When the low bitrate setting is not met, spectral smoothing and spectral noise shaping do not need to be performed.

[0068] (7) Scale factor coding module For the scale factors of multiple subbands, differential coding or entropy coding is performed based on the distribution of those scale factors.

[0069] (8) Bit allocation, MDCT spectral quantization, and entropy coding module Based on the scale factor obtained in step (4) and the adjustment factor obtained in step (6), the encoding is controlled to be a constant bit rate (CBR) encoding mode according to the coarse and fine bit allocation strategies, and quantization and entropy coding are performed on the MDCT spectral values.

[0070] (9) Residual coding module If the bit consumption in step (8) does not reach the target bits, further importance sorting is performed on the unencoded subbands, and bits are preferably allocated to encoding the MDCT spectral values ​​of important subbands.

[0071] (10) Stream packet header information packaging module Packet header information includes audio sampling rate (e.g., 44.1kHz / 48kHz / 88.2kHz / 96kHz), channel information (e.g., mono-channel and dual-channel), encoded frame length (e.g., 5ms and 10ms), and encoding mode (e.g., time-domain mode, frequency-domain mode, time-domain-frequency-domain mode, or frequency-domain-time-domain mode).

[0072] (11) Bitstream transmission module The bitstream includes a packet header, side information, and payload. The packet header carries packet header information, as described in step (10). The side information includes, for example, the encoded bitstream of the scale factor, information about the selected subband partitioning scheme, cutoff frequency information, a low bitrate flag, joint coding decision information (i.e., the MS transformation flag), and information such as the quantization step. The payload includes the encoded bitstream and the residual encoded bitstream of the MDCT spectrum.

[0073] The decoding procedure on the decoder side includes the following steps:

[0074] (1) Stream packet header information parsing module The stream packet header information parsing module parses packet header information from the received bitstream, where the packet header information includes information such as the sampling rate, channel information, encoded frame length, and encoding mode of the audio signal. The module then obtains the encoded bitrate through calculations based on the bitstream size, sampling rate, and encoded frame length, i.e., it obtains bitrate level information.

[0075] (2) Scale factor decoding module The scale factor decoding module decodes side information from the bitstream. Side information includes, for example, information about the selected subband partitioning scheme, cutoff frequency information, low bitrate flag, joint coding decision information, quantization step, and subband scale factor.

[0076] (3) Scale factor-based spectral noise shaping module At low bitrates (e.g., encoding bitrates lower than 300kbps, i.e., 150kbps / channel), further spectral noise shaping based on the scale factor is necessary to obtain the adjustment factor. The adjustment factor is, decrypt This is used to dequantize the resulting spectral values. The low bitrate setting is controlled by the low bitrate determination module. When the low bitrate setting is not met, spectral noise shaping does not need to be performed.

[0077] (4) MDCT spectral decoding module and residual decoding module The MDCT spectral decoding module decodes the MDCT spectral data in the decoded bitstream based on information about the subband partitioning scheme, quantization step information, and the scale factor obtained in step (2). At low bitrate levels, hole padding is performed, and if there are still bits remaining to be obtained through the computation, the residual decoding module performs residual decoding to obtain MDCT spectral data of another subband to obtain the final MDCT spectral data.

[0078] (5) LR Channel Conversion Module Based on the side information obtained in step (2), if it is determined, according to the joint coding decision information, that a dual-channel joint coding mode (for example, an coding bitrate of 300kbps / channel or higher and a sampling rate higher than 88.2kHz) should be used instead of a decoding low-energy mode, then an LR channel conversion is performed on the MDCT spectral data obtained in step (4).

[0079] (6) Inverse MDCT conversion module, low-latency synthesis window addition module, and superimposed addition module Based on steps (4) and (5), the inverse MDCT transform module performs an inverse MDCT transform on the acquired MDCT spectral data to obtain a time-domain aliased signal. Next, the low-latency synthesis window module adds a low-latency synthesis window to the time-domain aliased signal. The superposition addition module superimposes the time-domain aliased buffer signals of the current frame and the preceding frame to obtain a PCM signal, i.e., obtains the final PCM data based on the superposition addition method.

[0080] (7) PCM output module The PCM output module outputs PCM data for the corresponding channel based on the configured bit depth and channel decoding mode.

[0081] The audio encoding / decoding frameworks shown in Figures 3A and 3B are merely examples of terminals used in the embodiments of this application and are not intended to limit the embodiments of this application. Those skilled in the art can derive other encoding / decoding frameworks based on Figures 3A and 3B.

[0082] Figure 4 is a diagram showing the configuration of an electronic device according to one embodiment of this application. Optionally, the electronic device is one of the devices shown in Figure 1 and includes one or more processors 401, a communication bus 402, memory 403, and one or more communication interfaces 404.

[0083] The processor 401 is one or more integrated circuits configured to implement the solution of this application, such as a central processing unit (CPU), a network processor (NP), a microprocessor, or, for example, an application-specific integrated circuit (ASIC), a programmable logic device (PLD), or a combination thereof. Optionally, the PLD is a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.

[0084] Communications bus 402 It is configured to transmit information between the aforementioned components. Optionally, the communication bus 402 may be classified as an address bus, data bus, or control bus, etc. For ease of representation, the bus is shown using only one thick line in the diagram; however, this does not mean that there is only one bus or only one type of bus.

[0085] Optionally, memory 403 may be read-only memory (ROM), random access memory (RAM), electrically erasable programmable read-only memory (EEPROM), optical disc (including compact disc read-only memory, CD-ROM, compact disc, laser disc, digital multipurpose disc, Blu-ray disc, or similar), magnetic disk storage medium or other magnetic storage device, or any other medium that can be used to transport or store expected program code in the form of instructions or data structures and is accessible to the computer, but is not limited to these. Memory 403 may exist independently and be connected to the processor 401 via the communication bus 402, or memory 403 may be integrated with the processor 401.

[0086] The communication interface 404 is configured to communicate with another device or communication network using any device such as a transceiver. The communication interface 404 may include a wired communication interface, or optionally include a wireless communication interface. The wired communication interface is, for example, an Ethernet® interface. Optionally, the Ethernet interface may be an optical interface, an electrical interface, or a combination thereof. The wireless communication interface may be a wireless local area network (WLAN) interface, a cellular network communication interface, a combination thereof, or similar.

[0087] Optionally, in some embodiments, the electronic device includes multiple processors, such as processors 401 and 405 shown in Figure 4. Each of these processors is either a single-core or multi-core processor. Optionally, a processor here is one or more devices, circuits, and / or processing cores for processing data (e.g., computer program instructions).

[0088] In a specific implementation, in one embodiment, the electronic device further includes an output device 406 and an input device 407. The output device 406 communicates with the processor 401 and can display information in multiple ways. For example, the output device 406 is a liquid crystal display (LCD), a light-emitting diode (LED) display device, a cathode ray tube (CRT) display device, or a projector. The input device 407 communicates with the processor 401 and can receive user input in multiple ways. For example, the input device 407 is a mouse, a keyboard, a touchscreen device, or a sensing device.

[0089] In some embodiments, the memory 403 is configured to store program code 410 for executing the solution of this application, and the processor 401 can execute the program code 410 stored in the memory 403. The program code includes one or more software modules, and the electronic device can implement the audio signal processing method provided in the embodiment of Figure 5 below by using the processor 401 and the program code 410 in the memory 403.

[0090] Figure 5 is a flowchart of an audio signal processing method according to one embodiment of this application. This method is applied to the encoder side. Please refer to Figure 5. This method includes the following steps.

[0091] Step 501: Subband splitting is performed separately on the audio signal based on multiple subband splitting schemes and cutoff subbands corresponding to the multiple subband splitting schemes to obtain multiple candidate subband sets, each of which corresponds one-to-one with the multiple subband splitting schemes, and each candidate subband set contains multiple subbands.

[0092] In an embodiment of this application, the encoder side obtains a set of candidate subbands by separately performing subband division on the audio signal based on a plurality of subband division schemes and cutoff subbands corresponding to the plurality of subband division schemes in order to select the optimal subband division scheme from a plurality of subband division schemes.

[0093] Let's use any one of several subband division schemes as an example. The total number of subbands represented by the subband division scheme is 32, and the number of cutoff subbands corresponding to those subband division schemes is 16, indicating that the cutoff frequency of the audio signal is in the 16th subband. For example, the full bandwidth of the audio signal is 16 kilohertz (kHz). The cutoff subband indicates that the cutoff frequency of the audio signal is 5 kHz. After subband division is performed on the audio signal based on the subband division scheme, the resulting candidate subband set contains a total of 16 subbands, and the frequency range covered by these 16 subbands is from 0 kHz to 5 kHz, i.e., the range [0, cutoff frequency] is covered.

[0094] The subband splitting process is performed one audio frame at a time. The audio signals described in this specification can be considered as audio frames. Obviously, the encoder can perform subband splitting for each audio frame according to this solution.

[0095] In the embodiments of this application, there are multiple implementations in which the encoder acquires the cutoff subband. One of these implementations is described here.

[0096] If the encoding bitrate of the audio signal is lower than a first bitrate threshold, the encoder performs bandwidth detection on the spectrum of the audio signal to obtain the cutoff frequency of the audio signal. Based on the cutoff frequency, the encoder determines the cutoff subbands corresponding to each of the multiple subband division schemes. It should be understood that when the encoding bitrate is low, the number of encoding bits that can be allocated is small. Therefore, the encoder determines the cutoff frequency through bandwidth detection and then determines the cutoff subbands. In this case, spectral values ​​exceeding the cutoff frequency are not subsequently encoded, and as a result, the encoding effect is ensured while the requirements for the encoding bitrate are met.

[0097] There are several methods by which the encoder performs bandwidth detection. In one implementation, since the value of the frequency located after the cutoff frequency in the spectrum of the audio signal is zero, the encoder traverses the frequency values ​​in the spectrum sequentially from high frequency to low frequency, and the first traversed frequency value that is greater than the energy threshold is the cutoff frequency of the audio signal.

[0098] Optionally, the encoder takes the logarithm (e.g., log10) of the frequency values ​​in the spectrum, traverses the resulting frequency values ​​from high to low frequencies, and determines the first traversed frequency value greater than the energy threshold as the cutoff frequency of the audio signal. Optionally, the energy threshold can be -50dB, -80dB, or another value.

[0099] In addition, when the audio signal is a mono-channel signal, the encoder performs bandwidth detection on the mono-channel spectrum of the audio signal to obtain the audio signal's cutoff frequency. When the audio signal is a dual-channel signal, the encoder performs bandwidth detection separately on the left-channel spectrum and the right-channel spectrum of the audio signal to obtain the left-channel cutoff frequency and the right-channel cutoff frequency. If the left-channel cutoff frequency does not match the right-channel cutoff frequency, the encoder determines the larger of the two values ​​as the audio signal's cutoff frequency. If the left-channel cutoff frequency matches the right-channel cutoff frequency, the encoder determines the left-channel cutoff frequency as the audio signal's cutoff frequency.

[0100] In some other embodiments, the encoder may instead perform bandwidth detection on the spectrum in a different manner. This is not limited to this solution.

[0101] Optionally, after obtaining the cutoff frequency of the audio signal, the encoder determines the corresponding cutoff subbands for each of the multiple subband division schemes based on the position of the cutoff frequency within the full bandwidth of the audio signal.

[0102] For example, the cutoff frequency is located at the 30th frequency in the full bandwidth of the audio signal, and the 30th frequency is located within the kth subband of a set of subbands represented by a certain subband division scheme. In this case, the cutoff subband corresponding to the subband division scheme is k.

[0103] Optionally, in embodiments of this application, the first bitrate threshold is 150kbps or another value. For illustrative purposes, an example in which the first bitrate threshold is 150kbps is used below. Optionally, in this embodiment of this application, the encoded bitrate of the audio signal is the encoded bitrate of a single channel, i.e., the encoded bitrate of a single channel is compared to the first bitrate threshold. In some other embodiments, the first bitrate threshold may be another value instead. For example, the first bitrate threshold is 150kbps. When the audio signal is a dual-channel signal, the encoded bitrate of the audio signal is either the encoded bitrate of the left channel or the encoded bitrate of the right channel. The encoded bitrate of the left channel is usually the same as the encoded bitrate of the right channel. In this case, the encoder only needs to compare the encoded bitrate of the left channel to 150kbps.

[0104] Clearly, in some other embodiments, when the audio signal is a dual-channel signal, the encoding bitrate of the audio signal is the dual-channel encoding bitrate. Correspondingly, the first bitrate threshold is 300kbps.

[0105] Optionally, if the encoding bitrate of the audio signal is not lower than a first bitrate threshold, the encoder determines the last subband represented by each of the multiple subband division schemes as the cutoff subband corresponding to each subband division scheme. It should be understood that when the encoding bitrate is high, a large number of encoding bits can be allocated. Therefore, even without bandwidth detection by the encoder, the requirements for the encoding bitrate can still be met, and to some extent, encoding efficiency can be further improved. Certainly, in some other embodiments, the encoder may instead perform bandwidth detection on the spectrum of the audio signal where the encoding bitrate is not lower than a first bitrate threshold.

[0106] In embodiments of this application, before separately performing subband division on an audio signal based on a plurality of subband division schemes and cutoff subbands corresponding to the plurality of subband division schemes, the encoder performs feature analysis on the spectrum of the audio signal to obtain feature analysis results, and determines a plurality of subband division schemes from a plurality of candidate subband division schemes based on the feature analysis results and the encoded bitrate of the audio signal. In other words, the encoder pre-selects a plurality of subband division schemes from a plurality of candidate subband division schemes through frequency domain feature analysis, and then selects the optimal subband division scheme from the plurality of subband division schemes.

[0107] Optionally, the feature analysis results include a subjective signal flag or an objective signal flag, where the subjective signal flag indicates that the energy concentration of the audio signal is not greater than the concentration threshold, and the objective signal flag indicates that the energy concentration of the audio signal is greater than the concentration threshold. In other words, the feature analysis includes subjective and objective signal analysis, and the encoder pre-selects one of several subband division schemes based on the subjective and objective signal analysis results and the encoding bitrate.

[0108] The implementation of subjective and objective signal analysis is described below.

[0109] In embodiments of this application, the encoder performs subjective and objective signal analysis based on the portion of the audio signal spectrum that does not exceed the cutoff frequency, in order to reduce computational load and improve efficiency while ensuring accuracy.

[0110] The encoder takes the base-10 logarithm of each frequency value in the spectrum that does not exceed the cutoff frequency to obtain the logarithmic result for each frequency. The encoder normalizes the logarithmic result for each frequency to a dBFS scale to obtain the logarithmic result for each frequency on the dBFS scale. The encoder determines a first frequency number and a second frequency number. The first frequency number is the total number of frequencies whose logarithmic result is not greater than the energy threshold on the dBFS scale, and the second frequency number is the total number of frequencies in the spectrum that do not exceed the cutoff frequency. The encoder determines the ratio of the first frequency number to the second frequency number as the energy concentration of the audio signal. If the energy concentration of the audio signal is greater than the concentration threshold, the encoder determines that the audio signal is an objective signal and outputs an objective signal flag. If the total energy of the audio signal is not greater than the concentration threshold, the encoder determines that the audio signal is a subjective signal and outputs a subjective signal flag.

[0111] For example, the encoder side obtains the logarithmic result for each frequency by taking the base-10 logarithm of each value of the frequency in the spectrum that does not exceed the cutoff frequency, according to equation (1).

[0112]

number

[0113] In equation (1), X(k) represents the value of the kth frequency, i.e., the kth spectral value; cutOffFreq represents the frequency corresponding to the cutoff frequency, i.e., the second frequency; abs() represents the absolute value; and Xlg(k) represents the logarithmic result of the kth frequency.

[0114] The encoder normalizes the logarithmic result of each frequency to a dBFS scale according to equation (2) to obtain the logarithmic result of each frequency on the dBFS scale.

[0115]

number

[0116] In equation (2), XdBFS(k) represents the logarithmic result of the kth frequency on the dBFS scale, and X max This indicates the maximum spectral value within the spectrum that does not exceed the cutoff frequency.

[0117] The encoder collects statistics on the total number of frequencies whose logarithmic result is not greater than -80 dB on the dBFS scale to obtain a first frequency number lowEnergyCnt, where -80 dB represents the energy threshold, which is obtained through statistical collection or by other means. The encoder determines the energy concentration energyRate of the audio signal according to equation (3).

[0118]

number

[0119] The encoder outputs subjective and objective signal flags objFlag according to equation (4). When objFlag is 1, it indicates the objective signal flag. When objFlag is 0, it indicates the subjective signal flag.

[0120]

number

[0121] In equation (4), threshold represents the centrifugation threshold.

[0122] In this embodiment of this application, the concentration threshold is 0.6, and the concentration threshold is obtained through statistical collection or by other means. For example, the concentration threshold is a constant parameter obtained based on the signal distribution of different grades of bandwidth. Obviously, in some other embodiments, the concentration threshold may be other values ​​instead.

[0123] It should be understood that the aforementioned examples are used as implementations of subjective and objective signal analysis and are not intended to limit the embodiments of this application.

[0124] In another implementation, after obtaining the first and second frequencies, the encoder determines the ratio of the second frequency to the first frequency as the energy concentration of the audio signal. If the energy concentration of the audio signal is less than the concentration threshold, the encoder determines that the audio signal is an objective signal and outputs an objective signal flag. If the total energy of the audio signal is not less than the concentration threshold, the encoder determines that the audio signal is a subjective signal and outputs a subjective signal flag. The concentration threshold in this implementation is the reciprocal of the concentration threshold in the previous implementation. In other words, from the perspective that the proportion of non-background noise energy (i.e., the first frequency) is less than the threshold, it indicates that the frequency domain features of the objective signal are strong. The essence of this implementation is the same as that of the previous implementation.

[0125] In yet another implementation, the encoder does not normalize the logarithmic result of each frequency to a dBFS scale, but directly determines a third frequency number. The third frequency number is the total number of frequencies whose logarithmic result is not greater than the energy threshold. The encoder then uses the ratio of the third frequency number to the second frequency number to determine the energy threshold of the audio signal. Inside and This is how it is determined. Note that the energy threshold in this implementation is different from the energy threshold on the dBFS scale in the first implementation.

[0126] In yet another implementation, the encoder does not take the base-10 logarithm of the frequency values ​​that do not exceed the cutoff frequency in the spectrum, but instead directly collects statistics on the total number of frequencies within the spectral range that do not exceed the energy threshold and do not exceed the cutoff frequency in the spectrum, thereby obtaining a fourth frequency number. The encoder then calculates the ratio of the fourth frequency number to the second frequency number as the energy collection of the audio signal. Inside andThe decision is made based on this. Note that the energy threshold and condensation threshold in this implementation differ from those in some of the aforementioned implementations.

[0127] It should be understood that taking a base-10 logarithm and normalizing to a dBFS scale are done because operations are performed on different scales. Scaling is an optional operation on the encoder side, and the energy threshold and lumps threshold will differ on different scales.

[0128] The following describes the implementation process by which the encoder pre-selects multiple subband partitioning schemes based on the feature analysis results and encoding bitrate.

[0129] In this embodiment of the application, the feature analysis result includes a subjective signal flag or an objective signal flag. The frame length of the audio signal is 10 milliseconds (ms) and the sampling rate is 88.2 kilohertz (kHz) or 96 kHz, or the frame length of the audio signal is 5 ms and the sampling rate is 88.2 kHz or 96 kHz, or the frame length of the audio signal is 10 ms and the sampling rate is 44.1 kHz or 48 kHz. In this case, if the encoded bitrate of the audio signal is lower than the first bitrate threshold and the feature analysis result includes a subjective signal flag, the encoder determines a first group of subband division schemes from among a plurality of candidate subband division schemes as one of the plurality of subband division schemes. The first group of subband division schemes are as follows: { {0,1,2,3,4,6,8,10,13,16,20,24,28,33,38,45,52,61,70,79,88,100,112,127,142,160,178,196,217,238,259,280,480}, {0,1,2,3,5,7,9,12,15,18,22,26,30,35,41,48,56,65,74,84,94,106,118,134,150,166,184,202,220,240,260,280,480}, {0,1,2,3,4,5,7,9,11,14,17,21,25,29,34,40,46,52,60,68,76,86,98,110,126,144,162,180,200,224,250,280,480}, {0,2,4,6,8,12,16,21,26,31,36,41,46,51,56,61,66,71,77,83,89,95,103,111,121,131,147,163,179,203,240,280,480}, {0,1,2,3,5,7,9,12,15,19,23,27,32,37,43,49,57,66,76,86,98,110,125,140,158,176,194,216,238,264,290,320,480}, {0,1,2,3,5,7,10,13,17,21,25,30,35,41,47,54,62,70,80,90,102,114,130,146,162,180,198,218,240,264,290,320,480}, {0,1,2,4,6,8,11,14,18,22,26,30,36,42,50,58,66,76,88,100,112,128,144,160,182,204,226,256,286,316,352,400,480}, {0,1,2,4,6,8,11,14,18,22,26,30,36,42,50,58,68,78,90,102,116,132,148,166,186,208,234,262,292,324,360,400,480} }.

[0130] Figure 6 shows the relationship between the subbands represented by the eight subband division schemes included in the first group of subband division schemes and the initial frequencies of those subbands.

[0131] The audio signal frame length is 10 ms and the sampling rate is 88.2 kHz or 96 kHz, or the audio signal frame length is 5 ms and the sampling rate is 88.2 kHz or 96 kHz, or the audio signal frame length is 10 ms and the sampling rate is 44.1 kHz or 48 kHz. In this case, if the encoded bitrate of the audio signal is not lower than the first bitrate threshold and / or the feature analysis result includes an objective signal flag, the encoder determines the second group of subband division schemes from among the multiple candidate subband division schemes as one of the multiple subband division schemes. The subband division schemes of the second group are as follows: { {0,1,2,3,4,5,6,7,8,10,12,14,16,19,22,26,30,35,40,45,50,57,64,73,82,92,102,112,124,136,148,160,480}, {0,1,2,3,4,5,7,9,11,13,15,18,21,24,28,33,38,44,50,57,64,73,82,93,104,116,128,140,155,170,185,200,480}, {0,1,2,3,4,6,8,10,13,16,20,24,28,33,38,45,52,61,70,79,88,100,112,127,142,160,178,196,217,238,259,280,480}, {0,1,2,4,6,10,14,18,22,26,30,34,42,50,58,66,74,84,96,108,120,136,152,168,192,216,240,272,304,336,376,424,480}, {0,1,2,4,6,10,14,18,26,34,42,50,62,74,86,98,112,128,144,160,176,196,216,236,256,280,304,328,352,384,416,448,480}, {0,80,92,104,112,120,128,136,144,148,152,156,160,164,168,172,176,180,184,188,192,196,200,208,216,224,232,240,248,256,268,280,480}, {0,200,212,224,232,240,248,256,264,268,272,276,280,284,288,292,296,300,304,308,312,316,320,328,336,344,352,360,368,376,388,400,480}, {0,320,332,344,356,364,372,380,384,388,392,396,400,404,408,412,416,420,424,428,432,436,440,444,448,452,456,460,464,468,472,476,480} }.

[0132] Figure 7 shows the relationship between the subbands represented by the eight subband division schemes included in the second group of subband division schemes, and the initial frequencies of those subbands.

[0133] When the frame length of the audio signal is 10 ms and the sampling rate is 88.2 kHz or 96 kHz, the spectrum of each audio frame contained in the audio signal contains 960 frequencies. In the process of performing subband division based on the second group of subband division schemes, the encoder multiplies each subband division value in the second group of subband division schemes by 2 to obtain subband division values ​​corresponding to 960 frequencies, and performs subband division based on these subband division values. When the frame length of the audio signal is 5 ms and the sampling rate is 88.2 kHz or 96 kHz, or when the frame length of the audio signal is 10 ms and the sampling rate is 44.1 kHz or 48 kHz, the spectrum of each audio frame contained in the audio signal contains 480 frequencies, so the last subband division value in each subband division scheme included in the second group of subband division schemes is also 480. Therefore, the encoder directly performs subband division based on the second group of subband division schemes.

[0134] The audio signal has a frame length of 5ms and a sampling rate of 44.1kHz or 48kHz. In this case, if the encoded bitrate of the audio signal is lower than the first bitrate threshold and the feature analysis result includes a subjective signal flag, the encoder determines the third group of subband division schemes from among the multiple candidate subband division schemes as one of the multiple subband division schemes. The subband division schemes in the third group are as follows: { {0,1,2,3,4,5,6,7,8,9,10,12,14,16,19,22,26,30,35,39,44,50,56,63,71,80,89,98,108,119,129,140,240}, {0,1,2,3,4,5,6,7,8,9,11,13,15,17,20,24,28,32,37,42,47,53,59,67,75,83,92,101,110,120,130,140,240}, {0,1,2,3,4,5,6,7,8,9,10,11,12,14,17,20,23,26,30,34,38,43,49,55,63,72,81,90,100,112,125,140,240}, {0,1,2,3,4,6,8,10,13,15,18,20,23,25,28,30,33,35,38,41,44,47,51,55,60,65,73,81,89,101,120,140,240}, {0,1,2,3,4,5,6,7,9,11,13,14,16,18,21,24,28,33,38,43,49,55,62,70,79,88,97,108,119,132,145,160,240}, {0,1,2,3,4,5,6,7,8,10,12,14,17,20,23,27,31,35,40,45,51,57,65,73,81,90,99,109,120,132,145,160,240}, {0,1,2,3,4,5,6,7,9,11,13,15,18,21,25,29,33,38,44,50,56,64,72,80,91,102,113,128,143,158,176,200,240}, {0,1,2,3,4,5,6,7,9,11,13,15,18,21,25,29,34,39,45,51,58,66,74,83,93,104,117,131,146,162,180,200,240} }.

[0135] Figure 8 shows the relationship between the subbands represented by the eight subband division schemes included in the third group of subband division schemes, and the initial frequencies of those subbands.

[0136] The audio signal has a frame length of 5ms and a sampling rate of 44.1kHz or 48kHz. In this case, if the encoded bitrate of the audio signal is not lower than the first bitrate threshold and / or the feature analysis result includes an objective signal flag, the encoder determines the fourth group of subband division schemes from among the multiple candidate subband division schemes as one of the multiple subband division schemes. The subband division schemes in the fourth group are as follows: { {0,1,2,3,4,5,6,7,8,9,10,11,12,13,14,15,16,17,18,19,20,21,22,23,24,26,28,30,32,34,37,40,120}, {0,1,2,3,4,5,6,7,8,9,10,11,12,13,14,15,16,18,20,22,24,26,28,30,32,34,36,38,41,44,47,50,120}, {0,1,2,3,4,5,6,7,8,9,10,11,12,14,16,18,20,22,24,26,28,31,34,37,40,44,48,52,56,60,65,70,120}, {0,1,2,3,4,5,6,7,8,9,10,11,12,13,15,17,19,21,24,27,30,34,38,42,48,54,60,68,76,84,94,106,120}, {0,1,2,3,4,5,6,7,8,10,12,14,16,19,22,25,28,32,36,40,44,49,54,59,64,70,76,82,88,96,104,112,120}, {0,20,23,26,28,30,32,34,36,37,38,39,40,41,42,43,44,45,46,47,48,49,50,52,54,56,58,60,62,64,67,70,120}, {0,50,53,56,58,60,62,64,66,67,68,69,70,71,72,73,74,75,76,77,78,79,80,82,84,86,88,90,92,94,97,100,120}, {0,80,83,86,89,91,93,95,96,97,98,99,100,101,102,103,104,105,106,107,108,109,110,111,112,113,114,115,116,117,118,119,120} }.

[0137] Figure 9 shows the relationship between the subbands represented by the eight subband division schemes included in the fourth group of subband division schemes, and the initial frequencies of those subbands.

[0138] When the frame length of the audio signal is 5ms and the sampling rate is 44.1kHz or 48kHz, the spectrum of each audio frame contained in the audio signal contains 240 frequencies. In the process of performing subband division based on the fourth group subband division scheme, the encoder multiplies each subband division value in the fourth group subband division scheme by 2 to obtain subband division values ​​corresponding to the 240 frequencies, and performs subband division based on the subband division values ​​corresponding to the 240 frequencies.

[0139] Each subband division scheme provided in the embodiments of this application conforms to the Bark requirement. The Bark scale indicates that the spectral subband division policy divides the subbands in terms of hearing, based on the auditory characteristics of the human ear.

[0140] Step 502: Determine the total scale value for each candidate subband set based on the spectral values ​​of the audio signals in the subbands included in the candidate subband set, the encoded bitrate of the audio signals, and the subband bandwidth of the subbands included in the candidate subband set.

[0141] In an embodiment of this application, after obtaining a plurality of candidate subband sets corresponding one-to-one to a plurality of subband division schemes, the encoder determines the total scale value of each candidate subband set based on the spectral value of the audio signal in the subband included in the candidate subband set, the encoded bitrate of the audio signal, and the subband bandwidth of the subband included in the candidate subband set.

[0142] Optionally, the encoder determines the total scale value of each candidate subband set based on the spectral values ​​of the audio signals within the subbands included in the candidate subband set, the encoded bitrate of the audio signals, and the subband bandwidth of the subbands included in the candidate subband set. This implementation process includes determining the scale factor of each subband included in the first candidate subband set based on the spectral values ​​of the audio signals within the subbands included in the first candidate subband set. The first candidate subband set is one of the multiple candidate subband sets. The encoder then determines the total scale value of the first candidate subband set based on the encoded bitrate of the audio signals and the scale factor and subband bandwidth of each subband included in the first candidate subband set. For each of the candidate subband sets other than the first candidate subband set within the multiple candidate subband sets, the encoder determines the total scale value of each of those other candidate subband sets in the same manner as determining the total scale value of the first candidate subband set.

[0143] There are multiple implementations in which the encoder determines the scale factor of each subband. In one implementation, the encoder determines the scale factor of each subband included in the first candidate subband set based on the spectral values ​​of the audio signals within the subbands included in the first candidate subband set. This includes the encoder obtaining the maximum value among the absolute values ​​of all spectral values ​​of the audio signals within the first subband and determining the scale factor of the first subband based on this maximum value. The first subband is any of the subbands in the first candidate subband set. For each subband in the first candidate subband set other than the first subband, the encoder determines the scale factor of each of these other subbands in the same manner as determining the scale factor of the first subband.

[0144] For example, the encoder determines the scale factor of each subband in the first candidate subband set according to equation (5).

[0145]

number

[0146] In equation (5), X(k) represents the k-th spectral value of the audio signal, b represents the sequential number of the subband, I(b) represents the initial frequency of subband b, B represents the cutoff subband corresponding to the first candidate subband set, i.e., the total number of subbands included in the first candidate subband set, abs() represents obtaining the absolute value, max() represents obtaining the maximum value, ceil() represents rounding up, and E() represents the scale factor of the subband.

[0147] The following describes an implementation in which the encoder determines the total scale value of the first candidate subband set based on the encoding bitrate of the audio signal and the scale factor and subband bandwidth of each subband included in the first candidate subband set.

[0148] Optionally, the encoding bitrate of the audio signal is not lower than the first bitrate threshold, and / or the energy concentration of the audio signal is greater than the concentration threshold. The encoder determines an energy smoothing reference value based on the encoding bitrate of the audio signal and the second bitrate threshold. The encoder determines the total energy value of each subband included in the first candidate subband set based on the energy smoothing reference value and the scale factor and subband bandwidth of each subband included in the first candidate subband set. The encoder sums the total energy values ​​of the subbands included in the first candidate subband set to obtain the total scale value of the first candidate subband set. If the energy concentration of the audio signal is greater than the concentration threshold, it indicates that the audio signal is an objective signal. It should be understood that when the encoding bitrate is high and / or the audio signal is an objective signal, the encoder determines the total scale value based on the total energy value of each subband.

[0149] In the embodiments of this application, there are multiple implementations in which the encoder side determines the energy smoothing reference value. One of these implementations will be described here. In this implementation, the encoder side determines the energy smoothing reference value according to equation (6).

[0150]

number

[0151] In equation (6), E floor`<value>` indicates the energy smoothing reference value, `bpsPerChn` indicates the encoded bitrate of the audio signal, where the encoded bitrate of the audio signal is the encoded bitrate of a single channel, `200` indicates that the second bitrate threshold is 200kbps, `min()` indicates obtaining the minimum value, and `int()` indicates truncation. Note that the second bitrate threshold may be any other value.

[0152] There are several implementations in which the encoder determines the total energy value of each subband in the first candidate subband set based on an energy smoothing reference value and the scale factor and subband bandwidth of each subband included in the first candidate subband set. One of these implementations is described here. In this implementation, for the first subband included in the first candidate subband set, the encoder determines the larger of the scale factor and energy smoothing reference value of the first subband as the reference scale value of the first subband. The encoder determines the product of the reference scale value of the first subband and the subband bandwidth of the first subband as the total energy value of the first subband. The first subband is one of the subbands in the first candidate subband set. For each subband in the first candidate subband set other than the first subband, the encoder determines the total energy value of each of those other subbands in the same manner as determining the total energy value of the first subband.

[0153] The encoder determines the total energy value of each subband included in the first candidate subband set and the total scale value of the first candidate subband set according to equation (7).

[0154]

number

[0155] In equation (7), b represents the sequential number of the subband, B represents the cutoff subband corresponding to the first candidate subband set, bandWidth() represents the subband bandwidth, E(b) represents the scale factor of subband b, and E floor indicates the energy smoothing reference value, max() indicates obtaining the maximum value, and max[E(b),E floor ]*bandWidth(b) indicates the total energy value of subband b, E total This indicates the total scale value of the first subband set.

[0156] The above describes the implementation process by which the encoder determines the total scale value of the first candidate subband set when the encoded bitrate of the audio signal is not lower than the first bitrate threshold and / or the energy concentration of the audio signal is greater than the concentration threshold. Below, we describe the implementation process by which the encoder determines the total scale value of the first candidate subband set when the encoded bitrate of the audio signal is lower than the first bitrate threshold and the energy concentration of the audio signal is not greater than the concentration threshold.

[0157] If the encoded bitrate of the audio signal is lower than a first bitrate threshold and the energy concentration of the audio signal is not greater than a concentration threshold, the encoder determines the total scale value of the first candidate subband set based on the encoded bitrate of the audio signal and the scale factor and subband bandwidth of each subband included in the first candidate subband set. The implementation process includes the encoder determining an energy smoothing reference value based on the encoded bitrate of the audio signal and a second bitrate threshold. The encoder determines the scale difference values ​​of the subbands included in the first candidate subband set based on the energy smoothing reference value and the scale factor of each subband included in the first candidate subband set. The scale difference values ​​represent the difference between the scale factor of a corresponding subband and the scale factor of an adjacent subband of that corresponding subband. The encoder determines the total scale value of the first candidate subband set based on the scale difference values ​​of the subbands included in the first candidate subband set and the subband bandwidth of each subband. If the energy concentration of the audio signal is not greater than a concentration threshold, it indicates that the audio signal is a subjective signal. It should be understood that when the encoding bitrate is low and the audio signal is a subjective signal, the encoder determines the total scale value based on the difference between each subband and adjacent subbands.

[0158] For implementations where the encoder determines the energy smoothing reference value based on the encoding bitrate of the audio signal and a second bitrate threshold, please refer to the related explanation above. Further details will not be explained again here.

[0159] There are several implementations in which the encoder determines the scale difference values ​​of the subbands included in the first candidate subband set based on an energy smoothing reference value and the scale factor of each subband included in the first candidate subband set. One of these implementations is described here. In this implementation, for the first subband included in the first candidate subband set, the encoder determines the first smoothing value, the second smoothing value, and the third smoothing value of the first subband based on the energy smoothing reference value, the scale factor of the first subband, and the scale factor of the adjacent subbands of the first subband. The encoder then determines the scale difference value of the first subband based on the first smoothing value, the second smoothing value, and the third smoothing value of the first subband. The first subband is any of the subbands in the first candidate subband set.

[0160] Optionally, if the first subband is the first subband in the first candidate subband set, the encoder determines the larger of the scale factor and energy smoothing reference value of the first subband as the first smoothing value of the first subband. If the first subband is not the first subband in the first candidate subband set, the encoder determines the larger of the scale factor and energy smoothing reference value of the previous subband adjacent to the first subband as the first smoothing value of the first subband.

[0161] The encoder determines the second smoothing value of the first subband as the larger of the scale factor and the energy smoothing reference value of the first subband.

[0162] If the first subband is the last subband in the first candidate subband set, the encoder determines the larger of the scale factor and energy smoothing reference value of the first subband as the third smoothing value of the first subband. If the first subband is not the last subband in the first candidate subband set, the encoder determines the larger of the scale factor and energy smoothing reference value of the next subband adjacent to the first subband as the third smoothing value of the first subband.

[0163] In other words, the encoder determines the first smoothing value, the second smoothing value, and the third smoothing value for each subband according to equations (8), (9), and (10).

[0164]

number

[0165] In equations (8), (9), and (10), left(), center(), and right() represent the first smoothing value, the second smoothing value, and the third smoothing value, respectively. In this embodiment of the application, the first smoothing value, the second smoothing value, and the third smoothing value may also be referred to as the left smoothing value, the intermediate smoothing value, and the right smoothing value, respectively.

[0166] Optionally, an implementation process in which the encoder determines the scale difference value of a first subband after determining the first, second, and third smoothing values ​​of the first subband includes the encoder determining the first difference value and the second difference value of the first subband for the first subband included in the first candidate subband set. The first difference value is the absolute value of the difference between the first and second smoothing values ​​of the first subband, and the second difference value is the absolute value of the difference between the second and third smoothing values ​​of the first subband. Based on the first and second difference values ​​of the first subband, the encoder determines the scale difference value of the first subband. The first subband is any of the subbands in the first candidate subband set.

[0167] For example, the encoder determines the scale difference value of the first subband according to equation (11).

[0168]

number

[0169] An implementation process in which the encoder determines the scale difference value of each subband included in the first candidate subband set, and then determines the total scale value of the first candidate subband set based on the scale difference value and subband bandwidth of each subband, includes the encoder determining the smoothing weighting coefficient for each subband included in the first candidate subband set based on the number of subbands included in the first candidate subband set and the subband bandwidth of each subband. The encoder sums the smoothing weighting coefficients of the subbands included in the first candidate subband set to obtain the total smoothing weighting coefficient for the first candidate subband set. The encoder multiplies the scale difference value of the subbands included in the first candidate subband set by the smoothing weighting coefficient to obtain the weighted scale difference value for the subbands included in the first candidate subband set. The encoder sums the weighted scale difference values ​​of the subbands included in the first candidate subband set to obtain the total scale value for the first candidate subband set. The encoder divides the total scale value by the total smoothing weighting coefficient of the first candidate subband set to obtain the total scale value of the first candidate subband set.

[0170] The step by which the encoder determines the smoothing weighting coefficients for each subband and the total smoothing weighting coefficients for the first candidate subband set may instead be performed before the scale difference values ​​for each subband are determined. The order of the steps performed by the encoder is not limited to the embodiments of this application.

[0171] Optionally, the encoder determines a scaling factor for subband division based on the number of subbands included in the first candidate subband set. The encoder then determines a smoothing weighting factor for each subband included in the first candidate subband set based on the scaling factor for subband division and the subband bandwidth of each subband included in the first candidate subband set.

[0172] For example, the encoder determines the subband division scaling coefficient coef according to equation (12).

[0173]

number

[0174] The encoder determines the smoothing weighting coefficient frac for each subband according to equation (13).

[0175]

number

[0176] The encoder determines the sum of the smoothing weight coefficients (sum) for the first candidate subband set according to equation (14).

[0177]

number

[0178] The encoder side calculates the total scale value E' of the first candidate subband set according to equation (15). total To decide.

[0179]

number

[0180] E diff (b)*frac(b) shows the weighted scale difference value of subband b.

[0181] The encoder side calculates the total scale value E of the first candidate subband set according to equation (16). total To decide.

[0182]

number

[0183] When the audio signal is a mono-channel signal, the encoder can calculate the total scale value of each candidate subband set according to the formula described above. When the audio signal is a dual-channel signal, the spectrum of the audio signal includes the left channel spectrum and the right channel spectrum, and the encoder calculates the total scale value of each candidate subband set based on the left channel spectrum and the right channel spectrum. For example, the encoder obtains the total scale value of the candidate subband set by adding the total scale value calculated based on the left channel spectrum and the total scale value calculated based on the right channel spectrum. In one implementation, another sum Σ is added to the formula described above related to the sum Σ, and the added sum Σ is the sum of the associated data for the left channel and the associated data for the right channel.

[0184] Step 503: Based on the sum of the scale values ​​of each candidate subband set, select one candidate subband set from the multiple candidate subband sets as the target subband set, where each subband in the target subband set has a scale factor used to shape the spectral envelope of the audio signal.

[0185] In embodiments of this application, the encoder determines the candidate subband set having the smallest total scale value among a plurality of candidate subband sets as the target subband set. In some other embodiments, the encoder may instead determine the candidate subband set having the second smallest total scale value among a plurality of candidate subband sets as the target subband set. The second smallest total scale value is the smallest total scale value among the scale values ​​other than the smallest total scale value.

[0186] From the above, it can be seen that the encoder selects the optimal subband division method from multiple subband division methods based on the characteristics of the audio signal. In other words, the subband division method has signal adaptive characteristics and helps improve coding effect and compression efficiency.

[0187] To further improve coding effectiveness and compression efficiency, when the audio signal is a dual-channel signal, the encoder can further determine, based on the determined target subband set, whether coding performance can be improved by performing Mid / Side stereo transform coding (MS transformation) on the spectrum of the audio signal. Furthermore, if it is determined that MS transformation helps improve coding performance, the encoder performs subsequent coding procedures based on the spectrum obtained through MS transformation. If it is determined that MS transformation does not help improve coding performance, the encoder performs subsequent coding procedures based on the original spectrum of the audio signal. This will be discussed later.

[0188] In embodiments of this application, when the audio signal is a dual-channel signal, the encoder determines a first total scale value based on the scale factor and subband bandwidth of each subband included in the target subband set. The encoder performs an MS transformation on the spectrum of the dual-channel signal to obtain the transformed spectrum of the dual-channel signal. The encoder determines the transformed scale factor of each subband in the target subband set based on the transformed spectral values ​​of the dual-channel signal in each subband included in the target subband set. The encoder determines a second total scale value based on the transformed scale factor and subband bandwidth of each subband included in the target subband set. If the first total scale value is not greater than the second total scale value, the encoder determines the dual-channel signal (the dual-channel signal before MS transformation) to be the signal to be encoded.

[0189] It should be understood that the first total scale value is the total scale value before MS conversion, and the second total scale value is the total scale value obtained through MS conversion. A higher total scale value indicates a lower coding performance gain. If the first total scale value is not greater than the second total scale value, it indicates that MS conversion does not help improve coding performance. Therefore, the encoder determines the dual-channel signal before MS conversion to be the signal to be encoded.

[0190] Optionally, the spectrum of the dual-channel signal before MS conversion is referred to as the LR spectrum, and the spectrum of the dual-channel signal obtained through MS conversion is referred to as the MS spectrum. LR indicates the left and right channels.

[0191] When the audio signal is a dual-channel signal, the scale factor includes a left channel scale factor and a right channel scale factor. Optionally, an implementation process in which the encoder determines a first total scale value based on the scale factor and subband bandwidth of each subband included in the target subband set includes determining the product of the left channel scale factor and the subband bandwidth of each subband included in the target subband set as the left channel energy value of the corresponding subband, and determining the product of the right channel scale factor and the subband bandwidth of each subband included in the target subband set as the right channel energy value of the corresponding subband. The encoder then sums the left channel energy values ​​and right channel energy values ​​of all subbands included in the target subband set to obtain a first total scale value.

[0192] For example, the encoder determines the first total scale value according to equation (17).

[0193]

number

[0194] In equation (17), totalScale1 represents the first total scale value, and ch represents the sequential numbers of the left and right channels. When ch=0, E(b) represents the left channel scale factor. When ch=1, E(b) represents the right channel scale factor.

[0195] The encoder performs MS conversion according to equation (18).

[0196]

number

[0197] In equation (18), L and R represent the spectral values ​​of the left channel and the right channel before conversion, respectively. M and S represent the converted spectral values ​​of the left channel and the right channel, respectively. The encoder processes the spectral values ​​of the left channel and the right channel at the corresponding frequencies in order to obtain the spectral values ​​of the converted spectra of the left channel and the right channel, respectively. The converted spectral values ​​of the left channel and the right channel are the spectral values ​​of the two channels contained in the converted dual-channel signal. The converted left channel and the converted right channel can also be referred to as the converted M channel and the converted S channel.

[0198] The encoder determines the transformed scale factor for each subband according to equation (19), similar to equation (5).

[0199]

number

[0200] In equation (19), X_MS(k) represents the k-th transformed spectral value, and E_MS(b) represents the scale factor of subband b on the M channel or S channel, i.e., the scale factor of subband b on the transformed channel. The encoder calculates the scale factor of the M channel based on the spectral value of the M channel according to equation (19), and calculates the scale factor of the S channel based on the spectral value of the S channel according to equation (19).

[0201] The encoder determines the second total scale value according to equation (20).

[0202]

number

[0203] In Equation (20), totalScale2 represents the second total scale value, and ch represents the sequence numbers of the M channel and the S channel. When ch = 0, E_MS(b) represents the scale factor of the subband on the converted left channel. When ch = 1, E_MS(b) represents the scale factor of the subband on the converted right channel. That is, it is the scale factor of subband b on the M channel or the S channel.

[0204] Optionally, when the first total scale value is greater than the second total scale value, and the encoded bit rate of the audio signal is not lower than the first bit rate threshold and / or the energy concentration of the audio signal is greater than the concentration threshold, the encoder side determines the converted dual-channel signal as the signal to be encoded. It should be understood that the fact that the first total scale value is greater than the second total scale value indicates that the MS conversion can help improve the coding performance. Therefore, the encoder side determines the dual-channel signal obtained through the MS conversion as the signal to be encoded.

[0205] As can be seen from the above, when the audio signal is a dual-channel signal, the scale factor includes the left channel scale factor and the right channel scale factor. Optionally, when the first total scale value is greater than the second total scale value, the encoded bit rate of the audio signal is lower than the first bit rate threshold, and the energy concentration of the audio signal is not greater than the concentration threshold, the encoder side is based on the left channel scale factor and the right channel scale factor of each subband included in the target subband set, for each subband included in the target subband set, the left channel scale factor and the right channelDetermine the difference value from the scale factor. On the encoder side, based on the initial frequency and cut-off frequency of each sub-band included in the target sub-band set, determine the start - end frequency difference value of each sub-band included in the target sub-band set. Within the target sub-band set, for the left channel scale factor and the right channel scale factor, if there is at least one sub-band having a difference value therebetween and having a start - end frequency difference value within the first range, the encoder side determines the dual-channel signal before conversion as the signal to be encoded.

[0206] In other words, when the encoding bit rate is low and the audio signal is an objective signal, the encoder side determines whether the MS conversion can improve the coding performance based on the difference value between the left-channel scale factor and the right-channel scale factor and the start - end frequency difference value of the sub-band.

[0207] Optionally, the encoder side traverses all the sub-bands within the target sub-band set. For the left channel scale factor and the right channel scale factor, if there is a sub-band having a difference value therebetween and having a start - end frequency difference value within the first range, the encoder side determines the dual-channel signal before conversion as the signal to be encoded.

[0208] For example, on the encoder side, according to Equation (21), determine the difference value between the left channel scale factor and the right channel scale factor of each sub-band.

[0209]

Equation

[0210] In Equation (21), E_L() represents the left-channel scale factor, and E_R() represents right This shows the channel scale factor, and diffSFflag() is left channel Scale factor and right channel This shows the difference between the scale factor and the actual value.

[0211] The encoder side, according to equation (21), the left of each subband channel Scale factor and right channel When determining the difference value between the scale factor and the target value, the difference threshold is 3.

[0212] The encoder determines the subband center frequency of each subband according to equation (22).

[0213]

number

[0214] In equation (22), freq() represents the start-end frequency difference, bandstart() and bandend() represent the initial frequency and cutoff frequency, respectively, SamplingRate represents the sampling rate in Hz, and FrameLength represents the number of sampling points in each frame.

[0215] Optionally, when the encoder determines the subband center frequency of each subband according to equation (22), the first range is (3500, 12000).

[0216] Simply put, when the encoder uses equations (21) and (22), the encoder traverses all subbands within the target subband set. channel Scale factor and right channel When a subband exists that has a difference value diffSFflag between it and the scale factor, and has a subband center frequency freq within the range of (3500, 12000), the encoder determines the dual-channel signal before conversion to be the signal to be encoded.

[0217] If there is no such at least one sub - band within the target sub - band set, the encoder side determines the converted dual - channel signal as the signal to be encoded. The at least one sub - band refers to the left channel scale factor greater than the differential threshold and the right channel scale factor, and has a sub - band center frequency within the first range and a difference value between the two scale factors.

[0218] Hereinafter, referring to FIG. 10, the implementation process for the encoder side to determine whether to use the converted dual - channel signal as the signal to be encoded will be described again.

[0219] Please refer to FIG. 10. The encoder side calculates a first total scale value based on the selected target sub - band set and the left / right (LR) channel scale factors (scale factor, SF) of each sub - band within the target sub - band set. The first total scale value is the sum of the products of the LR channel scale factors of all sub - bands within the target sub - band set and the corresponding sub - band bandwidths. The encoder side converts the LR channel spectrum to the MS channel spectrum and calculates a second total scale value. The second total scale value is the sum of the products of the MS channel scale factors of all sub - bands within the target sub - band set and the corresponding sub - band bandwidths. If the first total scale value is not greater than the second total scale value, the encoder side determines the dual - channel signal before conversion as the signal to be encoded, sets MSFlag = 0, indicating that the execution of subsequent operations is not based on the spectral values obtained through MS conversion.

[0220] If the first total scale value is greater than the second total scale value, the encoder side determines whether the audio signal (that is, the dual - channel signal before conversion) satisfies the first condition. The first condition is that the encoding bit rate of the audio signal is lower than the first bit - rate threshold and the energy concentration of the audio signal Inside The bitrate must be less than the condensation threshold. If the audio signal satisfies the first condition, the encoder sets the high bitrate flag to 0. If the audio signal does not satisfy the first condition, the encoder sets the high bitrate flag to 1.

[0221] If the high bitrate flag is equal to 1, the encoder determines that the converted dual-channel signal is the signal to be encoded and sets MSFlag=1 to indicate that subsequent operations will be based on spectral values ​​obtained through MS conversion. If the high bitrate flag is equal to 0, the encoder calculates the LR channel SF difference value and the subband center frequency of each subband through traversal. If the subbands obtained through traversal satisfy the second condition, the encoder sets the SF difference flag to 1. The second condition is that the LR channel SF difference value of the corresponding subband is smaller than the difference threshold and the subband center frequency is within the first range. If the subbands obtained through traversal do not satisfy the second condition, the encoder sets the SF difference flag to 0.

[0222] If the SF difference flag is equal to 1, the encoder determines that the original dual-channel signal is the signal to be encoded and sets MSFlag=0. If the SF difference flag is equal to 0, the encoder determines that the converted dual-channel signal is the signal to be encoded and sets MSFlag=1.

[0223] In addition to the aforementioned implementation in which the encoder determines whether the converted dual-channel signal should be used as the signal to be encoded, the encoder may also make the determination using other methods. In other words, the aforementioned implementation is not intended to limit the embodiments of this application.

[0224] In short, in the embodiments of this application, the optimal subband division scheme is selected from a plurality of subband division schemes based on the characteristics of the audio signal. In other words, the subband division scheme has signal adaptive characteristics and can improve interference immunity by adapting to the encoding bitrate of the audio signal. Specifically, the audio signal is divided separately based on a plurality of subband division schemes, and a total scale value corresponding to each subband division scheme is determined based on the spectral values ​​of the audio signal in the subbands obtained through the division, the bandwidth of each subband, and the encoding bitrate of the audio signal. An optimal subband set is obtained by selecting the optimal target subband division scheme based on this total scale value. Subsequently, spectral envelope shaping is performed based on the scale factor of each subband in the optimal subband set, thereby improving coding effectiveness and compression efficiency.

[0225] Figure 11 shows the configuration of an audio signal processing device 1100 according to one embodiment of this application. The processing device 1100 can be implemented as part of or as part of an electronic device by using software, hardware, or a combination thereof. The electronic device can be any of the devices shown in Figure 1. Please refer to Figure 11. The device includes a subband division module 1101, a first determination module 1102, and a selection module 1103.

[0226] The subband splitting module 1101 is configured to obtain multiple candidate subband sets by separately performing subband splitting on an audio signal based on multiple subband splitting schemes and cutoff subbands corresponding to the multiple subband splitting schemes. The multiple candidate subband sets correspond one-to-one with the multiple subband splitting schemes, and each candidate subband set contains multiple subbands.

[0227] The first decision module 1102 is configured to determine the total scale value of each candidate subband set based on the spectral values ​​of the audio signals in the subbands included in the candidate subband set, the encoded bitrate of the audio signals, and the subband bandwidth of the subbands included in the candidate subband set.

[0228] The selection module 1103 is configured to select one candidate subband set as the target subband set from among several candidate subband sets based on the sum of the scale values ​​of each candidate subband set. Each subband included in the target subband set has a scale factor that is used to shape the spectral envelope of the audio signal.

[0229] Optionally, selection module 1103 is, The system is configured to determine the target subbandset from among multiple candidate subbandsets that has the smallest total scale value.

[0230] Optionally, the first decision module 1102 is: A first decision submodule configured to determine the scale factor of each subband included in a first candidate subband set based on the spectral values ​​of the audio signals in the subbands included in the first candidate subband set, wherein the first candidate subband set is one of the multiple candidate subband sets. The system includes a second determination submodule configured to determine the total scale value of a first candidate subband set based on the encoding bitrate of the audio signal and the scale factor and subband bandwidth of each subband included in the first candidate subband set.

[0231] Optionally, the second decision submodule is: For a first subband included in the first candidate subband set, obtain the maximum value among the absolute values ​​of all spectral values ​​of the audio signal within the first subband, and determine that the first subband is any of the subbands in the first candidate subband set. It is configured to determine the scale factor of the first subband based on the maximum value.

[0232] Optionally, the encoding bitrate of the audio signal is not lower than a first bitrate threshold, and / or the energy concentration of the audio signal is greater than a concentration threshold.

[0233] The second decision submodule is: An energy smoothing reference value is determined based on the encoding bitrate of the audio signal and a second bitrate threshold. Based on the energy smoothing reference value and the scale factor and subband bandwidth of each subband included in the first candidate subband set, the total energy value of each subband included in the first candidate subband set is determined. The system is configured to obtain the total scale value of the first candidate subband set by summing the total energy values ​​of the subbands included in the first candidate subband set.

[0234] Optionally, the second decision submodule is: For the first subband included in the first candidate subband set, the larger of the scale factor and the energy smoothing reference value of the first subband is determined as the reference scale value of the first subband, and the first subband is one of the subbands in the first candidate subband set. The system is configured to determine the total energy value of the first subband by multiplying the reference scale value of the first subband by the subband bandwidth of the first subband.

[0235] Optionally, the encoding bitrate of the audio signal is lower than the first bitrate threshold, and the energy concentration of the audio signal is not greater than the concentration threshold.

[0236] The second decision submodule is: An energy smoothing reference value is determined based on the encoding bitrate of the audio signal and a second bitrate threshold. Based on the energy smoothing reference value and the scale factor of each subband included in the first candidate subband set, the scale difference value of the subbands included in the first candidate subband set is determined, and this scale difference value represents the difference between the scale factor of the corresponding subband and the scale factor of the adjacent subband of the corresponding subband. The system is configured to determine the total scale value of the first candidate subband set based on the scale difference values ​​of the subbands included in the first candidate subband set and the subband bandwidth of each subband.

[0237] Optionally, the second decision submodule is: For a first subband included in the first candidate subband set, the first smoothing value, second smoothing value, and third smoothing value of the first subband are determined based on the energy smoothing reference value, the scale factor of the first subband, and the scale factor of the adjacent subbands of the first subband, and the first subband is one of the subbands in the first candidate subband set. The system is configured to determine the scale difference value of the first subband based on the first smoothing value, the second smoothing value, and the third smoothing value of the first subband.

[0238] Optionally, the second decision submodule is: If the first subband is the first subband in the first candidate subband set, the larger of the scale factor and the energy smoothing reference value of the first subband is determined as the first smoothing value of the first subband. If the first subband is not the first subband in the first candidate subband set, the larger of the scale factor and the energy smoothing reference value of the previous subband adjacent to the first subband is determined as the first smoothing value of the first subband. The larger of the scale factor and energy smoothing reference value of the first subband is determined as the second smoothing value of the first subband. If the first subband is the last subband in the first candidate subband set, the larger of the scale factor and energy smoothing reference value of the first subband is determined as the third smoothing value of the first subband. If the first subband is not the last subband in the first candidate subband set, the larger of the scale factor and energy smoothing reference value of the next subband adjacent to the first subband is determined as the third smoothing value of the first subband.

[0239] Optionally, the second decision submodule is: For a first subband included in the first candidate subband set, a first difference value and a second difference value of the first subband are determined, wherein the first difference value is the absolute value of the difference between the first smoothed value and the second smoothed value of the first subband, and the second difference value is the absolute value of the difference between the second smoothed value and the third smoothed value of the first subband, and the first subband is any subband in the first candidate subband set. The system is configured to determine the scale difference value of the first subband based on the first difference value and the second difference value of the first subband.

[0240] Optionally, the second decision submodule is: Based on the number of subbands included in the first candidate subband set and the subband bandwidth of each subband, the smoothing weighting coefficient for each subband included in the first candidate subband set is determined. The smoothing weighting coefficients of the subbands included in the first candidate subband set are summed to obtain the total smoothing weighting coefficient of the first candidate subband set. The scale difference values ​​of the subbands included in the first candidate subband set are multiplied by a smoothing weighting coefficient to obtain the weighted scale difference values ​​of the subbands included in the first candidate subband set. The weighted scale difference values ​​of the subbands included in the first candidate subband set are summed to obtain the total scale value of the first candidate subband set. The method is configured to obtain the total scale value of the first candidate subband set by dividing the total scale value by the total smoothing weight coefficient of the first candidate subband set.

[0241] Optionally, the device 1100 further: A bandwidth detection module configured to perform bandwidth detection on the spectrum of an audio signal to obtain the cutoff frequency of the audio signal when the encoding bitrate of the audio signal is lower than a first bitrate threshold, The system includes a second decision module configured to determine the cutoff subbands corresponding to each of a plurality of subband division schemes based on the cutoff frequencies.

[0242] Optionally, the device 1100 further: The system includes a third decision module configured to determine the last subband represented by each of several subband division schemes as the cutoff subband corresponding to each subband division scheme, provided that the encoded bitrate of the audio signal is not lower than a first bitrate threshold.

[0243] Optionally, the device 1100 further: A feature analysis module configured to perform feature analysis on the spectrum of an audio signal and obtain feature analysis results, The system includes a fourth decision module configured to determine multiple subband partitioning schemes from a plurality of candidate subband partitioning schemes based on feature analysis results and the encoding bitrate of the audio signal.

[0244] Optionally, the feature analysis results include a subjective signal flag or an objective signal flag, where the subjective signal flag indicates that the energy concentration of the audio signal is not greater than the concentration threshold, and the objective signal flag indicates that the energy concentration of the audio signal is greater than the concentration threshold.

[0245] Optionally, the audio signal frame length may be 10 milliseconds and the sampling rate 88.2 kHz or 96 kHz, or the audio signal frame length may be 5 milliseconds and the sampling rate 88.2 kHz or 96 kHz, or the audio signal frame length may be 10 milliseconds and the sampling rate 44.1 kHz or 48 kHz.

[0246] The fourth decision module is, The system includes a third decision submodule configured to determine a first group of subband partitioning schemes among a plurality of candidate subband partitioning schemes as a plurality of subband partitioning schemes when the encoded bitrate of the audio signal is lower than a first bitrate threshold and the feature analysis result includes a subjective signal flag.

[0247] The subband division scheme for the first group is as follows: { {0,1,2,3,4,6,8,10,13,16,20,24,28,33,38,45,52,61,70,79,88,100,112,127,142,160,178,196,217,238,259,280,480}, {0,1,2,3,5,7,9,12,15,18,22,26,30,35,41,48,56,65,74,84,94,106,118,134,150,166,184,202,220,240,260,280,480}, {0,1,2,3,4,5,7,9,11,14,17,21,25,29,34,40,46,52,60,68,76,86,98,110,126,144,162,180,200,224,250,280,480}, {0,2,4,6,8,12,16,21,26,31,36,41,46,51,56,61,66,71,77,83,89,95,103,111,121,131,147,163,179,203,240,280,480}, {0,1,2,3,5,7,9,12,15,19,23,27,32,37,43,49,57,66,76,86,98,110,125,140,158,176,194,216,238,264,290,320,480}, {0,1,2,3,5,7,10,13,17,21,25,30,35,41,47,54,62,70,80,90,102,114,130,146,162,180,198,218,240,264,290,320,480}, {0,1,2,4,6,8,11,14,18,22,26,30,36,42,50,58,66,76,88,100,112,128,144,160,182,204,226,256,286,316,352,400,480}, {0,1,2,4,6,8,11,14,18,22,26,30,36,42,50,58,68,78,90,102,116,132,148,166,186,208,234,262,292,324,360,400,480} }.

[0248] Optionally, the audio signal frame length may be 10 milliseconds and the sampling rate 88.2 kHz or 96 kHz, or the audio signal frame length may be 5 milliseconds and the sampling rate 88.2 kHz or 96 kHz, or the audio signal frame length may be 10 milliseconds and the sampling rate 44.1 kHz or 48 kHz.

[0249] The fourth decision module is, Includes a fourth decision submodule configured to determine a second group of subband partitioning schemes among a plurality of candidate subband partitioning schemes as a plurality of subband partitioning schemes if the encoded bitrate of the audio signal is not lower than a first bitrate threshold and / or the feature analysis result includes an objective signal flag.

[0250] The subband division scheme for the second group is as follows: { {0,1,2,3,4,5,6,7,8,10,12,14,16,19,22,26,30,35,40,45,50,57,64,73,82,92,102,112,124,136,148,160,480}, {0,1,2,3,4,5,7,9,11,13,15,18,21,24,28,33,38,44,50,57,64,73,82,93,104,116,128,140,155,170,185,200,480}, {0,1,2,3,4,6,8,10,13,16,20,24,28,33,38,45,52,61,70,79,88,100,112,127,142,160,178,196,217,238,259,280,480}, {0,1,2,4,6,10,14,18,22,26,30,34,42,50,58,66,74,84,96,108,120,136,152,168,192,216,240,272,304,336,376,424,480}, {0,1,2,4,6,10,14,18,26,34,42,50,62,74,86,98,112,128,144,160,176,196,216,236,256,280,304,328,352,384,416,448,480}, {0,80,92,104,112,120,128,136,144,148,152,156,160,164,168,172,176,180,184,188,192,196,200,208,216,224,232,240,248,256,268,280,480}, {0,200,212,224,232,240,248,256,264,268,272,276,280,284,288,292,296,300,304,308,312,316,320,328,336,344,352,360,368,376,388,400,480}, {0,320,332,344,356,364,372,380,384,388,392,396,400,404,408,412,416,420,424,428,432,436,440,444,448,452,456,460,464,468,472,476,480} }.

[0251] Optionally, the audio signal frame length is 5 milliseconds, and the sampling rate is 44.1 kHz or 48 kHz.

[0252] The fourth decision module is, The system includes a fifth decision submodule configured to determine a third group of subband partitioning schemes among a plurality of candidate subband partitioning schemes as a plurality of subband partitioning schemes when the encoded bitrate of the audio signal is lower than a first bitrate threshold and the feature analysis result includes a subjective signal flag.

[0253] The subband division scheme for the third group is as follows: { {0,1,2,3,4,5,6,7,8,9,10,12,14,16,19,22,26,30,35,39,44,50,56,63,71,80,89,98,108,119,129,140,240}, {0,1,2,3,4,5,6,7,8,9,11,13,15,17,20,24,28,32,37,42,47,53,59,67,75,83,92,101,110,120,130,140,240}, {0,1,2,3,4,5,6,7,8,9,10,11,12,14,17,20,23,26,30,34,38,43,49,55,63,72,81,90,100,112,125,140,240}, {0,1,2,3,4,6,8,10,13,15,18,20,23,25,28,30,33,35,38,41,44,47,51,55,60,65,73,81,89,101,120,140,240}, {0,1,2,3,4,5,6,7,9,11,13,14,16,18,21,24,28,33,38,43,49,55,62,70,79,88,97,108,119,132,145,160,240}, {0,1,2,3,4,5,6,7,8,10,12,14,17,20,23,27,31,35,40,45,51,57,65,73,81,90,99,109,120,132,145,160,240}, {0,1,2,3,4,5,6,7,9,11,13,15,18,21,25,29,33,38,44,50,56,64,72,80,91,102,113,128,143,158,176,200,240}, {0,1,2,3,4,5,6,7,9,11,13,15,18,21,25,29,34,39,45,51,58,66,74,83,93,104,117,131,146,162,180,200,240} }.

[0254] Optionally, the audio signal frame length is 5 milliseconds, and the sampling rate is 44.1 kHz or 48 kHz.

[0255] The fourth decision module is, Includes a sixth decision submodule configured to determine a fourth group of subband partitioning schemes among a plurality of candidate subband partitioning schemes as a plurality of subband partitioning schemes if the encoded bitrate of the audio signal is not lower than a first bitrate threshold and / or the feature analysis result includes an objective signal flag.

[0256] The subband division scheme for the fourth group is as follows: { {0,1,2,3,4,5,6,7,8,9,10,11,12,13,14,15,16,17,18,19,20,21,22,23,24,26,28,30,32,34,37,40,120}, {0,1,2,3,4,5,6,7,8,9,10,11,12,13,14,15,16,18,20,22,24,26,28,30,32,34,36,38,41,44,47,50,120}, {0,1,2,3,4,5,6,7,8,9,10,11,12,14,16,18,20,22,24,26,28,31,34,37,40,44,48,52,56,60,65,70,120}, {0,1,2,3,4,5,6,7,8,9,10,11,12,13,15,17,19,21,24,27,30,34,38,42,48,54,60,68,76,84,94,106,120}, {0,1,2,3,4,5,6,7,8,10,12,14,16,19,22,25,28,32,36,40,44,49,54,59,64,70,76,82,88,96,104,112,120}, {0,20,23,26,28,30,32,34,36,37,38,39,40,41,42,43,44,45,46,47,48,49,50,52,54,56,58,60,62,64,67,70,120}, {0,50,53,56,58,60,62,64,66,67,68,69,70,71,72,73,74,75,76,77,78,79,80,82,84,86,88,90,92,94,97,100,120}, {0,80,83,86,89,91,93,95,96,97,98,99,100,101,102,103,104,105,106,107,108,109,110,111,112,113,114,115,116,117,118,119,120} }.

[0257] Optionally, the audio signal is a dual-channel signal.

[0258] The device 1100 further, A fifth decision module configured to determine a first total scale value based on the scale factor and subband bandwidth of each subband included in the target subband set, A conversion module configured to perform mid / side stereo conversion coding on the spectrum of a dual-channel signal to obtain the converted spectrum of the dual-channel signal, A sixth decision module configured to determine the transformed scale factor of each subband in the target subband set based on the transformed spectral values ​​of the dual-channel signals within each subband included in the target subband set, A seventh decision module configured to determine a second total scale value based on the transformed scale factor and subband bandwidth of each subband included in the target subband set, The system includes an eighth decision module configured to determine a dual-channel signal as the signal to be encoded if the first total scale value is not greater than the second total scale value.

[0259] Optionally, the device 1100 further: The system is configured to determine that the converted dual-channel signal is the signal to be encoded if the first total scale value is greater than the second total scale value, and the encoded bitrate of the audio signal is not lower than the first bitrate threshold, and / or the energy concentration of the audio signal is greater than the concentration threshold.

[0260] Optionally, the scale factor includes the left channel scale factor and the right channel scale factor.

[0261] The device 1100 further, If the first total scale value is greater than the second total scale value, the encoded bitrate of the audio signal is lower than the first bitrate threshold, and the energy concentration of the audio signal is not greater than the concentration threshold, then, based on the left channel scale factor and right channel scale factor of each subband included in the target subband set, the left channel of each subband included in the target subband set channel Scale factor and right channel A ninth decision module configured to determine the difference value between the scale factor and, A tenth decision module configured to determine the subband center frequency of each subband included in the target subband set based on the initial frequency and cutoff frequency of each subband included in the target subband set, Within the target subband set, left-hand values ​​greater than the difference threshold. channel Scale factor and right channel The system includes an eleventh decision module configured to determine a dual-channel signal as the signal to be encoded if there is at least one subband having a difference value between it and a scale factor and having a subband center frequency within a first range.

[0262] Optionally, the device 1100 further: The system is configured to determine the converted dual-channel signal as the signal to be encoded if at least one subband is not present in the target subband set.

[0263] In embodiments of this application, the optimal subband division scheme is selected from a plurality of subband division schemes based on the characteristics of the audio signal. In other words, the subband division scheme has signal adaptive characteristics and can improve interference immunity by adapting to the encoding bitrate of the audio signal. Specifically, the audio signal is divided separately based on a plurality of subband division schemes, and a total scale value corresponding to each subband division scheme is determined based on the spectral values ​​of the audio signal in the subbands obtained through the division, the bandwidth of each subband, and the encoding bitrate of the audio signal. An optimal subband set is obtained by selecting the optimal target subband division scheme based on this total scale value. Subsequently, spectral envelope shaping is performed based on the scale factor of each subband in the optimal subband set, thereby improving coding effectiveness and compression efficiency.

[0264] Furthermore, when the audio signal processing device provided in the above-described embodiment processes an audio signal, the division of the functional modules described above is used merely as an example for illustrative purposes. In actual application, the functions described above may be assigned to different functional modules for implementation based on requirements. In other words, the internal configuration of the device is divided into different functional modules to implement all or some of the functions described above. Also, the audio signal processing device and the audio signal processing method embodiments provided in the above-described embodiment belong to the same concept. For specific implementation processes, please refer to the method embodiments. Details will not be described again here.

[0265] All or part of the embodiments described above may be implemented using software, hardware, firmware, or any combination thereof. When software is used for implementation, all or part of the embodiments may be implemented in the form of a computer program product. A computer program product includes one or more computer instructions. When the computer instructions are loaded onto a computer and executed, all or part of the procedures or functions according to the embodiments of this application are generated. The computer may be a general-purpose computer, a dedicated computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, computer instructions may be transmitted by wired means (e.g., coaxial cable, optical fiber, or digital subscriber line (DSL)) or wireless means (e.g., infrared, radio waves, or microwaves) from one website, computer, server, or data center to another. A computer-readable storage medium can be any usable medium accessible by a computer, or it can be a data storage device that integrates one or more usable media, such as a server or data center. The usable media may be magnetic media (e.g., floppy disks, hard disks, or magnetic tapes), optical media (e.g., digital versatile discs (DVDs)), semiconductor media (e.g., solid-state disks (SSDs)), etc. The computer-readable storage medium referred to in the embodiments of this application may be a non-volatile storage medium, i.e., a non-temporary storage medium.

[0266] It should be understood that in this specification, “at least one” refers to one or more, and “multiple” refers to two or more. In the description of embodiments of this application, unless otherwise specified, “ / ” means “or.” For example, A / B may represent A or B. In this specification, “and / or” describes only the relationship between the related objects and indicates that three relationships may exist. For example, A and / or B may represent three cases: only A exists, both A and B exist, and only B exists. Also, in order to clearly describe the technical solutions in embodiments of this application, terms such as “first” and “second” are used in embodiments of this application to distinguish between the same or similar items that have essentially the same function and purpose. As a person skilled in the art will understand, terms such as “first” and “second” do not limit the number or order of execution, and terms such as “first” and “second” do not indicate a clear difference.

[0267] Furthermore, the information (including, but not limited to, user device information and user personal information), data (including, but not limited to, data used for analysis, stored data, and displayed data) and signals in the embodiments of this application shall be used under the authorization of the user or the full authorization of all parties, and the capture, use, and processing of the relevant data shall comply with the relevant laws, regulations, and standards of the relevant countries and regions. For example, the audio signals relevant in the embodiments of this application shall be obtained under the full authorization.

[0268] The above description is an embodiment provided in this application and is not intended to limit this application. The original Any modification, equivalent substitution, or improvement made without deviation from the principle falls within the scope of protection of this application.

Claims

1. An audio signal processing method, Subband division is performed separately on the audio signal based on multiple subband division schemes and cutoff subbands corresponding to the multiple subband division schemes to obtain multiple candidate subband sets, each of which corresponds one-to-one with the multiple subband division schemes, and each candidate subband set includes multiple subbands. The total scale value of each candidate subband set is determined based on the spectral value of the audio signal within the subband included in the candidate subband set, the encoded bitrate of the audio signal, and the subband bandwidth of the subband included in the candidate subband set. Based on the sum scale value of each candidate subband set, one candidate subband set is selected from the plurality of candidate subband sets as the target subband set, and each subband included in the target subband set has a scale factor used to shape the spectral envelope of the audio signal. A method having the following characteristics.

2. Selecting one candidate subband set as the target subband set from the plurality of candidate subband sets based on the sum scale value of each candidate subband set is: Among the plurality of candidate subband sets, the candidate subband set having the smallest total scale value is determined as the target subband set. The method according to claim 1, wherein the method is as follows:

3. Determining the total scale value of each candidate subband set based on the spectral value of the audio signal in the subband included in the candidate subband set, the encoded bitrate of the audio signal, and the subband bandwidth of the subband included in the candidate subband set means that For a first candidate subband set among the plurality of candidate subband sets, the scale factor of each subband included in the first candidate subband set is determined based on the spectral value of the audio signal in the subband included in the first candidate subband set, and the first candidate subband set is one of the plurality of candidate subband sets. Based on the encoded bitrate of the audio signal and the scale factor and subband bandwidth of each subband included in the first candidate subband set, the total scale value of the first candidate subband set is determined. The method according to claim 2, wherein the method is as follows:

4. Determining the scale factor of each subband included in the first candidate subband set based on the spectral values ​​of the audio signals within the subbands included in the first candidate subband set is: For a first subband included in the first candidate subband set, the maximum value among the absolute values ​​of all spectral values ​​of the audio signal within the first subband is obtained, and the first subband is one of the subbands in the first candidate subband set. The scale factor of the first subband is determined based on the aforementioned maximum value. The method according to claim 3, wherein the above is achieved.

5. The encoded bitrate of the audio signal is not lower than the first bitrate threshold, and / or the energy concentration of the audio signal is greater than the concentration threshold. Determining the total scale value of the first candidate subband set based on the encoded bitrate of the audio signal and the scale factor and subband bandwidth of each subband included in the first candidate subband set is: An energy smoothing reference value is determined based on the encoded bitrate of the audio signal and a second bitrate threshold. Based on the energy smoothing reference value and the scale factor and subband bandwidth of each subband included in the first candidate subband set, the total energy value of each subband included in the first candidate subband set is determined. The total energy values ​​of the subbands included in the first candidate subband set are summed to obtain the total scale value of the first candidate subband set. Having, The method according to claim 4.

6. Determining the total energy value of each subband included in the first candidate subband set based on the energy smoothing reference value and the scale factor and subband bandwidth of each subband included in the first candidate subband set is: For the first subband included in the first candidate subband set, the larger of the scale factor and the energy smoothing reference value of the first subband is determined as the reference scale value of the first subband, and the first subband is one of the subbands in the first candidate subband set. The product of the reference scale value of the first subband and the subband bandwidth of the first subband is determined as the total energy value of the first subband. The method according to claim 5, wherein the method is as follows:

7. The encoded bitrate of the audio signal is lower than the first bitrate threshold, and the energy concentration of the audio signal is not greater than the concentration threshold. Determining the total scale value of the first candidate subband set based on the encoded bitrate of the audio signal and the scale factor and subband bandwidth of each subband included in the first candidate subband set is: An energy smoothing reference value is determined based on the encoded bitrate of the audio signal and a second bitrate threshold. Based on the energy smoothing reference value and the scale factor of each subband included in the first candidate subband set, the scale difference value of the subbands included in the first candidate subband set is determined, and the scale difference value represents the difference between the scale factor of the corresponding subband and the scale factor of the adjacent subband of the corresponding subband. The total scale value of the first candidate subband set is determined based on the scale difference value of the subbands included in the first candidate subband set and the subband bandwidth of each subband. Having, The method according to claim 4.

8. Determining the scale difference value of the subbands included in the first candidate subband set based on the energy smoothing reference value and the scale factor of each subband included in the first candidate subband set is: For the first subband included in the first candidate subband set, the first smoothing value, second smoothing value, and third smoothing value of the first subband are determined based on the energy smoothing reference value, the scale factor of the first subband, and the scale factor of the adjacent subband of the first subband, and the first subband is one of the subbands in the first candidate subband set. Based on the first smoothed value, the second smoothed value, and the third smoothed value of the first subband, the scale difference value of the first subband is determined. The method according to claim 7, wherein the method is as follows:

9. Determining the first smoothing value, second smoothing value, and third smoothing value of the first subband based on the energy smoothing reference value, the scale factor of the first subband, and the scale factor of the adjacent subband of the first subband is as follows: If the first subband is the first subband in the first candidate subband set, the larger of the scale factor of the first subband and the energy smoothing reference value is determined as the first smoothing value of the first subband; if the first subband is not the first subband in the first candidate subband set, the larger of the scale factor of the preceding subband adjacent to the first subband and the energy smoothing reference value is determined as the first smoothing value of the first subband. The larger of the scale factor and the energy smoothing reference value of the first subband is determined as the second smoothing value of the first subband. If the first subband is the last subband in the first candidate subband set, the larger of the scale factor and the energy smoothing reference value of the first subband is determined as the third smoothing value of the first subband. If the first subband is not the last subband in the first candidate subband set, the larger of the scale factor and the energy smoothing reference value of the next subband adjacent to the first subband is determined as the third smoothing value of the first subband. The method according to claim 8, wherein the above is achieved.

10. Determining the scale difference value of the first subband based on the first smoothed value, the second smoothed value, and the third smoothed value of the first subband is: For the first subband included in the first candidate subband set, a first difference value and a second difference value of the first subband are determined, wherein the first difference value is the absolute value of the difference between the first smoothed value and the second smoothed value of the first subband, and the second difference value is the absolute value of the difference between the second smoothed value and the third smoothed value of the first subband, and the first subband is any subband in the first candidate subband set. Based on the first difference value and the second difference value of the first subband, the scale difference value of the first subband is determined. The method according to claim 9, wherein the method is as follows:

11. Determining the total scale value of the first candidate subband set based on the scale difference value of the subbands included in the first candidate subband set and the subband bandwidth of each subband is: Based on the number of subbands included in the first candidate subband set and the bandwidth of each subband, the smoothing weighting coefficient for each subband included in the first candidate subband set is determined. The smoothing weighting coefficients of the subbands included in the first candidate subband set are summed up to obtain the total smoothing weighting coefficient of the first candidate subband set. The scale difference value of the subbands included in the first candidate subband set is multiplied by the smoothing weighting coefficient to obtain the weighted scale difference value of the subbands included in the first candidate subband set. The weighted scale difference values ​​of the subbands included in the first candidate subband set are summed to obtain the total scale value of the first candidate subband set. The total scale value of the first candidate subband set is obtained by dividing the total scale value by the total smoothing weighting coefficient of the first candidate subband set. The method according to claim 10, wherein the above is achieved.

12. This method further, If the encoded bitrate of the audio signal is lower than the first bitrate threshold, bandwidth detection is performed on the spectrum of the audio signal to obtain the cutoff frequency of the audio signal. Based on the aforementioned cutoff frequencies, the cutoff subbands corresponding to each of the plurality of subband division schemes are determined. The method according to claim 1, wherein the method is as follows:

13. This method further, If the encoded bitrate of the audio signal is not lower than the first bitrate threshold, the last subband indicated by each of the plurality of subband division schemes is determined as the cutoff subband corresponding to each subband division scheme. The method according to claim 1, wherein the method is as follows:

14. This method further, Feature analysis is performed on the spectrum of the aforementioned audio signal to obtain the feature analysis results. Based on the feature analysis results and the encoding bitrate of the audio signal, the plurality of subband division schemes are determined from the plurality of candidate subband division schemes. The method according to claim 1, wherein the method is as follows:

15. The method according to claim 14, wherein the feature analysis result has a subjective signal flag or an objective signal flag, the subjective signal flag indicating that the energy concentration of the audio signal is not greater than the concentration threshold, and the objective signal flag indicating that the energy concentration of the audio signal is greater than the concentration threshold.

16. The frame length of the audio signal is 10 milliseconds and the sampling rate is 88.2 kilohertz or 96 kilohertz, or the frame length of the audio signal is 5 milliseconds and the sampling rate is 88.2 kilohertz or 96 kilohertz, or the frame length of the audio signal is 10 milliseconds and the sampling rate is 44.1 kilohertz or 48 kilohertz. Determining the plurality of subband division schemes from a plurality of candidate subband division schemes based on the feature analysis results and the encoding bitrate of the audio signal is: If the encoded bitrate of the audio signal is lower than the first bitrate threshold and the feature analysis result has the subjective signal flag, the first group of subband division schemes among the plurality of candidate subband division schemes is determined to be the plurality of subband division schemes. Having, The subband division scheme for the first group is as follows: { {0,1,2,3,4,6,8,10,13,16,20,24,28,33,38,45,52,61,70,79,88,100,112,127,142,160,178,196,217,238,259,280,480}, {0,1,2,3,5,7,9,12,15,18,22,26,30,35,41,48,56,65,74,84,94,106,118,134,150,166,184,202,220,240,260,280,480}, {0,1,2,3,4,5,7,9,11,14,17,21,25,29,34,40,46,52,60,68,76,86,98,110,126,144,162,180,200,224,250,280,480}, {0,2,4,6,8,12,16,21,26,31,36,41,46,51,56,61,66,71,77,83,89,95,103,111,121,131,147,163,179,203,240,280,480}, {0,1,2,3,5,7,9,12,15,19,23,27,32,37,43,49,57,66,76,86,98,110,125,140,158,176,194,216,238,264,290,320,480}, {0,1,2,3,5,7,10,13,17,21,25,30,35,41,47,54,62,70,80,90,102,114,130,146,162,180,198,218,240,264,290,320,480}, {0,1,2,4,6,8,11,14,18,22,26,30,36,42,50,58,66,76,88,100,112,128,144,160,182,204,226,256,286,316,352,400,480}, {0,1,2,4,6,8,11,14,18,22,26,30,36,42,50,58,68,78,90,102,116,132,148,166,186,208,234,262,292,324,360,400,480} }、 The method according to claim 15.

17. The frame length of the audio signal is 10 milliseconds and the sampling rate is 88.2 kilohertz or 96 kilohertz, or the frame length of the audio signal is 5 milliseconds and the sampling rate is 88.2 kilohertz or 96 kilohertz, or the frame length of the audio signal is 10 milliseconds and the sampling rate is 44.1 kilohertz or 48 kilohertz. Determining the plurality of subband division schemes from a plurality of candidate subband division schemes based on the feature analysis results and the encoding bitrate of the audio signal is: If the encoded bitrate of the audio signal is not lower than the first bitrate threshold, and / or the feature analysis result has the objective signal flag, then the second group of subband division schemes among the plurality of candidate subband division schemes is determined to be the plurality of subband division schemes. Having, The subband division scheme for the second group is as follows: { {0,1,2,3,4,5,6,7,8,10,12,14,16,19,22,26,30,35,40,45,50,57,64,73,82,92,102,112,124,136,148,160,480}, {0,1,2,3,4,5,7,9,11,13,15,18,21,24,28,33,38,44,50,57,64,73,82,93,104,116,128,140,155,170,185,200,480}, {0,1,2,3,4,6,8,10,13,16,20,24,28,33,38,45,52,61,70,79,88,100,112,127,142,160,178,196,217,238,259,280,480}, {0,1,2,4,6,10,14,18,22,26,30,34,42,50,58,66,74,84,96,108,120,136,152,168,192,216,240,272,304,336,376,424,480}, {0,1,2,4,6,10,14,18,26,34,42,50,62,74,86,98,112,128,144,160,176,196,216,236,256,280,304,328,352,384,416,448,480}, {0,80,92,104,112,120,128,136,144,148,152,156,160,164,168,172,176,180,184,188,192,196,200,208,216,224,232,240,248,256,268,280,480}, {0,200,212,224,232,240,248,256,264,268,272,276,280,284,288,292,296,300,304,308,312,316,320,328,336,344,352,360,368,376,388,400,480}, {0,320,332,344,356,364,372,380,384,388,392,396,400,404,408,412,416,420,424,428,432,436,440,444,448,452,456,460,464,468,472,476,480} }、 The method according to claim 15.

18. The frame length of the audio signal is 5 milliseconds, and the sampling rate is 44.1 kilohertz or 48 kilohertz. Determining the plurality of subband division schemes from a plurality of candidate subband division schemes based on the feature analysis results and the encoding bitrate of the audio signal is: If the encoded bitrate of the audio signal is lower than the first bitrate threshold and the feature analysis result has the subjective signal flag, the third group of subband division schemes among the plurality of candidate subband division schemes is determined to be the plurality of subband division schemes. Having, The subband division scheme for the third group is as follows: { {0,1,2,3,4,5,6,7,8,9,10,12,14,16,19,22,26,30,35,39,44,50,56,63,71,80,89,98,108,119,129,140,240}, {0,1,2,3,4,5,6,7,8,9,11,13,15,17,20,24,28,32,37,42,47,53,59,67,75,83,92,101,110,120,130,140,240}, {0,1,2,3,4,5,6,7,8,9,10,11,12,14,17,20,23,26,30,34,38,43,49,55,63,72,81,90,100,112,125,140,240}, {0,1,2,3,4,6,8,10,13,15,18,20,23,25,28,30,33,35,38,41,44,47,51,55,60,65,73,81,89,101,120,140,240}, {0,1,2,3,4,5,6,7,9,11,13,14,16,18,21,24,28,33,38,43,49,55,62,70,79,88,97,108,119,132,145,160,240}, {0,1,2,3,4,5,6,7,8,10,12,14,17,20,23,27,31,35,40,45,51,57,65,73,81,90,99,109,120,132,145,160,240}, {0,1,2,3,4,5,6,7,9,11,13,15,18,21,25,29,33,38,44,50,56,64,72,80,91,102,113,128,143,158,176,200,240}, {0,1,2,3,4,5,6,7,9,11,13,15,18,21,25,29,34,39,45,51,58,66,74,83,93,104,117,131,146,162,180,200,240} }、 The method according to claim 15.

19. The frame length of the audio signal is 5 milliseconds, and the sampling rate is 44.1 kilohertz or 48 kilohertz. Determining the plurality of subband division schemes from a plurality of candidate subband division schemes based on the feature analysis results and the encoding bitrate of the audio signal is: If the encoded bitrate of the audio signal is not lower than the first bitrate threshold, and / or the feature analysis result has the objective signal flag, the fourth group of subband division schemes among the plurality of candidate subband division schemes is determined to be the plurality of subband division schemes. Having, The subband division scheme for the fourth group is as follows: { {0,1,2,3,4,5,6,7,8,9,10,11,12,13,14,15,16,17,18,19,20,21,22,23,24,26,28,30,32,34,37,40,120}, {0,1,2,3,4,5,6,7,8,9,10,11,12,13,14,15,16,18,20,22,24,26,28,30,32,34,36,38,41,44,47,50,120}, {0,1,2,3,4,5,6,7,8,9,10,11,12,14,16,18,20,22,24,26,28,31,34,37,40,44,48,52,56,60,65,70,120}, {0,1,2,3,4,5,6,7,8,9,10,11,12,13,15,17,19,21,24,27,30,34,38,42,48,54,60,68,76,84,94,106,120}, {0,1,2,3,4,5,6,7,8,10,12,14,16,19,22,25,28,32,36,40,44,49,54,59,64,70,76,82,88,96,104,112,120}, {0,20,23,26,28,30,32,34,36,37,38,39,40,41,42,43,44,45,46,47,48,49,50,52,54,56,58,60,62,64,67,70,120}, {0,50,53,56,58,60,62,64,66,67,68,69,70,71,72,73,74,75,76,77,78,79,80,82,84,86,88,90,92,94,97,100,120}, {0,80,83,86,89,91,93,95,96,97,98,99,100,101,102,103,104,105,106,107,108,109,110,111,112,113,114,115,116,117,118,119,120} }、 The method according to claim 15.

20. The aforementioned audio signal is a dual-channel signal. This method further, Based on the scale factor and subband bandwidth of each subband included in the target subband set, a first total scale value is determined. Mid / side stereo conversion coding is performed on the spectrum of the dual-channel signal to obtain the converted spectrum of the dual-channel signal. Based on the transformed spectral values ​​of the dual-channel signals within each subband included in the target subband set, the transformed scale factor of each subband in the target subband set is determined. Based on the converted scale factor and subband bandwidth of each subband included in the target subband set, a second total scale value is determined. If the first total scale value is not greater than the second total scale value, the dual-channel signal is determined to be the signal to be encoded. The method according to claim 14, wherein the above is achieved.

21. This method further, The converted dual-channel signal is determined to be the signal to be encoded if the first total scale value is greater than the second total scale value, and the encoded bitrate of the audio signal is not lower than the first bitrate threshold and / or the energy concentration of the audio signal is greater than the concentration threshold. The method according to claim 20, wherein the above is achieved.

22. The scale factor has a left channel scale factor and a right channel scale factor, This method further, If the first total scale value is greater than the second total scale value, the encoded bitrate of the audio signal is lower than the first bitrate threshold, and the energy concentration of the audio signal is not greater than the concentration threshold, then the difference between the left channel scale factor and the right channel scale factor of each subband included in the target subband set is determined based on the left channel scale factor and the right channel scale factor of each subband included in the target subband set. Based on the initial frequency and cutoff frequency of each subband included in the target subband set, the subband center frequency of each subband included in the target subband set is determined. The dual-channel signal is determined to be the signal to be encoded if, within the target subband set, there is at least one subband having a difference value between the left channel scale factor and the right channel scale factor that is greater than the difference threshold, and having a subband center frequency within a first range. The method according to claim 20, wherein the above is achieved.

23. This method further, If at least one subband is not present in the target subband set, the converted dual-channel signal is determined to be the signal to be encoded. The method according to claim 22, wherein the above is achieved.

24. An audio signal processing method, The bitstream is decoded to obtain side information, which includes information about a target subband partitioning scheme, a quantization step, and a scale factor for each subband in the target subband set, wherein the target subband partitioning scheme indicates a subband partitioning scheme for the audio signal in the bitstream and relates to the encoding bitrate of the audio signal, and the target subband set is obtained by partitioning the audio signal according to the target subband partitioning scheme. Based on the information regarding the target subband division scheme, the quantization step, and the scale factor of each subband in the target subband set, the spectral data of each subband in the target subband set is decoded, and the spectral data of the audio signal includes the spectral data of each subband in the target subband set. A method having the following characteristics.

25. Decoding the spectral data of the audio signal based on the information regarding the target subband division scheme and the quantization step further involves After decoding the spectral data of each subband in the target subband set, if bits still remain in the bitstream, residual decoding is performed on the bitstream to obtain spectral data of another subband, and the spectral data of the audio signal further includes the spectral data of the other subband. The method according to claim 24, wherein the above is achieved.

26. The aforementioned side information further includes a low bitrate flag, Before decoding the spectral data of the audio signal based on the information about the target subband division scheme and the quantization step, the method further: When the low bitrate flag indicates a low bitrate of the bitstream, spectral noise shaping is performed based on the scale factor to obtain an adjustment factor, which is used to dequantize the spectral data to be decoded. The method according to claim 24, wherein the above is achieved.

27. The side information further includes a low bitrate flag, and after decoding the spectral data of the audio signal, the method further When the low bitrate flag indicates a low bitrate of the bitstream, hole padding is performed at the low bitrate level of the bitstream. The method according to claim 24, wherein the above is achieved.

28. The side information further includes joint coding decision information, and after decoding the spectral data of the audio signal, the method further When a dual-channel joint coding mode is determined according to the joint coding determination information, an LR channel conversion is performed on the spectral data of the audio signal, and the dual-channel joint coding mode corresponds to an coding bitrate greater than or equal to a first bitrate threshold, and the sampling rate is higher than the sampling rate threshold. The method according to claim 27, wherein the above is achieved.

29. This method further, The packet header information is decoded from the bitstream, and the packet header information includes the sampling rate and frame length of the audio signal. The encoded bitrate is obtained based on the bitstream size of the bitstream, the sampling rate of the bitstream, the sampling rate of the audio signal, and the frame length. The method according to claim 28, wherein the method is as follows:

30. The method according to claim 28, wherein the first bitrate threshold is 150 kbps / channel and the sampling rate threshold is 88.2 kHz.

31. The frame length of the audio signal is 10 milliseconds and the sampling rate is 88.2 kilohertz or 96 kilohertz, or the frame length of the audio signal is 5 milliseconds and the sampling rate is 88.2 kilohertz or 96 kilohertz, or the frame length of the audio signal is 10 milliseconds and the sampling rate is 44.1 kilohertz or 48 kilohertz. The aforementioned target subband division scheme belongs to the following first group of subband division schemes: { {0,1,2,3,4,6,8,10,13,16,20,24,28,33,38,45,52,61,70,79,88,100,112,127,142,160,178,196,217,238,259,280,480}, {0,1,2,3,5,7,9,12,15,18,22,26,30,35,41,48,56,65,74,84,94,106,118,134,150,166,184,202,220,240,260,280,480}, {0,1,2,3,4,5,7,9,11,14,17,21,25,29,34,40,46,52,60,68,76,86,98,110,126,144,162,180,200,224,250,280,480}, {0,2,4,6,8,12,16,21,26,31,36,41,46,51,56,61,66,71,77,83,89,95,103,111,121,131,147,163,179,203,240,280,480}, {0,1,2,3,5,7,9,12,15,19,23,27,32,37,43,49,57,66,76,86,98,110,125,140,158,176,194,216,238,264,290,320,480}, {0,1,2,3,5,7,10,13,17,21,25,30,35,41,47,54,62,70,80,90,102,114,130,146,162,180,198,218,240,264,290,320,480}, {0,1,2,4,6,8,11,14,18,22,26,30,36,42,50,58,66,76,88,100,112,128,144,160,182,204,226,256,286,316,352,400,480}, {0,1,2,4,6,8,11,14,18,22,26,30,36,42,50,58,68,78,90,102,116,132,148,166,186,208,234,262,292,324,360,400,480} }, or The aforementioned target subband division scheme belongs to the second group of subband division schemes, as follows: { {0,1,2,3,4,5,6,7,8,10,12,14,16,19,22,26,30,35,40,45,50,57,64,73,82,92,102,112,124,136,148,160,480}, {0,1,2,3,4,5,7,9,11,13,15,18,21,24,28,33,38,44,50,57,64,73,82,93,104,116,128,140,155,170,185,200,480}, {0,1,2,3,4,6,8,10,13,16,20,24,28,33,38,45,52,61,70,79,88,100,112,127,142,160,178,196,217,238,259,280,480}, {0,1,2,4,6,10,14,18,22,26,30,34,42,50,58,66,74,84,96,108,120,136,152,168,192,216,240,272,304,336,376,424,480}, {0,1,2,4,6,10,14,18,26,34,42,50,62,74,86,98,112,128,144,160,176,196,216,236,256,280,304,328,352,384,416,448,480}, {0,80,92,104,112,120,128,136,144,148,152,156,160,164,168,172,176,180,184,188,192,196,200,208,216,224,232,240,248,256,268,280,480}, {0,200,212,224,232,240,248,256,264,268,272,276,280,284,288,292,296,300,304,308,312,316,320,328,336,344,352,360,368,376,388,400,480}, {0,320,332,344,356,364,372,380,384,388,392,396,400,404,408,412,416,420,424,428,432,436,440,444,448,452,456,460,464,468,472,476,480} }、 The method according to claim 24.

32. The frame length of the audio signal is 5 milliseconds, the sampling rate is 44.1 kilohertz or 48 kilohertz, and the target subband division scheme belongs to the third group of subband division schemes as follows: { {0,1,2,3,4,5,6,7,8,9,10,12,14,16,19,22,26,30,35,39,44,50,56,63,71,80,89,98,108,119,129,140,240}, {0,1,2,3,4,5,6,7,8,9,11,13,15,17,20,24,28,32,37,42,47,53,59,67,75,83,92,101,110,120,130,140,240}, {0,1,2,3,4,5,6,7,8,9,10,11,12,14,17,20,23,26,30,34,38,43,49,55,63,72,81,90,100,112,125,140,240}, {0,1,2,3,4,6,8,10,13,15,18,20,23,25,28,30,33,35,38,41,44,47,51,55,60,65,73,81,89,101,120,140,240}, {0,1,2,3,4,5,6,7,9,11,13,14,16,18,21,24,28,33,38,43,49,55,62,70,79,88,97,108,119,132,145,160,240}, {0,1,2,3,4,5,6,7,8,10,12,14,17,20,23,27,31,35,40,45,51,57,65,73,81,90,99,109,120,132,145,160,240}, {0,1,2,3,4,5,6,7,9,11,13,15,18,21,25,29,33,38,44,50,56,64,72,80,91,102,113,128,143,158,176,200,240}, {0,1,2,3,4,5,6,7,9,11,13,15,18,21,25,29,34,39,45,51,58,66,74,83,93,104,117,131,146,162,180,200,240} }、 The method according to claim 24.

33. An audio signal processing device having memory and a processor, The memory is configured to store a computer program, and the computer program has program instructions. The processor is configured to call the computer program and carry out the method described in any one of claims 1 to 32. Audio signal processing device.

34. A computer-readable storage medium storing a computer program, wherein when the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 32 are performed.

35. A computer program having computer instructions, wherein when the computer instructions are executed by a processor, the steps of the method according to any one of claims 1 to 32 are performed.