Audio signal processing method and device, storage medium, and computer program product
By employing subband decomposition and spectral envelope shaping based on spectral values and coding bit rates, the method enhances coding efficiency and sound quality in audio signal processing, addressing interference in Bluetooth scenarios.
Patent Information
- Application Number
- JP2025504297
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-09-19
- Filing Date
- 2023-05-04
- Publication Date
- 2025-09-02
- Estimated Expiration
- 2043-05-04
AI Technical Summary
Existing audio encoding technologies face challenges in improving compression efficiency while maintaining sound quality, particularly in Bluetooth interconnection scenarios where interference is common.
The method involves performing subband decomposition using multiple schemes, determining optimal subband sets based on spectral values, coding bit rates, and spectral envelope shaping to enhance coding and compression efficiency.
This approach improves coding efficiency and maintains sound quality by adaptively selecting subband decomposition schemes that resist interference, optimizing scale factors for spectral envelopes.
Smart Images

Figure 2025528736000001_ABST
Abstract
Description
[Technical Field]
[0001] This application claims priority to Chinese Patent Application No. 202210894324.9, filed on July 27, 2022, entitled "Audio Signal Processing Method and Apparatus, Storage Medium, and Computer Program Product," and Chinese Patent Application No. 202211139940.X, filed on September 19, 2022, entitled "Audio Signal Processing Method and Apparatus, Storage Medium, and Computer Program Product," both of which are incorporated herein by reference in their entirety.
[0002] This application relates to the field of audio encoding / decoding, and in particular to an audio signal processing method and apparatus, a storage medium, and a computer program product. [Background technology]
[0003] As the quality of life improves, people increasingly demand high-quality audio. In order to better transmit audio signals within a limited bandwidth, an encoder typically performs data compression on the audio signal to obtain a bitstream. The bitstream is then transmitted to a decoder. The decoder then decodes the received bitstream to reconstruct the audio signal. The reconstructed audio signal is used for playback. However, in the process of compressing an audio signal, the sound quality of the audio signal may be affected. Therefore, how to improve the compression efficiency of an audio signal while maintaining the sound quality of the audio signal is an urgent technical problem that needs to be solved. Summary of the Invention
[0004] This application provides an audio signal processing method and apparatus, a storage medium, and a computer program product for improving coding effect and compression efficiency. The technical solutions are as follows:
[0005] According to a first aspect, there is provided an audio signal processing method, the method comprising: The method includes: performing subband decomposition on the audio signal separately based on a plurality of subband decomposition schemes and cutoff subbands corresponding to the plurality of subband decomposition schemes to obtain a plurality of candidate subband sets, the plurality of candidate subband sets corresponding one-to-one to the plurality of subband decomposition schemes, each candidate subband set including a plurality of subbands; determining a total scale value for each candidate subband set based on spectral values of the audio signal in subbands included in the candidate subband set, an encoding bit rate of the audio signal, and subband bandwidths of the subbands included in the candidate subband set; and selecting one candidate subband set from the plurality of candidate subband sets as a target subband set based on the total scale value for each candidate subband set, wherein each subband included in the target subband set has a scale factor used to shape the spectral envelope of the audio signal.
[0006] In this application, an optimal subband decomposition scheme is selected from multiple subband decomposition schemes based on the characteristics of the audio signal. In other words, the subband decomposition scheme has signal adaptive properties and can adapt to the coding bit rate of the audio signal to improve interference resistance. Specifically, the audio signal is separately divided into multiple subband decomposition schemes, and a total scale factor corresponding to each subband decomposition scheme is determined based on the spectral values of the audio signal in the subbands obtained through the division, the bandwidth of each subband, and the coding bit rate of the audio signal. An optimal target subband decomposition scheme is selected based on the total scale factor to obtain an optimal subband set. Then, spectral envelope shaping is performed based on the scale factor of each subband in the optimal subband set, thereby improving coding efficiency and compression efficiency.
[0007] Optionally, selecting one candidate subband set from the plurality of candidate subband sets as the target subband set based on the total scale value of each candidate subband set includes: determining a candidate subband set having a minimum total scale value among the plurality of candidate subband sets as a target subband set; This includes:
[0008] Optionally, determining a total scale value for each candidate subband set based on spectral values of the audio signal in subbands included in the candidate subband set, a coding bit rate of the audio signal, and subband bandwidths of the subbands included in the candidate subband set comprises: For a first candidate subband set among the plurality of candidate subband sets, determine a scale factor for each subband included in the first candidate subband set based on spectral values of the audio signal in the subbands included in the first candidate subband set, the first candidate subband set being any one of the plurality of candidate subband sets; determining a total scale factor for the first candidate subband set based on the coding bit rate of the audio signal and the scale factor and subband bandwidth of each subband included in the first candidate subband set; This includes:
[0009] Optionally, determining a scale factor for each subband included in the first candidate subband set based on spectral values of the audio signal within the subbands included in the first candidate subband set comprises: For a first subband included in a first candidate subband set, obtain a maximum value among absolute values of all spectral values of the audio signal in the first subband, the first subband being any subband in the first candidate subband set; determining a scale factor for the first subband based on the maximum value; This includes:
[0010] Optionally, the coding bit rate of the audio signal is not lower than a first bit rate threshold and / or the energy concentration of the audio signal is greater than a concentration threshold.
[0011] determining a total scale factor for the first candidate subband set based on a coding bit rate of the audio signal and a scale factor and a subband bandwidth of each subband included in the first candidate subband set; determining an energy smoothing reference value based on an encoding bit rate of the audio signal and a second bit rate threshold; determining a total energy value for each subband in the first candidate subband set based on the energy smoothing criterion and the scale factor and subband bandwidth of each subband in the first candidate subband set; summing the total energy values of the subbands included in the first candidate subband set to obtain a total scale value for the first candidate subband set; This includes:
[0012] Optionally, determining a total energy value for each subband in the first candidate subband set based on the energy smoothing reference value and the scale factor and subband bandwidth of each subband in the first candidate subband set comprises: For a first subband included in a first candidate subband set, determine a larger value of the scale factor of the first subband and the energy smoothing reference value as a reference scale value of the first subband, where the first subband is any subband in the first candidate subband set; determining a product of the reference scale value of the first subband and the subband bandwidth of the first subband as a total energy value of the first subband; This includes:
[0013] Optionally, the coding bit rate of the audio signal is lower than a first bit rate threshold and the energy concentration of the audio signal is not greater than a concentration threshold.
[0014] determining a total scale factor for the first candidate subband set based on a coding bit rate of the audio signal and a scale factor and a subband bandwidth of each subband included in the first candidate subband set; determining an energy smoothing reference value based on an encoding bit rate of the audio signal and a second bit rate threshold; determining scale difference values for the subbands included in the first candidate subband set based on the energy smoothing reference value and the scale factors of each subband included in the first candidate subband set, the scale difference values indicating differences between the scale factors of the corresponding subbands and the scale factors of the adjacent subbands of the corresponding subbands; determining a total scale value for the first candidate subband set based on the scale difference values of the subbands included in the first candidate subband set and the subband bandwidth of each subband; This includes:
[0015] Optionally, determining scale difference values for subbands in the first candidate subband set based on the energy smoothing reference value and the scale factors of each subband in the first candidate subband set comprises: For a first subband included in a first candidate subband set, determine a first smoothed value, a second smoothed value, and a third smoothed value for the first subband based on the energy smoothing reference value, a scale factor of the first subband, and scale factors of adjacent subbands of the first subband, where the first subband is any subband in the first candidate subband set; determining a scaled difference value for the first subband based on the first smoothed value, the second smoothed value, and the third smoothed value for the first subband; This includes:
[0016] Optionally, determining the first smoothed value, the second smoothed value, and the third smoothed value for the first subband based on the energy smoothed reference value, the scale factor of the first subband, and the scale factors of adjacent subbands of the first subband includes: If the first subband is the first subband in the first candidate subband set, determine the larger value of the scale factor of the first subband and the energy smoothing reference value as the first smoothed value of the first subband; if the first subband is not the first subband in the first candidate subband set, determine the larger value of the scale factor of a previous subband adjacent to the first subband and the energy smoothing reference value as the first smoothed value of the first subband; determining a second smoothed value for the first subband as a larger value of the scale factor for the first subband and the energy smoothed reference value; If the first subband is the last subband in the first candidate subband set, determine the larger of the scale factor of the first subband and the energy smoothing reference value as the third smoothed value of the first subband; if the first subband is not the last subband in the first candidate subband set, determine the larger of the scale factor of the next subband adjacent to the first subband and the energy smoothing reference value as the third smoothed value of the first subband. This includes:
[0017] Optionally, determining the scale difference value for the first subband based on the first smoothed value, the second smoothed value, and the third smoothed value for the first subband includes: determining a first difference value and a second difference value for a first subband included in a first candidate subband set, the first difference value being an absolute value of a difference value between the first smoothed value and the second smoothed value for the first subband, and the second difference value being an absolute value of a difference value between the second smoothed value and the third smoothed value for the first subband, the first subband being any subband in the first candidate subband set; determining a scaled difference value for the first subband based on the first difference value and the second difference value for the first subband; This includes:
[0018] Optionally, determining a total scale value for the first candidate subband set based on the scale difference values of the subbands included in the first candidate subband set and the subband bandwidth of each subband includes: determining a smoothing weighting factor for each subband included in the first candidate subband set based on the number of subbands included in the first candidate subband set and the subband bandwidth of each subband; summing the smoothed weighting factors of the subbands included in the first candidate subband set to obtain a total smoothed weighting factor for the first candidate subband set; multiplying the scaled difference values of the subbands included in the first candidate subband set by the smoothing weighting factors to obtain weighted scaled difference values of the subbands included in the first candidate subband set; summing the weighted scale difference values of the subbands included in the first candidate subband set to obtain a total scale value for the first candidate subband set; Dividing the total scale value by the total smoothing weighting factor of the first candidate subband set to obtain a total scale value of the first candidate subband set; This includes:
[0019] Optionally, the method further comprises: When an encoding bit rate of the audio signal is lower than a first bit rate threshold, performing bandwidth detection on the spectrum of the audio signal to obtain a cut-off frequency of the audio signal; determining cutoff subbands corresponding to the plurality of subband division schemes based on the cutoff frequencies; This includes:
[0020] Optionally, the method further comprises: determining a last subband indicated by each of the plurality of subband decomposition schemes as a cutoff subband corresponding to each subband decomposition scheme when the encoding bit rate of the audio signal is not lower than a first bit rate threshold; Optionally, the method further comprises: performing a feature analysis on the spectrum of the audio signal to obtain a feature analysis result; determining a plurality of subband decomposition schemes from a plurality of candidate subband decomposition schemes based on the feature analysis result and the coding bit rate of the audio signal; This includes:
[0021] Optionally, the feature analysis result includes a subjective signal flag or an objective signal flag, the subjective signal flag indicating that the energy concentration of the audio signal is not greater than a concentration threshold, and the objective signal flag indicating that the energy concentration of the audio signal is greater than a concentration threshold.
[0022] Optionally, the audio signal has a frame length of 10 milliseconds and a sampling rate of 88.2 kilohertz or 96 kilohertz, or a frame length of 5 milliseconds and a sampling rate of 88.2 kilohertz or 96 kilohertz, or a frame length of 10 milliseconds and a sampling rate of 44.1 kilohertz or 48 kilohertz.
[0023] determining a plurality of subband decomposition schemes from a plurality of candidate subband decomposition schemes based on the feature analysis result and the coding bit rate of the audio signal; determining a first group of subband decomposition schemes from among the plurality of candidate subband decomposition schemes as the plurality of subband decomposition schemes when the encoding bit rate of the audio signal is lower than a first bit rate threshold and the feature analysis result includes a subjective signal flag; This includes:
[0024] The subband division method for the first group is as follows: { {0,1,2,3,4,6,8,10,13,16,20,24,28,33,38,45,52,61,70,79,88,100,112,127,142,160,178,196,217,238,259,280,480}, {0,1,2,3,5,7,9,12,15,18,22,26,30,35,41,48,56,65,74,84,94,106,118,134,150,166,184,202,220,240,260,280,480}, {0,1,2,3,4,5,7,9,11,14,17,21,25,29,34,40,46,52,60,68,76,86,98,110,126,144,162,180,200,224,250,280,480}, {0,2,4,6,8,12,16,21,26,31,36,41,46,51,56,61,66,71,77,83,89,95,103,111,121,131,147,163,179,203,240,280,480}, {0,1,2,3,5,7,9,12,15,19,23,27,32,37,43,49,57,66,76,86,98,110,125,140,158,176,194,216,238,264,290,320,480}, {0,1,2,3,5,7,10,13,17,21,25,30,35,41,47,54,62,70,80,90,102,114,130,146,162,180,198,218,240,264,290,320,480}, {0,1,2,4,6,8,11,14,18,22,26,30,36,42,50,58,66,76,88,100,112,128,144,160,182,204,226,256,286,316,352,400,480}, {0,1,2,4,6,8,11,14,18,22,26,30,36,42,50,58,68,78,90,102,116,132,148,166,186,208,234,262,292,324,360,400,480} }。
[0025] Optionally, the audio signal has a frame length of 10 milliseconds and a sampling rate of 88.2 kilohertz or 96 kilohertz, or a frame length of 5 milliseconds and a sampling rate of 88.2 kilohertz or 96 kilohertz, or a frame length of 10 milliseconds and a sampling rate of 44.1 kilohertz or 48 kilohertz.
[0026] determining a plurality of subband decomposition schemes from a plurality of candidate subband decomposition schemes based on the feature analysis result and the coding bit rate of the audio signal; determining a second group of subband decomposition schemes from the plurality of candidate subband decomposition schemes as the plurality of subband decomposition schemes when the encoding bit rate of the audio signal is not lower than a first bit rate threshold and / or the feature analysis result includes an objective signal flag; This includes:
[0027] The subband division scheme for the second group is as follows: { {0,1,2,3,4,5,6,7,8,10,12,14,16,19,22,26,30,35,40,45,50,57,64,73,82,92,102,112,124,136,148,160,480}, {0,1,2,3,4,5,7,9,11,13,15,18,21,24,28,33,38,44,50,57,64,73,82,93,104,116,128,140,155,170,185,200,480}, {0,1,2,3,4,6,8,10,13,16,20,24,28,33,38,45,52,61,70,79,88,100,112,127,142,160,178,196,217,238,259,280,480}, {0,1,2,4,6,10,14,18,22,26,30,34,42,50,58,66,74,84,96,108,120,136,152,168,192,216,240,272,304,336,376,424,480}, {0,1,2,4,6,10,14,18,26,34,42,50,62,74,86,98,112,128,144,160,176,196,216,236,256,280,304,328,352,384,416,448,480}, {0,80,92,104,112,120,128,136,144,148,152,156,160,164,168,172,176,180,184,188,192,196,200,208,216,224,232,240,248,256,268,280,480}, {0,200,212,224,232,240,248,256,264,268,272,276,280,284,288,292,296,300,304,308,312,316,320,328,336,344,352,360,368,376,388,400,480}, {0,320,332,344,356,364,372,380,384,388,392,396,400,404,408,412,416,420,424,428,432,436,440,444,448,452,456,460,464,468,472,476,480} }.
[0028] Optionally, the frame length of the audio signal is 5 milliseconds and the sampling rate is 44.1 kHz or 48 kHz.
[0029] determining a plurality of subband decomposition schemes from a plurality of candidate subband decomposition schemes based on the feature analysis result and the coding bit rate of the audio signal; determining a third group of subband decomposition schemes from among the plurality of candidate subband decomposition schemes as the plurality of subband decomposition schemes when the encoding bit rate of the audio signal is lower than a first bit rate threshold and the feature analysis result includes a subjective signal flag; This includes:
[0030] The subband division method of the third group is as follows: { {0,1,2,3,4,5,6,7,8,9,10,12,14,16,19,22,26,30,35,39,44,50,56,63,71,80,89,98,108,119,129,140,240}, {0,1,2,3,4,5,6,7,8,9,11,13,15,17,20,24,28,32,37,42,47,53,59,67,75,83,92,101,110,120,130,140,240}, {0,1,2,3,4,5,6,7,8,9,10,11,12,14,17,20,23,26,30,34,38,43,49,55,63,72,81,90,100,112,125,140,240}, {0,1,2,3,4,6,8,10,13,15,18,20,23,25,28,30,33,35,38,41,44,47,51,55,60,65,73,81,89,101,120,140,240}, {0,1,2,3,4,5,6,7,9,11,13,14,16,18,21,24,28,33,38,43,49,55,62,70,79,88,97,108,119,132,145,160,240}, {0,1,2,3,4,5,6,7,8,10,12,14,17,20,23,27,31,35,40,45,51,57,65,73,81,90,99,109,120,132,145,160,240}, {0,1,2,3,4,5,6,7,9,11,13,15,18,21,25,29,33,38,44,50,56,64,72,80,91,102,113,128,143,158,176,200,240}, {0,1,2,3,4,5,6,7,9,11,13,15,18,21,25,29,34,39,45,51,58,66,74,83,93,104,117,131,146,162,180,200,240} }.
[0031] Optionally, the frame length of the audio signal is 5 milliseconds and the sampling rate is 44.1 kHz or 48 kHz.
[0032] determining a plurality of subband decomposition schemes from a plurality of candidate subband decomposition schemes based on the feature analysis result and the coding bit rate of the audio signal; determining a fourth group of subband decomposition schemes from the plurality of candidate subband decomposition schemes as the plurality of subband decomposition schemes when the encoding bit rate of the audio signal is not lower than the first bit rate threshold and / or the feature analysis result includes an objective signal flag; This includes:
[0033] The subband division method for the fourth group is as follows: { {0,1,2,3,4,5,6,7,8,9,10,11,12,13,14,15,16,17,18,19,20,21,22,23,24,26,28,30,32,34,37,40,120}, {0,1,2,3,4,5,6,7,8,9,10,11,12,13,14,15,16,18,20,22,24,26,28,30,32,34,36,38,41,44,47,50,120}, {0,1,2,3,4,5,6,7,8,9,10,11,12,14,16,18,20,22,24,26,28,31,34,37,40,44,48,52,56,60,65,70,120}, {0,1,2,3,4,5,6,7,8,9,10,11,12,13,15,17,19,21,24,27,30,34,38,42,48,54,60,68,76,84,94,106,120}, {0,1,2,3,4,5,6,7,8,10,12,14,16,19,22,25,28,32,36,40,44,49,54,59,64,70,76,82,88,96,104,112,120}, {0,20,23,26,28,30,32,34,36,37,38,39,40,41,42,43,44,45,46,47,48,49,50,52,54,56,58,60,62,64,67,70,120}, {0,50,53,56,58,60,62,64,66,67,68,69,70,71,72,73,74,75,76,77,78,79,80,82,84,86,88,90,92,94,97,100,120}, {0,80,83,86,89,91,93,95,96,97,98,99,100,101,102,103,104,105,106,107,108,109,110,111,112,113,114,115,116,117,118,119,120} }.
[0034] Optionally, the audio signal is a dual channel signal.
[0035] The method further comprises: determining a first total scale value based on the scale factor and subband bandwidth of each subband included in the target subband set; performing mid / side stereo transform coding on the spectrum of the dual-channel signal to obtain a transformed spectrum of the dual-channel signal; determining a transformed scale factor for each subband in the target subband set based on the transformed spectral values of the dual-channel signal in each subband included in the target subband set; determining a second total scale value based on the transformed scale factor and the subband bandwidth of each subband included in the target subband set; determining the dual-channel signal as a signal to be encoded if the first total scale value is not greater than the second total scale value; This includes:
[0036] Optionally, the method further comprises: determining the converted dual-channel signal as a signal to be encoded if the first total scale value is greater than the second total scale value, and the encoding bit rate of the audio signal is not lower than a first bit rate threshold and / or the energy concentration of the audio signal is greater than a concentration threshold; This includes:
[0037] Optionally, the scale factors include a left channel scale factor and a right channel scale factor.
[0038] The method further comprises: When the first total scale value is greater than the second total scale value, the encoding bit rate of the audio signal is lower than a first bit rate threshold, and the energy concentration of the audio signal is not greater than a concentration threshold, a left channel scale factor and a right channel scale factor of each subband included in the target subband set are calculated based on the left channel scale factor and the right channel scale factor of each subband included in the target subband set. channel Scale Factor and Right channel determining a difference value between the scale factors; determining a subband center frequency for each subband included in the target subband set based on the initial frequency and the cutoff frequency of each subband included in the target subband set; If there are any subbands in the target subband set that are larger than the difference threshold, channel Scale Factor and Right channel determining the dual-channel signal as a signal to be coded if there is at least one subband having a difference value between the scale factor and the first subband and having a subband center frequency within the first range; This includes:
[0039] Optionally, the method further comprises: determining the transformed dual-channel signal as the signal to be coded if at least one subband does not exist in the target subband set; This includes:
[0040] According to a second aspect, there is provided an audio signal processing device having a function of implementing the audio signal processing method of the first aspect, the audio signal processing device including one or more modules configured to implement the audio signal processing method of the first aspect.
[0041] According to a third aspect, there is provided an audio signal processing device. The audio signal processing device includes a processor and a memory. The memory is configured to store a program used to execute the audio signal processing method of the first aspect and to store data used to implement the audio signal processing method of the first aspect. The processor is configured to execute the program stored in the memory. The audio signal processing device may further include a communication bus, the communication bus being configured to establish a connection between the processor and the memory.
[0042] According to a fourth aspect, there is provided a computer-readable storage medium storing instructions that, when executed on a computer, enable the computer to perform the audio signal processing method of the first aspect.
[0043] According to a fifth aspect, there is provided a computer program product comprising instructions which, when run on a computer, enable the computer to carry out the audio signal processing method of the first aspect.
[0044] The technical effects achieved in the second, third, fourth and fifth aspects are the same as those achieved by the corresponding technical means in the first aspect, and the details will not be described again here. [Brief explanation of the drawings]
[0045] [Figure 1] FIG. 1 is a diagram of a Bluetooth interconnection scenario according to an embodiment of the present application. [Figure 2] 1 is a diagram of a system architecture relating to an audio signal processing method according to an embodiment of the present application; [Figure 3A] 3A and 3B are diagrams of an overall audio encoding / decoding framework according to one embodiment of this application. [Figure 3B] 3A and 3B are diagrams of an overall audio encoding / decoding framework according to one embodiment of this application. [Figure 4] FIG. 1 is a diagram of an electronic device configuration according to one embodiment of the present application. [Figure 5] 1 is a flowchart of an audio signal processing method according to an embodiment of the present application. [Figure 6] 1 is a diagram of the relationship between subbands and initial frequencies shown in a first group of subband division schemes according to one embodiment of the present application. [Figure 7] FIG. 10 is a diagram of the relationship between subbands and initial frequencies shown in the second group of subband division schemes according to one embodiment of the present application. [Figure 8] FIG. 10 is a diagram of the relationship between subbands and initial frequencies shown in the third group of subband division schemes according to one embodiment of the present application. [Figure 9] FIG. 10 is a diagram of the relationship between subbands and initial frequencies shown in the fourth group of subband division schemes according to one embodiment of the present application. [Figure 10] 1 is a flowchart of a method for determining whether an MS conversion has gain, according to one embodiment of the present application. [Figure 11] 1 is a diagram of the configuration of an audio signal processing apparatus according to an embodiment of the present application; DETAILED DESCRIPTION OF THE INVENTION
[0046] To make the objectives, technical solutions and advantages of the embodiments of this application clearer, the following further describes the implementation of this application in detail with reference to the accompanying drawings.
[0047] First, an implementation environment and background knowledge related to the embodiments of this application will be described.
[0048] As wireless Bluetooth devices, such as true wireless stereo (TWS) headsets, smart speakers, and smart watches, become widely popular and used in people's daily lives, people's demands for high-quality audio playback experiences in various scenarios become increasingly urgent, especially in environments where Bluetooth signals are vulnerable to interference, such as subways, airports, and train stations. This is because in Bluetooth interconnection scenarios, the Bluetooth channel connecting an audio transmitting device and an audio receiving device limits the size of transmission data. An audio encoder in the audio transmitting device performs data compression on an audio signal before transmitting the audio signal to the audio receiving device. The compressed audio signal can only be played back after being decoded by an audio decoder in the audio receiving device. It is clear that the popularity of wireless Bluetooth devices also promotes the flourishing of various audio codecs.
[0049] Currently, Bluetooth audio codecs include sub-band coding (SBC), advanced audio coding (AAC), aptX series encoders, low-latency high-definition audio codec (LHDC), low-energy low-latency LC3 audio codec, and LC3plus.
[0050] It should be understood that the audio provided in the embodiments of this application Signal Processing The method can be applied to an audio transmitting device (ie, encoder side) and an audio receiving device (ie, decoder side) in a Bluetooth interconnection scenario.
[0051] FIG. 1 is a diagram of a Bluetooth interconnection scenario according to one embodiment of this application. Please refer to FIG. 1. The audio transmitting device in the Bluetooth interconnection scenario may be a mobile phone, a computer, a tablet computer, or the like. The computer may be a notebook computer, a desktop computer, or the like, and the tablet computer may be a handheld tablet computer, an in-car tablet computer, or the like. The audio receiving device in the Bluetooth interconnection scenario may be a TWS headset, a smart speaker, wireless headphones, a wireless neckband headset, a smart watch, smart glasses, a smart in-car device, or the like. In some other embodiments, the audio receiving device in the Bluetooth interconnection scenario may instead be a mobile phone, a computer, a tablet computer, or the like.
[0052] In addition to the Bluetooth interconnection scenario, the audio Signal Processing The method can also be applied to other device interconnection scenarios. In other words, the system architecture and service scenarios described in the embodiments of this application are intended to more clearly explain the technical solutions in the embodiments of this application, and do not constitute limitations on the technical solutions provided in the embodiments of this application. It is understood by those skilled in the art that the technical solutions provided in the embodiments of this application can also be applied to similar technical problems as system architectures evolve and new service scenarios emerge.
[0053] Fig. 2 is a diagram of a system architecture related to an audio signal processing method according to one embodiment of this application. Please refer to Fig. 2. The system includes an encoder side and a decoder side. The encoder side includes an input module, an encoding module, and a transmission module. The decoder side includes a receiving module, an input module, a decoding module, and a playback module.
[0054] At the encoder side, a user selects one of two encoding modes, a low-delay encoding mode and a high-quality encoding mode, based on a usage scenario. The encoding frame lengths of these two encoding modes are 5 ms and 10 ms, respectively. For example, if the usage scenario is playing a game, live broadcasting, or making a phone call, the user can select the low-delay encoding mode; or if the usage scenario is enjoying music through a headset or speaker, the user can select the high-quality encoding mode. The user also needs to provide the audio signal to be encoded (pulse code modulation (PCM) data shown in FIG. 2) to the encoder side. In addition, the user also needs to set a target bit rate of the bit stream obtained through encoding, that is, the encoding bit rate of the audio signal. A higher target bit rate indicates better sound quality but poorer interference resistance performance of the bit stream in a short-distance transmission process. A lower target bit rate indicates poorer sound quality but better interference resistance performance of the bit stream in a short-distance transmission process. In brief, the input module on the encoder side takes the encoding frame length, the encoding bit rate, and the audio signal to be encoded, submitted by the user.
[0055] An input module on the encoder side inputs user-submitted data into a frequency-domain encoder of the encoding module.
[0056] The frequency domain encoder of the encoding module performs encoding based on the received data to obtain a bitstream. The frequency domain encoder analyzes the audio signal to be encoded to obtain signal characteristics (including mono / dual channel signal, stable / unstable signal, full bandwidth / narrow bandwidth signal, and subjective / objective signal, etc.). The audio signal enters a corresponding encoding processing sub-module based on the signal characteristics and bit rate level (i.e., encoding bit rate). The encoding processing sub-module encodes the audio signal and packages a bitstream packet header (including sampling rate, channel number, encoding mode, frame length, etc.) to finally obtain a bitstream.
[0057] A transmitting module on the encoder side transmits the bitstream to the decoder side. Optionally, the transmitting module may be the short-range transmitting module shown in Fig. 2 or another type of transmitting module, which is not limited in the embodiment of this application.
[0058] On the decoder side, after receiving the bitstream, the decoder side receiving module sends the bitstream to the frequency domain decoder of the decoding module, and notifies the decoder side input module to obtain the configured bit depth, the configured channel decoding mode, or the like. Optionally, the receiving module may be the short-range receiving module shown in FIG. 2 or another type of receiving module, which is not limited in the embodiment of this application.
[0059] An input module on the decoder side inputs the obtained information, such as bit depth and audio channel decoding mode, into a frequency domain decoder of the decoding module.
[0060] The frequency domain decoder of the decoding module decodes the bitstream based on the bit depth, channel decoding mode, etc. to obtain the required audio data (PCM data shown in Figure 2), and sends the obtained audio data to the playback module. The playback module plays the audio. The audio channel decoding mode indicates the channel that needs to be decoded.
[0061] 3A and 3B are diagrams of an overall audio encoding / decoding framework according to one embodiment of this application. Please refer to Fig. 3A and Fig. 3B. The encoding procedure at the encoder side includes the following steps:
[0062] (1) PCM input module PCM data is input. The PCM data is mono-channel data or dual-channel data, and the bit depth can be 16 bits, 24 bits, 32-bit floating-point numbers, or 32-bit fixed-point numbers. Optionally, the PCM input module converts the input PCM data to the same bit depth, for example, 24 bits, performs deinterleaving on the PCM data, and then places the deinterleaved PCM data on the left channel and the right channel.
[0063] (2) Low-latency analysis window addition module and modified discrete cosine transform (MDCT) module A low-delay analysis window is added to the PCM data processed in step (1), and an MDCT transform is performed to obtain spectral data in the MDCT domain. The window is added to prevent spectral leakage.
[0064] (3) MDCT domain signal analysis module and adaptive bandwidth detection module The MDCT domain signal analysis module is effective in full bitrate scenarios, and the adaptive bandwidth detection module is activated in low bitrates (e.g., bitrates < 150 kbps / channel). First, bandwidth detection is performed on the spectral data in the MDCT domain obtained in step (2) to obtain the cutoff frequency or effective bandwidth. Then, signal analysis is performed on the spectral data within the effective bandwidth, i.e., whether the frequency distribution is concentrated or flat is analyzed to obtain energy concentration. Then, based on the energy concentration, a flag is obtained indicating whether the audio signal to be coded is an objective signal or a subjective signal (the flag for the objective signal is 1, and the flag for the subjective signal is 0). If the audio signal is an objective signal, spectral noise shaping (SNS) processing and MDCT spectrum smoothing are not performed on the scale factor at low bitrates because this would reduce the coding effect of the objective signal. Then, based on the bandwidth detection result, the flag of the subjective signal and the flag of the objective signal, it is determined whether to perform a subband cutoff operation in the MDCT domain. If the audio signal is an objective signal, the subband cutoff operation is not performed; if the audio signal is a subjective signal and the bandwidth detection result is identified as 0 (within the full bandwidth), the subband cutoff operation is determined based on the bit rate; if the audio signal is a subjective signal and the bandwidth detection result is not identified as 0 (i.e., the bandwidth is less than half of the limited bandwidth of the sampling rate), the subband cutoff operation is determined based on the bandwidth detection result.
[0065] (4) Subband division selection and scale factor calculation module Based on the bit rate level, the subjective signal flag, the objective signal flag, and the cutoff frequency obtained in step (3), an optimal subband decomposition scheme is selected from multiple subband decomposition schemes, and the total number of subbands for encoding the audio signal is obtained. In addition, the spectral envelope is obtained through calculation, i.e., the scale factor corresponding to the selected subband decomposition scheme is calculated.
[0066] (5) MS channel conversion module For the dual-channel PCM data, a joint encoding decision is made based on the scale factors calculated in step (4), ie, whether to perform MS channel transform on the left channel data and the right channel data.
[0067] (6) Spectral smoothing module and scale factor-based spectral noise shaping module The spectral smoothing module performs MDCT spectral smoothing based on a low bit rate setting (e.g., bit rate < 150 kbps / channel), and the spectral noise shaping module performs spectral noise shaping on the data on which the spectral smoothing is performed based on a scale factor to obtain an adjustment factor used to quantize the spectral values of the audio signal. The low bit rate setting is controlled by the low bit rate determination module. When the low bit rate setting is not met, the spectral smoothing and spectral noise shaping do not need to be performed.
[0068] (7) Scale Factor Encoding Module Differential or entropy coding is performed on the scale factors of the multiple subbands based on the distribution of the scale factors.
[0069] (8) Bit allocation and MDCT spectrum quantization and entropy coding module Based on the scale factors obtained in step (4) and the adjustment factors obtained in step (6), the coding is controlled to be in a constant bit rate (CBR) coding mode according to a coarse estimation and fine estimation bit allocation strategy, and quantization and entropy coding are performed on the MDCT spectral values.
[0070] (9) Residual coding module If the bit consumption in step (8) does not reach the target bits, further importance sorting is performed on the uncoded subbands, and bits are preferably allocated to coding the MDCT spectral values of the important subbands.
[0071] (10) Stream packet header information packaging module The packet header information includes audio sampling rate (e.g., 44.1 kHz / 48 kHz / 88.2 kHz / 96 kHz), channel information (e.g., mono channel and dual channel), coding frame length (e.g., 5 ms and 10 ms), and coding mode (e.g., time domain mode, frequency domain mode, time domain-frequency domain mode, or frequency domain-time domain mode).
[0072] (11) Bitstream transmission module The bitstream includes a packet header, side information, and a payload. The packet header carries packet header information, which is as described in step (10). The side information includes, for example, an encoded bitstream of scale factors, information about a selected subband decomposition scheme, cutoff frequency information, a low bitrate flag, joint encoding decision information (i.e., MS transform flag), and a quantization step. The payload includes an encoded bitstream and a residual encoded bitstream of the MDCT spectrum.
[0073] The decoding procedure at the decoder side includes the following steps:
[0074] (1) Stream packet header information parsing module The stream packet header information parsing module parses packet header information from the received bitstream, where the packet header information includes information such as sampling rate, channel information, encoding frame length, and encoding mode of the audio signal, and obtains the encoding bitrate through calculation based on the bitstream size, sampling rate, and encoding frame length, i.e., obtains bitrate level information.
[0075] (2) Scale factor decoding module The scale factor decoding module decodes side information from the bitstream, including, for example, information about the selected subband decomposition scheme, cutoff frequency information, low bit rate flags, joint coding decision information, quantization steps, and subband scale factors.
[0076] (3) Scale factor-based spectral noise shaping module At low bit rates (e.g., coding bit rates below 300 kbps, i.e., 150 kbps / channel), it is necessary to further perform spectral noise shaping based on the scale factor to obtain the adjustment factor, which is Decryption The low bit rate setting is controlled by the low bit rate decision module. When the low bit rate setting is not met, there is no need to perform spectral noise shaping.
[0077] (4) MDCT spectrum decoding module and residual decoding module The MDCT spectrum decoding module decodes the MDCT spectrum data in the decoded bitstream based on the information about the subband division scheme, the quantization step information, and the scale factor obtained in step (2). At a low bit rate level, hole padding is performed, and if there are still bits remaining to be obtained through calculation, the residual decoding module performs residual decoding to obtain the MDCT spectrum data of another subband, thereby obtaining the final MDCT spectrum data.
[0078] (5) LR channel conversion module Based on the side information obtained in step (2), if it is determined according to the joint coding decision information that a dual-channel joint coding mode (e.g., the coding bit rate is 300 kbps / channel or more and the sampling rate is higher than 88.2 kHz) is used instead of the decoding low-energy mode, an LR channel transform is performed on the MDCT spectrum data obtained in step (4).
[0079] (6) Inverse MDCT transform module, low-delay synthesis window addition module, and overlap addition module Based on step (4) and step (5), the inverse MDCT transform module performs an inverse MDCT transform on the acquired MDCT spectrum data to obtain a time-domain aliased signal. Then, the low-delay synthesis window module adds a low-delay synthesis window to the time-domain aliased signal. The overlap-and-add module overlaps the time-domain aliased buffer signals of the current frame and the previous frame to obtain a PCM signal, i.e., obtains final PCM data based on the overlap-and-add method.
[0080] (7) PCM output module The PCM output module outputs the PCM data of the corresponding channel based on the configured bit depth and channel decoding mode.
[0081] It should be noted that the audio encoding / decoding framework shown in Figures 3A and 3B is only used as an example of a terminal in the embodiment of this application and is not intended to limit the embodiment of this application. Those skilled in the art can obtain other encoding / decoding frameworks based on Figures 3A and 3B.
[0082] 4 is a diagram of an electronic device configuration according to one embodiment of the present application, optionally any of the devices shown in FIG. 1, including one or more processors 401, a communication bus 402, a memory 403, and one or more communication interfaces 404.
[0083] The processor 401 may be a general-purpose central processing unit (CPU), a network processor (NP), a microprocessor, or one or more integrated circuits configured to implement the solutions of this application, such as an application-specific integrated circuit (ASIC), a programmable logic device (PLD), or a combination thereof. Optionally, the PLD may be a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.
[0084] communication bus 402 is configured to transmit information between the aforementioned components. Optionally, the communication bus 402 may be classified as an address bus, a data bus, a control bus, etc. For ease of representation, the diagram uses only one thick line to represent a bus, but this does not imply that there is only one bus or only one type of bus.
[0085] Optionally, memory 403 is read-only memory (ROM), random access memory (RAM), electrically erasable programmable read-only memory (EEPROM), optical disk (including compact disc read-only memory (CD-ROM), compact disc, laser disc, digital versatile disc, Blu-ray disc, or the like), magnetic disk storage medium or other magnetic storage device, or any other medium that can be used to carry or store expected program code in the form of instructions or data structures and that is accessible to a computer, but is not limited to such. Memory 403 may exist independently and be connected to processor 401 via communication bus 402, or memory 403 may be integrated with processor 401.
[0086] The communication interface 404 is configured to communicate with another device or a communication network by using any device such as a transceiver. The communication interface 404 may include a wired communication interface or, optionally, a wireless communication interface. The wired communication interface may be, for example, an Ethernet interface. Optionally, the Ethernet interface may be an optical interface, an electrical interface, or a combination thereof. The wireless communication interface may be a wireless local area network (WLAN) interface, a cellular network communication interface, a combination thereof, or the like.
[0087] Optionally, in some embodiments, the electronic device includes multiple processors, such as processor 401 and processor 405 shown in Figure 4. Each of these processors may be a single-core processor or a multi-core processor. Optionally, a processor herein may be one or more devices, circuits, and / or processing cores for processing data (e.g., computer program instructions).
[0088] In a specific implementation, in one embodiment, the electronic device further includes an output device 406 and an input device 407. The output device 406 communicates with the processor 401 and can display information in multiple ways. For example, the output device 406 is a liquid crystal display (LCD), a light emitting diode (LED) display device, a cathode ray tube (CRT) display device, or a projector. The input device 407 communicates with the processor 401 and can receive user input in multiple ways. For example, the input device 407 is a mouse, a keyboard, a touchscreen device, or a sensing device.
[0089] In some embodiments, the memory 403 is configured to store program code 410 for implementing the solution of this application, and the processor 401 can execute the program code 410 stored in the memory 403. The program code includes one or more software modules, and the electronic device can implement the audio signal processing method provided in the embodiment of FIG. 5 below by using the processor 401 and the program code 410 in the memory 403.
[0090] 5 is a flowchart of an audio signal processing method according to an embodiment of this application, which is applied on the encoder side. Please refer to FIG. 5. The method includes the following steps:
[0091] Step 501: Based on a plurality of subband decomposition schemes and cutoff subbands corresponding to the plurality of subband decomposition schemes, perform subband decomposition on the audio signal separately to obtain a plurality of candidate subband sets, where the plurality of candidate subband sets correspond one-to-one to the plurality of subband decomposition schemes, and each candidate subband set includes a plurality of subbands.
[0092] In an embodiment of this application, the encoder side performs subband decomposition on the audio signal separately based on the multiple subband decomposition schemes and the cutoff subbands corresponding to the multiple subband decomposition schemes to obtain multiple candidate subband sets, in order to select the optimal subband decomposition scheme from the multiple subband decomposition schemes.
[0093] Take any one of multiple subband decomposition schemes as an example. The total number of subbands indicated by the subband decomposition schemes is 32, and the number of cutoff subbands corresponding to these subband decomposition schemes is 16, indicating that the cutoff frequency of the audio signal is within the 16th subband. For example, the full bandwidth of the audio signal is 16 kilohertz (kHz). The cutoff subband indicates that the cutoff frequency of the audio signal is 5 kHz. After the subband decomposition is performed on the audio signal based on the subband decomposition scheme, the obtained candidate subband set includes a total of 16 subbands, and the frequency range covered by the 16 subbands is from 0 kHz to 5 kHz, i.e., the range of [0, cutoff frequency] is covered.
[0094] It should be noted that the subband decomposition process is performed for each audio frame. The audio signal described in this specification can be regarded as an audio frame. Obviously, the encoder side can perform subband decomposition for each audio frame according to this solution.
[0095] In the embodiment of this application, there are several implementations in which the encoder side obtains the cutoff subbands, and one of these implementations is described here.
[0096] When the encoding bit rate of the audio signal is lower than the first bit rate threshold, the encoder performs bandwidth detection on the spectrum of the audio signal to obtain a cutoff frequency of the audio signal. The encoder determines cutoff subbands corresponding to the plurality of subband division schemes based on the cutoff frequencies. It should be understood that when the encoding bit rate is low, the number of coding bits that can be allocated is small. Therefore, the encoder determines the cutoff frequency through bandwidth detection and then determines the cutoff subbands. In this case, spectral values exceeding the cutoff frequency are not subsequently encoded, thereby ensuring the encoding effect and satisfying the encoding bit rate requirement.
[0097] There are several ways that the encoder can perform bandwidth detection. In one implementation, since the values of frequencies in the spectrum of the audio signal after the cutoff frequency are zero, the encoder traverses the frequency values in the spectrum from high to low, and the first traversed frequency value that is greater than the energy threshold is the cutoff frequency of the audio signal.
[0098] Optionally, the encoder side takes the logarithm (e.g., log10) of the frequency values in the spectrum, traverses the frequency values obtained after taking the logarithm from high to low frequencies, and determines the first traversed frequency value obtained after taking the logarithm that is greater than an energy threshold as the cutoff frequency of the audio signal. Optionally, the energy threshold is -50 dB, -80 dB, or another value.
[0099] In addition, when the audio signal is a mono-channel signal, the encoder side performs bandwidth detection on the mono-channel spectrum of the audio signal to obtain the cutoff frequency of the audio signal. When the audio signal is a dual-channel signal, the encoder side performs bandwidth detection separately on the left-channel spectrum and the right-channel spectrum of the audio signal to obtain the left-channel cutoff frequency and the right-channel cutoff frequency. If the left-channel cutoff frequency does not match the right-channel cutoff frequency, the encoder side determines the larger value of the left-channel cutoff frequency and the right-channel cutoff frequency as the cutoff frequency of the audio signal. If the left-channel cutoff frequency matches the right-channel cutoff frequency, the encoder side determines the left-channel cutoff frequency as the cutoff frequency of the audio signal.
[0100] In some other embodiments, the encoder side may instead perform bandwidth detection on the spectrum in another manner, which is not a limitation of this solution.
[0101] Optionally, after obtaining the cutoff frequency of the audio signal, the encoder side determines cutoff sub-bands corresponding to the plurality of sub-band division schemes respectively based on the positions of the cutoff frequencies in the complete bandwidth of the audio signal.
[0102] For example, the cutoff frequency is located at the 30th frequency in the complete bandwidth of the audio signal, and the 30th frequency is located within the kth subband among the subbands represented by a subband decomposition scheme, where the cutoff subband corresponding to the subband decomposition scheme is k.
[0103] Optionally, in an embodiment of this application, the first bitrate threshold is 150 kbps or another value. Hereinafter, for the sake of explanation, an example in which the first bitrate threshold is 150 kbps is used. Optionally, in this embodiment of this application, the encoding bitrate of the audio signal is the encoding bitrate of a single channel, i.e., the encoding bitrate of a single channel is compared with the first bitrate threshold. In some other embodiments, the first bitrate threshold may instead be another value. For example, the first bitrate threshold is 150 kbps. When the audio signal is a dual-channel signal, the encoding bitrate of the audio signal is the encoding bitrate of the left channel or the encoding bitrate of the right channel. The encoding bitrate of the left channel is usually the same as the encoding bitrate of the right channel. In this case, the encoder side only needs to compare the encoding bitrate of the left channel with 150 kbps.
[0104] Obviously, in some other embodiments, when the audio signal is a dual-channel signal, the encoding bit rate of the audio signal is the encoding bit rate of the dual channel, and correspondingly, the first bit rate threshold is 300 kbps.
[0105] Optionally, if the encoding bit rate of the audio signal is not lower than the first bit rate threshold, the encoder side determines the last subband indicated by each of the multiple subband division schemes as the cutoff subband corresponding to each subband division scheme. It should be understood that when the encoding bit rate is high, the number of encoding bits that can be allocated is large. Therefore, even if the encoder side does not perform bandwidth detection, the requirement for the encoding bit rate can still be met, and the encoding efficiency can be further improved to a certain extent. Indeed, in some other embodiments, the encoder side may instead perform bandwidth detection on the spectrum of the audio signal whose encoding bit rate is not lower than the first bit rate threshold.
[0106] In an embodiment of this application, before separately performing subband decomposition on the audio signal based on a plurality of subband decomposition schemes and the cutoff subbands corresponding to the plurality of subband decomposition schemes, the encoder side performs feature analysis on the spectrum of the audio signal to obtain the feature analysis result, and determines a plurality of subband decomposition schemes from a plurality of candidate subband decomposition schemes based on the feature analysis result and the coding bit rate of the audio signal. In other words, the encoder side pre-selects a plurality of subband decomposition schemes from a plurality of candidate subband decomposition schemes through frequency domain feature analysis, and then selects an optimal subband decomposition scheme from the plurality of subband decomposition schemes.
[0107] Optionally, the feature analysis result includes a subjective signal flag or an objective signal flag, where the subjective signal flag indicates that the energy concentration of the audio signal is not greater than a concentration threshold, and the objective signal flag indicates that the energy concentration of the audio signal is greater than the concentration threshold. In other words, the feature analysis includes subjective and objective signal analyses, and the encoder pre-selects a plurality of subband decomposition schemes based on the subjective and objective signal analysis results and the encoding bit rate.
[0108] Below we describe the implementation of subjective and objective signal analysis.
[0109] In an embodiment of this application, the encoder side performs subjective and objective signal analysis based on the portion of the spectrum of the audio signal that does not exceed the cutoff frequency in order to reduce the amount of calculation and improve efficiency while ensuring accuracy.
[0110] The encoder side takes the logarithm of each frequency value in the spectrum that does not exceed the cutoff frequency to obtain a logarithmic result for each frequency. The encoder side normalizes the logarithmic result for each frequency to a dBFS scale to obtain a logarithmic result for each frequency on the dBFS scale. The encoder side determines a first frequency number and a second frequency number. The first frequency number is the total number of frequencies whose logarithmic results are not greater than an energy threshold on the dBFS scale, and the second frequency number is the total number of frequencies in the spectrum that do not exceed the cutoff frequency. The encoder side determines a ratio of the first frequency number to the second frequency number as an energy concentration of the audio signal. If the energy concentration of the audio signal is greater than the concentration threshold, the encoder side determines the audio signal to be an objective signal and outputs an objective signal flag. If the total energy of the audio signal is not greater than the concentration threshold, the encoder side determines the audio signal to be a subjective signal and outputs a subjective signal flag.
[0111] For example, the encoder side obtains the logarithmic result of each frequency by taking the logarithm to the base 10 of each value of the frequency in the spectrum that does not exceed the cutoff frequency according to equation (1).
[0112]
number
[0113] In equation (1), X(k) represents the value of the kth frequency, i.e., the kth spectrum value, cutOffFreq represents the frequency corresponding to the cutoff frequency, i.e., the second frequency number, abs() represents the absolute value, and Xlg(k) represents the logarithmic result of the kth frequency.
[0114] The encoder side normalizes the logarithmic result of each frequency to the dBFS scale according to equation (2) to obtain the logarithmic result of each frequency on the dBFS scale.
[0115]
number
[0116] In equation (2), XdBFS(k) denotes the logarithmic result of the kth frequency on the dBFS scale, and X max indicates the maximum spectral value in the spectrum that does not exceed the cutoff frequency.
[0117] The encoder side collects statistics on the total number of frequencies whose logarithmic result is not greater than -80 dB on the dBFS scale to obtain a first frequency number lowEnergyCnt, where -80 dB indicates the energy threshold, which is obtained through statistics collection or in other ways. The encoder side determines the energy concentration energyRate of the audio signal according to equation (3).
[0118]
number
[0119] The encoder side outputs the subjective and objective signal flag objFlag according to equation (4): When objFlag is 1, it indicates the objective signal flag; when objFlag is 0, it indicates the subjective signal flag.
[0120]
number
[0121] In equation (4), threshold denotes the concentration threshold.
[0122] In this embodiment of the present application, the concentration threshold is 0.6, and the concentration threshold is obtained through statistical collection or in other ways. For example, the concentration ratio threshold is a certain parameter obtained based on the signal distribution of different grades of bandwidth. Obviously, in some other embodiments, the concentration threshold may be other values instead.
[0123] It should be understood that the above example is used as one implementation of subjective and objective signal analysis, and is not intended to limit the embodiments of this application.
[0124] In another implementation, after obtaining the first frequency number and the second frequency number, the encoder determines the ratio of the second frequency number to the first frequency number as the energy concentration of the audio signal. If the energy concentration of the audio signal is less than the concentration threshold, the encoder determines that the audio signal is an objective signal and outputs an objective signal flag. If the total energy of the audio signal is not less than the concentration threshold, the encoder determines that the audio signal is a subjective signal and outputs a subjective signal flag. The concentration threshold in this implementation is the reciprocal of the concentration threshold in the previous implementation. In other words, if the proportion of the amount of non-background noise energy (i.e., the first frequency number) is less than the threshold, it indicates that the frequency domain characteristics of the objective signal are strong. The essence of this implementation is the same as that of the previous implementation.
[0125] In yet another implementation, the encoder does not normalize the logarithmic results of each frequency to the dBFS scale, but directly determines a third frequency number, which is the total number of frequencies whose logarithmic results are not greater than the energy threshold. The encoder then calculates the ratio of the third frequency number to the second frequency number as the energy concentration of the audio signal. Inside and Note that the energy threshold in this implementation is different from the energy threshold on the dBFS scale in the first implementation.
[0126] In yet another implementation, the encoder side does not take the base 10 logarithm of the values of the frequencies in the spectrum that do not exceed the cutoff frequency, but directly collects statistics on the total number of frequencies in the spectral range that does not exceed the cutoff frequency in the spectrum that do not exceed the energy threshold to obtain the fourth frequency number. The encoder side then calculates the ratio of the fourth frequency number to the second frequency number as the energy collection of the audio signal. Inside andNote that the energy and concentration thresholds in this implementation are different from those in some of the implementations described above.
[0127] It should be understood that taking the logarithm to the base 10 and normalizing to the dBFS scale is done because we are operating on different scales. The scale conversion is an optional operation on the encoder side, and the energy and concentration thresholds on different scales are different.
[0128] The following describes an implementation process in which the encoder pre-selects among multiple subband decomposition methods based on the feature analysis result and the encoding bit rate.
[0129] In this embodiment of the application, the feature analysis result includes a subjective signal flag or an objective signal flag. The frame length of the audio signal is 10 milliseconds (ms) and the sampling rate is 88.2 kilohertz (kHz) or 96 kHz, or the frame length of the audio signal is 5 ms and the sampling rate is 88.2 kHz or 96 kHz, or the frame length of the audio signal is 10 ms and the sampling rate is 44.1 kHz or 48 kHz. In this case, if the encoding bit rate of the audio signal is lower than a first bit rate threshold and the feature analysis result includes the subjective signal flag, the encoder side determines a first group of subband decomposition schemes from among a plurality of candidate subband decomposition schemes as the plurality of subband decomposition schemes. The subband decomposition schemes of the first group are as follows: { {0,1,2,3,4,6,8,10,13,16,20,24,28,33,38,45,52,61,70,79,88,100,112,127,142,160,178,196,217,238,259,280,480}, {0,1,2,3,5,7,9,12,15,18,22,26,30,35,41,48,56,65,74,84,94,106,118,134,150,166,184,202,220,240,260,280,480}, {0,1,2,3,4,5,7,9,11,14,17,21,25,29,34,40,46,52,60,68,76,86,98,110,126,144,162,180,200,224,250,280,480}, {0,2,4,6,8,12,16,21,26,31,36,41,46,51,56,61,66,71,77,83,89,95,103,111,121,131,147,163,179,203,240,280,480}, {0,1,2,3,5,7,9,12,15,19,23,27,32,37,43,49,57,66,76,86,98,110,125,140,158,176,194,216,238,264,290,320,480}, {0,1,2,3,5,7,10,13,17,21,25,30,35,41,47,54,62,70,80,90,102,114,130,146,162,180,198,218,240,264,290,320,480}, {0,1,2,4,6,8,11,14,18,22,26,30,36,42,50,58,66,76,88,100,112,128,144,160,182,204,226,256,286,316,352,400,480}, {0,1,2,4,6,8,11,14,18,22,26,30,36,42,50,58,68,78,90,102,116,132,148,166,186,208,234,262,292,324,360,400,480} }.
[0130] FIG. 6 shows the relationship between the subbands indicated by the eight subband division schemes included in the first group of subband division schemes and the initial frequencies of those subbands.
[0131] The frame length of the audio signal is 10 ms and the sampling rate is 88.2 kHz or 96 kHz, or the frame length of the audio signal is 5 ms and the sampling rate is 88.2 kHz or 96 kHz, or the frame length of the audio signal is 10 ms and the sampling rate is 44.1 kHz or 48 kHz. In this case, if the encoding bit rate of the audio signal is not lower than the first bit rate threshold and / or the feature analysis result includes an objective signal flag, the encoder side determines a second group of subband decomposition schemes from the plurality of candidate subband decomposition schemes as the plurality of subband decomposition schemes. The subband decomposition schemes of the second group are as follows: { {0,1,2,3,4,5,6,7,8,10,12,14,16,19,22,26,30,35,40,45,50,57,64,73,82,92,102,112,124,136,148,160,480}, {0,1,2,3,4,5,7,9,11,13,15,18,21,24,28,33,38,44,50,57,64,73,82,93,104,116,128,140,155,170,185,200,480}, {0,1,2,3,4,6,8,10,13,16,20,24,28,33,38,45,52,61,70,79,88,100,112,127,142,160,178,196,217,238,259,280,480}, {0,1,2,4,6,10,14,18,22,26,30,34,42,50,58,66,74,84,96,108,120,136,152,168,192,216,240,272,304,336,376,424,480}, {0,1,2,4,6,10,14,18,26,34,42,50,62,74,86,98,112,128,144,160,176,196,216,236,256,280,304,328,352,384,416,448,480}, {0,80,92,104,112,120,128,136,144,148,152,156,160,164,168,172,176,180,184,188,192,196,200,208,216,224,232,240,248,256,268,280,480}, {0,200,212,224,232,240,248,256,264,268,272,276,280,284,288,292,296,300,304,308,312,316,320,328,336,344,352,360,368,376,388,400,480}, {0,320,332,344,356,364,372,380,384,388,392,396,400,404,408,412,416,420,424,428,432,436,440,444,448,452,456,460,464,468,472,476,480} }.
[0132] FIG. 7 shows the relationship between the subbands represented by the eight subband division schemes included in the second group of subband division schemes and the initial frequencies of those subbands.
[0133] When the frame length of an audio signal is 10 ms and the sampling rate is 88.2 kHz or 96 kHz, the spectrum of each audio frame included in the audio signal includes 960 frequencies. In the process of performing subband decomposition based on the subband decomposition scheme of the second group, the encoder multiplies each subband decomposition value in the subband decomposition scheme of the second group by 2 to obtain subband decomposition values corresponding to the 960 frequencies, and performs subband decomposition based on the subband decomposition values corresponding to the 960 frequencies. When the frame length of an audio signal is 5 ms and the sampling rate is 88.2 kHz or 96 kHz, or when the frame length of an audio signal is 10 ms and the sampling rate is 44.1 kHz or 48 kHz, the spectrum of each audio frame included in the audio signal includes 480 frequencies, and therefore the last subband decomposition value in each subband decomposition scheme of the second group is also 480. Therefore, the encoder directly performs subband decomposition based on the subband decomposition scheme of the second group.
[0134] The frame length of the audio signal is 5 ms, and the sampling rate is 44.1 kHz or 48 kHz. In this case, if the encoding bit rate of the audio signal is lower than the first bit rate threshold and the feature analysis result includes the subjective signal flag, the encoder side determines the subband decomposition schemes of the third group from the plurality of candidate subband decomposition schemes as the plurality of subband decomposition schemes. The subband decomposition schemes of the third group are as follows: { {0,1,2,3,4,5,6,7,8,9,10,12,14,16,19,22,26,30,35,39,44,50,56,63,71,80,89,98,108,119,129,140,240}, {0,1,2,3,4,5,6,7,8,9,11,13,15,17,20,24,28,32,37,42,47,53,59,67,75,83,92,101,110,120,130,140,240}, {0,1,2,3,4,5,6,7,8,9,10,11,12,14,17,20,23,26,30,34,38,43,49,55,63,72,81,90,100,112,125,140,240}, {0,1,2,3,4,6,8,10,13,15,18,20,23,25,28,30,33,35,38,41,44,47,51,55,60,65,73,81,89,101,120,140,240}, {0,1,2,3,4,5,6,7,9,11,13,14,16,18,21,24,28,33,38,43,49,55,62,70,79,88,97,108,119,132,145,160,240}, {0,1,2,3,4,5,6,7,8,10,12,14,17,20,23,27,31,35,40,45,51,57,65,73,81,90,99,109,120,132,145,160,240}, {0,1,2,3,4,5,6,7,9,11,13,15,18,21,25,29,33,38,44,50,56,64,72,80,91,102,113,128,143,158,176,200,240}, {0,1,2,3,4,5,6,7,9,11,13,15,18,21,25,29,34,39,45,51,58,66,74,83,93,104,117,131,146,162,180,200,240} }.
[0135] FIG. 8 shows the relationship between the subbands indicated by the eight subband division schemes included in the third group of subband division schemes and the initial frequencies of those subbands.
[0136] The frame length of the audio signal is 5 ms, and the sampling rate is 44.1 kHz or 48 kHz. In this case, if the encoding bit rate of the audio signal is not lower than the first bit rate threshold and / or the feature analysis result includes an objective signal flag, the encoder side determines a fourth group of subband decomposition schemes from the plurality of candidate subband decomposition schemes as the plurality of subband decomposition schemes. The subband decomposition schemes of the fourth group are as follows: { {0,1,2,3,4,5,6,7,8,9,10,11,12,13,14,15,16,17,18,19,20,21,22,23,24,26,28,30,32,34,37,40,120}, {0,1,2,3,4,5,6,7,8,9,10,11,12,13,14,15,16,18,20,22,24,26,28,30,32,34,36,38,41,44,47,50,120}, {0,1,2,3,4,5,6,7,8,9,10,11,12,14,16,18,20,22,24,26,28,31,34,37,40,44,48,52,56,60,65,70,120}, {0,1,2,3,4,5,6,7,8,9,10,11,12,13,15,17,19,21,24,27,30,34,38,42,48,54,60,68,76,84,94,106,120}, {0,1,2,3,4,5,6,7,8,10,12,14,16,19,22,25,28,32,36,40,44,49,54,59,64,70,76,82,88,96,104,112,120}, {0,20,23,26,28,30,32,34,36,37,38,39,40,41,42,43,44,45,46,47,48,49,50,52,54,56,58,60,62,64,67,70,120}, {0,50,53,56,58,60,62,64,66,67,68,69,70,71,72,73,74,75,76,77,78,79,80,82,84,86,88,90,92,94,97,100,120}, {0,80,83,86,89,91,93,95,96,97,98,99,100,101,102,103,104,105,106,107,108,109,110,111,112,113,114,115,116,117,118,119,120} }.
[0137] FIG. 9 shows the relationship between the subbands indicated by the eight subband division schemes included in the fourth group of subband division schemes and the initial frequencies of those subbands.
[0138] When the frame length of an audio signal is 5 ms and the sampling rate is 44.1 kHz or 48 kHz, the spectrum of each audio frame included in the audio signal includes 240 frequencies. In the process of performing subband decomposition based on the subband decomposition method of the fourth group, the encoder side multiplies each subband decomposition value in the subband decomposition method of the fourth group by 2 to obtain subband decomposition values corresponding to the 240 frequencies, and performs subband decomposition based on the subband decomposition values corresponding to the 240 frequencies.
[0139] Each subband division scheme provided in the embodiments of this application complies with the Bark requirements. The Bark scale indicates that the spectral subband division policy divides the subbands from the perspective of hearing ability based on the auditory characteristics of the human ear.
[0140] Step 502: Determine a total scale value for each candidate subband set based on the spectral values of the audio signal in the subbands included in the candidate subband set, the encoding bit rate of the audio signal, and the subband bandwidth of the subbands included in the candidate subband set.
[0141] In an embodiment of this application, after obtaining multiple candidate subband sets that correspond one-to-one to multiple subband division methods, the encoder side determines the total scale value of each candidate subband set based on the spectral values of the audio signal in the subbands included in the candidate subband set, the encoding bit rate of the audio signal, and the subband bandwidth of the subbands included in the candidate subband set.
[0142] Optionally, the encoder side may determine the total scale factor of each candidate subband set based on the spectral values of the audio signal in the subbands included in the candidate subband set, the coding bitrate of the audio signal, and the subband bandwidths of the subbands included in the candidate subband set, including: for a first candidate subband set among the plurality of candidate subband sets, determining a scale factor of each subband included in the first candidate subband set based on the spectral values of the audio signal in the subbands included in the first candidate subband set. The first candidate subband set is one of the plurality of candidate subband sets. Then, the encoder side determines the total scale factor of the first candidate subband set based on the coding bitrate of the audio signal and the scale factor and subband bandwidth of each subband included in the first candidate subband set. Note that for each candidate subband set among the plurality of candidate subband sets other than the first candidate subband set, the encoder side determines the total scale factor of each of the other candidate subband sets in the same manner as determining the total scale factor of the first candidate subband set.
[0143] There are several implementations in which the encoder determines the scale factor for each subband. In one implementation, determining the scale factor for each subband included in the first candidate subband set based on the spectral values of the audio signal in the subbands included in the first candidate subband set includes, for the first subband included in the first candidate subband set, the encoder obtains the maximum value among the absolute values of all the spectral values of the audio signal in the first subband and determines the scale factor for the first subband based on the maximum value. The first subband is one of the subbands in the first candidate subband set. Note that for each subband other than the first subband in the first candidate subband set, the encoder determines the scale factor for each of the other subbands in the first candidate subband set in the same manner as determining the scale factor for the first subband.
[0144] For example, the encoder side determines the scale factor for each subband in the first candidate subband set according to equation (5).
[0145]
number
[0146] In equation (5), X(k) represents the kth spectral value of the audio signal, b represents the sequential number of the subband, I(b) represents the initial frequency of subband b, B represents the cutoff subband corresponding to the first candidate subband set, i.e., the total number of subbands included in the first candidate subband set, abs() represents obtaining the absolute value, max() represents obtaining the maximum value, ceil() represents rounding up, and E() represents the scale factor of the subband.
[0147] Below, we describe an implementation in which the encoder determines the total scale value of the first candidate subband set based on the encoding bit rate of the audio signal and the scale factor and subband bandwidth of each subband included in the first candidate subband set.
[0148] Optionally, the encoding bitrate of the audio signal is not lower than a first bitrate threshold, and / or the energy concentration of the audio signal is greater than a concentration threshold. The encoder side determines an energy smoothing reference value based on the encoding bitrate of the audio signal and a second bitrate threshold. The encoder side determines a total energy value of each subband included in the first candidate subband set based on the energy smoothing reference value and the scale factor and subband bandwidth of each subband included in the first candidate subband set. The encoder side sums the total energy values of the subbands included in the first candidate subband set to obtain a total scale value for the first candidate subband set. If the energy concentration of the audio signal is greater than the concentration threshold, it indicates that the audio signal is an objective signal. It should be understood that when the encoding bitrate is high and / or the audio signal is an objective signal, the encoder side determines the total scale value based on the total energy value of each subband.
[0149] In the embodiment of this application, there are several implementations in which the encoder side determines the energy smoothing reference value. One of these implementations is described here. In this implementation, the encoder side determines the energy smoothing reference value according to Equation (6).
[0150]
number
[0151] In equation (6), E floordenotes the energy smoothing reference value, bpsPerChn denotes the coding bit rate of the audio signal, where the coding bit rate of the audio signal is the coding bit rate of a single channel, 200 denotes that the second bit rate threshold is 200 kbps, min() denotes obtaining the minimum value, and int() denotes rounding down. Note that the second bit rate threshold may be other values instead.
[0152] There are several implementations in which the encoder determines the total energy value of each subband included in the first candidate subband set based on the energy smoothing reference value and the scale factor and subband bandwidth of each subband included in the first candidate subband set. One of these implementations is described here. In this implementation, for a first subband included in the first candidate subband set, the encoder determines the larger of the scale factor of the first subband and the energy smoothing reference value as the reference scale value of the first subband. The encoder determines the total energy value of the first subband as the product of the reference scale value of the first subband and the subband bandwidth of the first subband. The first subband is one of the subbands in the first candidate subband set. Note that for each subband in the first candidate subband set other than the first subband, the encoder determines the total energy value of each of the other subbands in the same manner as determining the total energy value of the first subband.
[0153] The encoder side determines the total energy value of each subband included in the first candidate subband set and the total scale factor of the first candidate subband set according to equation (7).
[0154]
number
[0155] In equation (7), b denotes the sequential number of the subband, B denotes the cutoff subband corresponding to the first candidate subband set, bandWidth() denotes the subband bandwidth, E(b) denotes the scale factor of subband b, and E floor indicates the energy smoothing reference value, max() indicates obtaining the maximum value, and max[E(b),E floor ]*bandWidth(b) denotes the total energy value of subband b, and E total denotes the total scale value of the first subband set.
[0156] The above describes an implementation process in which the encoder determines the sum scale value of the first candidate subband set when the encoding bitrate of the audio signal is not lower than the first bitrate threshold and / or the energy concentration of the audio signal is greater than the concentration threshold. The following describes an implementation process in which the encoder determines the sum scale value of the first candidate subband set when the encoding bitrate of the audio signal is lower than the first bitrate threshold and the energy concentration of the audio signal is not greater than the concentration threshold.
[0157] When the encoding bitrate of the audio signal is lower than a first bitrate threshold and the energy concentration of the audio signal is not greater than the concentration threshold, the encoder determines a total scale factor for the first candidate subband set based on the encoding bitrate of the audio signal and the scale factor and subband bandwidth of each subband included in the first candidate subband set. The encoder determines an energy smoothing reference value based on the encoding bitrate of the audio signal and a second bitrate threshold. The encoder determines scale difference values for the subbands included in the first candidate subband set based on the energy smoothing reference value and the scale factor of each subband included in the first candidate subband set. The scale difference values indicate differences between the scale factor of a corresponding subband and the scale factor of an adjacent subband of the corresponding subband. The encoder determines a total scale factor for the first candidate subband set based on the scale difference values of the subbands included in the first candidate subband set and the subband bandwidth of each subband. When the energy concentration of the audio signal is not greater than the concentration threshold, it indicates that the audio signal is a subjective signal. It should be understood that when the encoding bit rate is low and the audio signal is a subjective signal, the encoder side determines the total scale value based on the difference between each subband and its neighboring subbands.
[0158] For the implementation in which the encoder determines the energy smoothing reference value based on the encoding bit rate of the audio signal and the second bit rate threshold, please refer to the related description above, and the details will not be described again here.
[0159] There are several implementations in which the encoder determines the scale difference values of the subbands included in the first candidate subband set based on the energy smoothed reference value and the scale factors of each subband included in the first candidate subband set. One of these implementations is described here. In this implementation, for a first subband included in the first candidate subband set, the encoder determines a first smoothed value, a second smoothed value, and a third smoothed value for the first subband based on the energy smoothed reference value, the scale factor of the first subband, and the scale factors of the first subband's adjacent subbands. The encoder determines the scale difference value of the first subband based on the first smoothed value, the second smoothed value, and the third smoothed value for the first subband. The first subband is one of the subbands in the first candidate subband set.
[0160] Optionally, if the first subband is the first subband in the first candidate subband set, the encoder side determines the larger of the scale factor of the first subband and the energy smoothing reference value as the first smoothed value of the first subband, and if the first subband is not the first subband in the first candidate subband set, the encoder side determines the larger of the scale factor of the previous subband adjacent to the first subband and the energy smoothing reference value as the first smoothed value of the first subband.
[0161] The encoder determines the larger value of the scale factor of the first subband and the energy smoothing reference value as the second smoothed value of the first subband.
[0162] If the first subband is the last subband in the first candidate subband set, the encoder determines the larger of the scale factor and the energy smoothing reference value of the first subband as the third smoothed value of the first subband; if the first subband is not the last subband in the first candidate subband set, the encoder determines the larger of the scale factor and the energy smoothing reference value of the next subband adjacent to the first subband as the third smoothed value of the first subband.
[0163] In other words, the encoder side determines the first smoothed value, the second smoothed value, and the third smoothed value for each subband according to equations (8), (9), and (10).
[0164]
number
[0165] In Equation (8), Equation (9), and Equation (10), left(), center(), and right() denote the first smoothed value, the second smoothed value, and the third smoothed value, respectively. In this embodiment of this application, the first smoothed value, the second smoothed value, and the third smoothed value may also be referred to as the left smoothed value, the middle smoothed value, and the right smoothed value, respectively.
[0166] Optionally, after the encoder side determines the first smoothed value, the second smoothed value, and the third smoothed value for the first subband, the implementation process for determining the scaled difference value for the first subband includes the encoder side determining, for the first subband included in the first candidate subband set, a first difference value and a second difference value for the first subband. The first difference value is the absolute value of the difference value between the first smoothed value and the second smoothed value for the first subband, and the second difference value is the absolute value of the difference value between the second smoothed value and the third smoothed value for the first subband. The encoder side determines the scaled difference value for the first subband based on the first difference value and the second difference value for the first subband. The first subband is any subband in the first candidate subband set.
[0167] For example, the encoder side determines the scale difference value of the first subband according to equation (11).
[0168]
number
[0169] After the encoder determines the scale difference value of each subband included in the first candidate subband set, the encoder determines a total scale value for the first candidate subband set based on the scale difference value of each subband and the subband bandwidth. The encoder determines a smooth weighting factor for each subband included in the first candidate subband set based on the number of subbands included in the first candidate subband set and the subband bandwidth of each subband. The encoder sums the smooth weighting factors of the subbands included in the first candidate subband set to obtain a total smooth weighting factor for the first candidate subband set. The encoder multiplies the scale difference value of the subbands included in the first candidate subband set by the smooth weighting factor to obtain a weighted scale difference value for the subbands included in the first candidate subband set. The encoder sums the weighted scale difference values of the subbands included in the first candidate subband set to obtain a total scale value for the first candidate subband set. The encoder side divides the total scale value by the total smoothing weighting factor of the first candidate subband set to obtain a total scale value for the first candidate subband set.
[0170] Alternatively, the step of the encoder determining the smoothed weighting factor for each subband and the total smoothed weighting factor for the first candidate subband set may be performed before the scaled difference value for each subband is determined. The order of the steps performed by the encoder is not limited in the embodiments of this application.
[0171] Optionally, the encoder side determines a scaling factor for the subband division based on the number of subbands included in the first candidate subband set, and determines a smoothing weighting factor for each subband included in the first candidate subband set based on the scaling factor for the subband division and the subband bandwidth of each subband included in the first candidate subband set.
[0172] For example, the encoder determines the scaling coefficient coef for the subband division according to equation (12).
[0173]
number
[0174] The encoder determines the smoothing weighting factor frac for each subband according to equation (13).
[0175]
number
[0176] The encoder determines the sum smoothing weighting factor sum for the first candidate subband set according to equation (14).
[0177]
number
[0178] The encoder calculates the total scale value E′ of the first candidate subband set according to equation (15). total Determine.
[0179]
number
[0180] E diff (b)*frac(b) denotes the weighted scale differential value for subband b.
[0181] The encoder calculates the total scale value E for the first candidate subband set according to equation (16). total Determine.
[0182]
number
[0183] When the audio signal is a mono-channel signal, the encoder side can calculate the total scale value of each candidate subband set according to the above formula. When the audio signal is a dual-channel signal, the spectrum of the audio signal includes a left channel spectrum and a right channel spectrum, and the encoder side calculates the total scale value of each candidate subband set based on the left channel spectrum and the right channel spectrum. For example, the encoder side adds the total scale value calculated based on the left channel spectrum and the total scale value calculated based on the right channel spectrum to obtain the total scale value of the candidate subband set. In one implementation, another sum Σ is added to the above formula related to the sum Σ, and the added sum Σ indicates that the related data of the left channel and the related data of the right channel are added.
[0184] Step 503: Select one candidate subband set from the plurality of candidate subband sets as a target subband set based on the total scale value of each candidate subband set, and each subband included in the target subband set has a scale factor used to shape the spectral envelope of the audio signal.
[0185] In an embodiment of this application, the encoder side determines the candidate subband set with the smallest total scale value among the plurality of candidate subband sets as the target subband set. In some other embodiments, the encoder side may instead determine the candidate subband set with the second smallest total scale value among the plurality of candidate subband sets as the target subband set. The second smallest total scale value is the smallest total scale value among the scale values other than the smallest total scale value.
[0186] It can be seen from the above that the encoder selects the most suitable subband decomposition method from among multiple subband decomposition methods based on the characteristics of the audio signal, which means that the subband decomposition method has signal adaptive properties, which helps to improve the coding effect and compression efficiency.
[0187] To further improve coding effect and compression efficiency, when the audio signal is a dual-channel signal, the encoder side can further determine whether coding performance can be improved by performing mid / side stereo transform coding (MS transform) on the spectrum of the audio signal based on the determined target subband set. If it is determined that the MS transform helps improve coding performance, the encoder side performs subsequent encoding procedures based on the spectrum obtained through the MS transform. If it is determined that the MS transform does not help improve coding performance, the encoder side performs subsequent encoding procedures based on the original spectrum of the audio signal, as will be described later.
[0188] In an embodiment of this application, when the audio signal is a dual-channel signal, the encoder side determines a first total scale value based on the scale factor and subband bandwidth of each subband included in a target subband set. The encoder side performs MS transform on the spectrum of the dual-channel signal to obtain a transformed spectrum of the dual-channel signal. The encoder side determines a transformed scale factor for each subband in the target subband set based on the transformed spectral value of the dual-channel signal in each subband included in the target subband set. The encoder side determines a second total scale value based on the transformed scale factor and subband bandwidth of each subband included in the target subband set. If the first total scale value is not greater than the second total scale value, the encoder side determines the dual-channel signal (the dual-channel signal before MS transform) as the signal to be encoded.
[0189] It should be understood that the first total scale value is the total scale value before MS conversion, and the second total scale value is the total scale value obtained through MS conversion. A higher total scale value indicates a lower coding performance gain. If the first total scale value is not greater than the second total scale value, it indicates that MS conversion does not help improve coding performance. Therefore, the encoder side determines the dual-channel signal before MS conversion as the signal to be coded.
[0190] Optionally, the spectrum of the dual-channel signal before MS conversion is referred to as the LR spectrum, and the spectrum of the dual-channel signal obtained through MS conversion is referred to as the MS spectrum, where LR denotes left and right channels.
[0191] When the audio signal is a dual-channel signal, the scale factors include a left channel scale factor and a right channel scale factor. Optionally, the encoder side may determine the first total scale value based on the scale factors and subband bandwidths of each subband included in the target subband set by: determining a product of the left channel scale factor of each subband included in the target subband set and the subband bandwidth of the corresponding subband as the left channel energy value of the corresponding subband; and determining a product of the right channel scale factor of each subband included in the target subband set and the subband bandwidth of the corresponding subband as the right channel energy value of the corresponding subband. The encoder side may sum the left channel energy values and the right channel energy values of all subbands included in the target subband set to obtain the first total scale value.
[0192] For example, the encoder side determines the first total scale value according to equation (17).
[0193]
number
[0194] In equation (17), totalScale1 denotes the first total scale value, and ch denotes the sequential number of the left and right channels. When ch=0, E(b) denotes the left channel scale factor. When ch=1, E(b) denotes the right channel scale factor.
[0195] The encoder side performs MS conversion according to equation (18).
[0196]
number
[0197] In equation (18), L and R respectively represent the spectral values of the left channel and the right channel before transformation. M and S respectively represent the transformed spectral values of the left channel and the transformed spectral values of the right channel. The encoder processes the spectral values of corresponding frequencies in the left channel spectrum and the right channel spectrum according to equation (18) to obtain the spectral values of corresponding frequencies in the transformed spectrum of the left channel and the transformed spectrum of the right channel. The transformed spectral values of the left channel and the transformed spectral values of the right channel are the spectral values of the two channels included in the transformed dual-channel signal. The transformed left channel and the transformed right channel may also be referred to as the transformed M channel and the transformed S channel.
[0198] The encoder side determines the transformed scale factor for each subband according to equation (19), which is similar to equation (5).
[0199]
number
[0200] In equation (19), X_MS(k) denotes the transformed k-th spectral value, and E_MS(b) denotes the scale factor of subband b on the M channel or S channel, i.e., the scale factor of subband b on the transformed channel. Note that the encoder calculates the scale factor of the M channel based on the spectral value of the M channel according to equation (19), and calculates the scale factor of the S channel based on the spectral value of the S channel according to equation (19).
[0201] The encoder side determines the second total scale value according to equation (20).
[0202]
number
[0203] In equation (20), totalScale2 denotes the second total scale value, ch denotes the sequential number of the M channel and the S channel, when ch=0, E_MS(b) denotes the scale factor of the subband on the transformed left channel, when ch=1, E_MS(b) denotes the scale factor of the subband on the transformed right channel, that is, the scale factor of subband b on the M channel or the S channel.
[0204] Optionally, if the first total scale value is greater than the second total scale value, and the encoding bit rate of the audio signal is not lower than the first bit rate threshold and / or the energy concentration of the audio signal is greater than the concentration threshold, the encoder side determines the transformed dual-channel signal as the signal to be encoded. It should be understood that the first total scale value being greater than the second total scale value indicates that MS transformation can help improve coding performance. Therefore, the encoder side determines the dual-channel signal obtained through MS transformation as the signal to be encoded.
[0205] It can be seen from the above that when the audio signal is a dual-channel signal, the scale factors include a left channel scale factor and a right channel scale factor. Optionally, when the first total scale value is greater than the second total scale value, the encoding bit rate of the audio signal is lower than a first bit rate threshold, and the energy concentration of the audio signal is not greater than a concentration threshold, the encoder side determines a left channel scale factor of each subband included in the target subband set based on the left channel scale factor and the right channel scale factor of each subband included in the target subband set. channel Scale Factor and Right channelThe encoder determines a difference value between the start and end frequency of each subband included in the target subband set based on the initial frequency and cutoff frequency of each subband included in the target subband set. If there are any subbands in the target subband set that are greater than the difference threshold, the encoder determines a difference value between the start and end frequency of each subband included in the target subband set based on the initial frequency and cutoff frequency of each subband included in the target subband set. channel Scale Factor and Right channel If there is at least one subband that has a difference value between the scale factor and the start-end frequency difference value that is within the first range, the encoder side determines the pre-converted dual-channel signal as the signal to be encoded.
[0206] In other words, when the coding bit rate is low and the audio signal is an objective signal, the encoder side determines whether MS conversion improves coding performance based on the difference value between the left channel scale factor and the right channel scale factor and the start-end frequency difference value of the subband.
[0207] Optionally, the encoder traverses all subbands in the target subband set. channel Scale Factor and Right channel When there is a subband having a difference value between the scale factor and the start-end frequency difference value that is within the first range, the encoder side determines the pre-transformed dual-channel signal as the signal to be encoded.
[0208] For example, the encoder side calculates the left and right subbands according to equation (21). channel Scale Factor and Right channel The difference value between the scale factor is determined.
[0209]
number
[0210] In equation (21), E_L() denotes the left channel scale factor, and E_R() is right Indicates the channel scale factor, diffSFflag() is left channel Scale Factor and Right channel Indicates the difference between the scale factor.
[0211] The encoder calculates the left half of each subband according to equation (21). channel Scale Factor and Right channel When determining the difference value between the scale factors, the difference threshold is 3.
[0212] The encoder determines the subband center frequency of each subband according to equation (22).
[0213]
number
[0214] In equation (22), freq() denotes the start-end frequency difference value, bandstart() and bandend() denote the initial frequency and cutoff frequency, respectively, SamplingRate denotes the sampling rate in Hz, and FrameLength denotes the number of sampling points in each frame.
[0215] Optionally, when the encoder side determines the subband center frequency of each subband according to equation (22), the first range is (3500, 12000].
[0216] Simply put, when the encoder uses equations (21) and (22), it traverses all subbands in the target subband set. channel Scale Factor and Right channel If there exists a subband having a difference value diffSFflag between the scale factor and the subband center frequency freq within the range of (3500, 12000), the encoder determines the pre-conversion dual-channel signal as the signal to be coded.
[0217] If the at least one subband does not exist in the target subband set, the encoder side determines the transformed dual-channel signal as the signal to be coded. The at least one subband is defined as a subband having a left-hand difference greater than a difference threshold. channel Scale Factor and Right channel The subbands have a difference value between the scale factor and the subband center frequency that is within the first range.
[0218] In the following, with reference to FIG. 10, the implementation process in which the encoder side determines whether the converted dual-channel signal should be used as the signal to be encoded will be described again.
[0219] See FIG. 10. The encoder calculates a first total scale value based on the selected target subband set and the left / right (LR) channel scale factor (SF) of each subband in the target subband set. The first total scale value is the sum of the products of the LR channel scale factors of all subbands in the target subband set and the corresponding subband bandwidths. The encoder transforms the LR channel spectrum into an MS channel spectrum and calculates a second total scale value. The second total scale value is the sum of the products of the MS channel scale factors of all subbands in the target subband set and the corresponding subband bandwidths. If the first total scale value is not greater than the second total scale value, the encoder determines the pre-transformed dual-channel signal as the signal to be coded and sets MSFlag=0 to indicate that the execution of subsequent operations will not be based on the spectral values obtained through MS transformation.
[0220] If the first total scale value is greater than the second total scale value, the encoder determines whether the audio signal (i.e., the dual-channel signal before conversion) satisfies a first condition. The first condition is whether the encoding bit rate of the audio signal is lower than a first bit rate threshold and the energy concentration of the audio signal is greater than a first bit rate threshold. Inside The first condition is that the audio signal is smaller than the convergence threshold. If the audio signal satisfies the first condition, the encoder sets the high bitrate flag to 0. If the audio signal does not satisfy the first condition, the encoder sets the high bitrate flag to 1.
[0221] If the high bitrate flag is equal to 1, the encoder determines the converted dual-channel signal as the signal to be coded and sets MSFlag=1 to indicate that subsequent operations will be based on the spectral values obtained through MS conversion. If the high bitrate flag is equal to 0, the encoder calculates the LR channel SF difference value and the subband center frequency of each subband through traversal. If the subband obtained through traversal satisfies a second condition, the encoder sets the SF difference flag to 1. The second condition is that the LR channel SF difference value of the corresponding subband is smaller than the difference threshold and the subband center frequency is within a first range. If the subband obtained through traversal does not satisfy the second condition, the encoder sets the SF difference flag to 0.
[0222] If the SF difference flag is equal to 1, the encoder side determines the pre-conversion dual-channel signal as the signal to be coded and sets MSFlag = 0. If the SF difference flag is equal to 0, the encoder side determines the converted dual-channel signal as the signal to be coded and sets MSFlag = 1.
[0223] It should be noted that in addition to the above implementation in which the encoder side determines whether to use the converted dual-channel signal as the signal to be encoded, the encoder side may instead make the determination in other ways, in other words, the above implementation is not intended to limit the embodiments of this application.
[0224] In summary, in an embodiment of this application, an optimal subband decomposition scheme is selected from multiple subband decomposition schemes based on the characteristics of the audio signal. In other words, the subband decomposition scheme has signal adaptive properties and can adapt to the coding bit rate of the audio signal to improve interference resistance. Specifically, the audio signal is separately divided based on multiple subband decomposition schemes, and a total scale factor corresponding to each subband decomposition scheme is determined based on the spectral values of the audio signal in the subbands obtained through the division, the bandwidth of each subband, and the coding bit rate of the audio signal. An optimal target subband decomposition scheme is selected based on the total scale factor to obtain an optimal subband set. Then, spectral envelope shaping is performed based on the scale factor of each subband in the optimal subband set, thereby improving coding efficiency and compression efficiency.
[0225] 11 is a diagram of a configuration of an audio signal processing device 1100 according to an embodiment of the present application. The processing device 1100 may be implemented as part of or as a whole of an electronic device by using software, hardware, or a combination thereof. The electronic device may be any of the devices shown in FIG. 1. See FIG. 11. The device includes a subband decomposition module 1101, a first determination module 1102, and a selection module 1103.
[0226] The subband decomposition module 1101 is configured to perform subband decomposition on the audio signal separately based on a plurality of subband decomposition schemes and cutoff subbands corresponding to the plurality of subband decomposition schemes to obtain a plurality of candidate subband sets, where the plurality of candidate subband sets correspond one-to-one to the plurality of subband decomposition schemes, and each candidate subband set includes a plurality of subbands.
[0227] The first determination module 1102 is configured to determine a total scale value for each candidate subband set based on the spectral values of the audio signal in the subbands included in the candidate subband set, the encoding bit rate of the audio signal, and the subband bandwidth of the subbands included in the candidate subband set.
[0228] The selection module 1103 is configured to select one candidate subband set from the plurality of candidate subband sets as a target subband set based on the total scale factor of each candidate subband set, where each subband included in the target subband set has a scale factor used to shape the spectral envelope of the audio signal.
[0229] Optionally, the selection module 1103 is The candidate subband set having the smallest total scale value among the plurality of candidate subband sets is determined as the target subband set.
[0230] Optionally, the first determination module 1102: a first determining sub-module configured to, for a first candidate subband set among a plurality of candidate subband sets, determine a scale factor for each subband included in the first candidate subband set based on spectral values of an audio signal in the subbands included in the first candidate subband set, the first candidate subband set being any one of the plurality of candidate subband sets; and a second determination submodule configured to determine a total scale value of the first candidate subband set based on the encoding bit rate of the audio signal and the scale factor and subband bandwidth of each subband included in the first candidate subband set.
[0231] Optionally, the second decision submodule: For a first subband included in a first candidate subband set, obtain a maximum value among absolute values of all spectral values of the audio signal in the first subband, the first subband being any subband in the first candidate subband set; and determining a scale factor for the first subband based on the maximum value.
[0232] Optionally, the coding bit rate of the audio signal is not lower than a first bit rate threshold and / or the energy concentration of the audio signal is greater than a concentration threshold.
[0233] The second decision submodule is determining an energy smoothing reference value based on an encoding bit rate of the audio signal and a second bit rate threshold; determining a total energy value for each subband in the first candidate subband set based on the energy smoothing criterion and the scale factor and subband bandwidth of each subband in the first candidate subband set; The method is configured to sum the total energy values of the subbands included in the first candidate subband set to obtain a total scale value for the first candidate subband set.
[0234] Optionally, the second decision submodule: For a first subband included in a first candidate subband set, determine a larger value of the scale factor of the first subband and the energy smoothing reference value as a reference scale value of the first subband, where the first subband is any subband in the first candidate subband set; The audio signal processing unit is configured to determine a product of the reference scale value of the first subband and the subband bandwidth of the first subband as a total energy value of the first subband.
[0235] Optionally, the coding bit rate of the audio signal is lower than a first bit rate threshold and the energy concentration of the audio signal is not greater than a concentration threshold.
[0236] The second decision submodule is determining an energy smoothing reference value based on an encoding bit rate of the audio signal and a second bit rate threshold; determining scale difference values for the subbands included in the first candidate subband set based on the energy smoothing reference value and the scale factors of each subband included in the first candidate subband set, the scale difference values indicating differences between the scale factors of the corresponding subbands and the scale factors of the adjacent subbands of the corresponding subbands; The first candidate subband set is configured to determine a total scale value for the first candidate subband set based on the scale difference values of the subbands included in the first candidate subband set and the subband bandwidth of each subband.
[0237] Optionally, the second decision submodule: For a first subband included in a first candidate subband set, determine a first smoothed value, a second smoothed value, and a third smoothed value for the first subband based on the energy smoothing reference value, a scale factor of the first subband, and scale factors of adjacent subbands of the first subband, where the first subband is any subband in the first candidate subband set; and determining a scaled difference value for the first subband based on the first smoothed value, the second smoothed value, and the third smoothed value for the first subband.
[0238] Optionally, the second decision submodule: If the first subband is the first subband in the first candidate subband set, determine the larger value of the scale factor of the first subband and the energy smoothing reference value as the first smoothed value of the first subband; if the first subband is not the first subband in the first candidate subband set, determine the larger value of the scale factor of a previous subband adjacent to the first subband and the energy smoothing reference value as the first smoothed value of the first subband; determining a second smoothed value for the first subband as a larger value of the scale factor for the first subband and the energy smoothed reference value; If the first subband is the last subband in the first candidate subband set, the larger of the scale factor of the first subband and the energy smoothed reference value is determined as the third smoothed value of the first subband; if the first subband is not the last subband in the first candidate subband set, the larger of the scale factor of the next subband adjacent to the first subband and the energy smoothed reference value is determined as the third smoothed value of the first subband.
[0239] Optionally, the second decision submodule: determining a first difference value and a second difference value for a first subband included in a first candidate subband set, the first difference value being an absolute value of a difference value between the first smoothed value and the second smoothed value for the first subband, and the second difference value being an absolute value of a difference value between the second smoothed value and the third smoothed value for the first subband, the first subband being any subband in the first candidate subband set; and determining a scaled difference value for the first subband based on the first difference value and the second difference value for the first subband.
[0240] Optionally, the second decision submodule: determining a smoothing weighting factor for each subband included in the first candidate subband set based on the number of subbands included in the first candidate subband set and the subband bandwidth of each subband; summing the smoothed weighting factors of the subbands included in the first candidate subband set to obtain a total smoothed weighting factor for the first candidate subband set; multiplying the scaled difference values of the subbands included in the first candidate subband set by the smoothing weighting factors to obtain weighted scaled difference values of the subbands included in the first candidate subband set; summing the weighted scale difference values of the subbands included in the first candidate subband set to obtain a total scale value for the first candidate subband set; and dividing the total scale value by the total smoothed weighting factor of the first candidate subband set to obtain a total scale value of the first candidate subband set.
[0241] Optionally, the device 1100 further comprises: a bandwidth detection module configured to perform bandwidth detection on a spectrum of the audio signal to obtain a cut-off frequency of the audio signal when an encoding bit rate of the audio signal is lower than a first bit rate threshold; and a second determining module configured to determine cutoff subbands corresponding to the plurality of subband division schemes, respectively, based on the cutoff frequencies.
[0242] Optionally, the device 1100 further comprises: and a third determination module configured to determine the last subband indicated by each of the plurality of subband decomposition schemes as the cutoff subband corresponding to each subband decomposition scheme when the encoding bitrate of the audio signal is not lower than the first bitrate threshold.
[0243] Optionally, the device 1100 further comprises: a feature analysis module configured to perform feature analysis on the spectrum of the audio signal to obtain a feature analysis result; and a fourth determination module configured to determine a plurality of subband decomposition schemes from a plurality of candidate subband decomposition schemes based on the feature analysis result and the coding bit rate of the audio signal.
[0244] Optionally, the feature analysis result includes a subjective signal flag or an objective signal flag, the subjective signal flag indicating that the energy concentration of the audio signal is not greater than a concentration threshold, and the objective signal flag indicating that the energy concentration of the audio signal is greater than a concentration threshold.
[0245] Optionally, the audio signal has a frame length of 10 milliseconds and a sampling rate of 88.2 kilohertz or 96 kilohertz, or a frame length of 5 milliseconds and a sampling rate of 88.2 kilohertz or 96 kilohertz, or a frame length of 10 milliseconds and a sampling rate of 44.1 kilohertz or 48 kilohertz.
[0246] The fourth decision module is and a third determination submodule configured to determine a first group of subband decomposition schemes from among a plurality of candidate subband decomposition schemes as a plurality of subband decomposition schemes when the encoding bit rate of the audio signal is lower than a first bit rate threshold and the feature analysis result includes a subjective signal flag.
[0247] The subband division method for the first group is as follows: { {0,1,2,3,4,6,8,10,13,16,20,24,28,33,38,45,52,61,70,79,88,100,112,127,142,160,178,196,217,238,259,280,480}, {0,1,2,3,5,7,9,12,15,18,22,26,30,35,41,48,56,65,74,84,94,106,118,134,150,166,184,202,220,240,260,280,480}, {0,1,2,3,4,5,7,9,11,14,17,21,25,29,34,40,46,52,60,68,76,86,98,110,126,144,162,180,200,224,250,280,480}, {0,2,4,6,8,12,16,21,26,31,36,41,46,51,56,61,66,71,77,83,89,95,103,111,121,131,147,163,179,203,240,280,480}, {0,1,2,3,5,7,9,12,15,19,23,27,32,37,43,49,57,66,76,86,98,110,125,140,158,176,194,216,238,264,290,320,480}, {0,1,2,3,5,7,10,13,17,21,25,30,35,41,47,54,62,70,80,90,102,114,130,146,162,180,198,218,240,264,290,320,480}, {0,1,2,4,6,8,11,14,18,22,26,30,36,42,50,58,66,76,88,100,112,128,144,160,182,204,226,256,286,316,352,400,480}, {0,1,2,4,6,8,11,14,18,22,26,30,36,42,50,58,68,78,90,102,116,132,148,166,186,208,234,262,292,324,360,400,480} }.
[0248] Optionally, the audio signal has a frame length of 10 milliseconds and a sampling rate of 88.2 kilohertz or 96 kilohertz, or a frame length of 5 milliseconds and a sampling rate of 88.2 kilohertz or 96 kilohertz, or a frame length of 10 milliseconds and a sampling rate of 44.1 kilohertz or 48 kilohertz.
[0249] The fourth decision module is and a fourth determination sub-module configured to determine a second group of subband decomposition schemes from the plurality of candidate subband decomposition schemes as the plurality of subband decomposition schemes when the encoding bit rate of the audio signal is not lower than the first bit rate threshold and / or the feature analysis result includes an objective signal flag.
[0250] The subband division scheme for the second group is as follows: { {0,1,2,3,4,5,6,7,8,10,12,14,16,19,22,26,30,35,40,45,50,57,64,73,82,92,102,112,124,136,148,160,480}, {0,1,2,3,4,5,7,9,11,13,15,18,21,24,28,33,38,44,50,57,64,73,82,93,104,116,128,140,155,170,185,200,480}, {0,1,2,3,4,6,8,10,13,16,20,24,28,33,38,45,52,61,70,79,88,100,112,127,142,160,178,196,217,238,259,280,480}, {0,1,2,4,6,10,14,18,22,26,30,34,42,50,58,66,74,84,96,108,120,136,152,168,192,216,240,272,304,336,376,424,480}, {0,1,2,4,6,10,14,18,26,34,42,50,62,74,86,98,112,128,144,160,176,196,216,236,256,280,304,328,352,384,416,448,480}, {0,80,92,104,112,120,128,136,144,148,152,156,160,164,168,172,176,180,184,188,192,196,200,208,216,224,232,240,248,256,268,280,480}, {0,200,212,224,232,240,248,256,264,268,272,276,280,284,288,292,296,300,304,308,312,316,320,328,336,344,352,360,368,376,388,400,480}, {0,320,332,344,356,364,372,380,384,388,392,396,400,404,408,412,416,420,424,428,432,436,440,444,448,452,456,460,464,468,472,476,480} }.
[0251] Optionally, the frame length of the audio signal is 5 milliseconds and the sampling rate is 44.1 kHz or 48 kHz.
[0252] The fourth decision module is and a fifth determination submodule configured to determine a third group of subband decomposition schemes from the plurality of candidate subband decomposition schemes as the plurality of subband decomposition schemes when the encoding bit rate of the audio signal is lower than a first bit rate threshold and the feature analysis result includes a subjective signal flag.
[0253] The subband division method of the third group is as follows: { {0,1,2,3,4,5,6,7,8,9,10,12,14,16,19,22,26,30,35,39,44,50,56,63,71,80,89,98,108,119,129,140,240}, {0,1,2,3,4,5,6,7,8,9,11,13,15,17,20,24,28,32,37,42,47,53,59,67,75,83,92,101,110,120,130,140,240}, {0,1,2,3,4,5,6,7,8,9,10,11,12,14,17,20,23,26,30,34,38,43,49,55,63,72,81,90,100,112,125,140,240}, {0,1,2,3,4,6,8,10,13,15,18,20,23,25,28,30,33,35,38,41,44,47,51,55,60,65,73,81,89,101,120,140,240}, {0,1,2,3,4,5,6,7,9,11,13,14,16,18,21,24,28,33,38,43,49,55,62,70,79,88,97,108,119,132,145,160,240}, {0,1,2,3,4,5,6,7,8,10,12,14,17,20,23,27,31,35,40,45,51,57,65,73,81,90,99,109,120,132,145,160,240}, {0,1,2,3,4,5,6,7,9,11,13,15,18,21,25,29,33,38,44,50,56,64,72,80,91,102,113,128,143,158,176,200,240}, {0,1,2,3,4,5,6,7,9,11,13,15,18,21,25,29,34,39,45,51,58,66,74,83,93,104,117,131,146,162,180,200,240} }.
[0254] Optionally, the frame length of the audio signal is 5 milliseconds and the sampling rate is 44.1 kHz or 48 kHz.
[0255] The fourth decision module is and a sixth determination sub-module configured to determine a fourth group of subband decomposition schemes from among the plurality of candidate subband decomposition schemes as the plurality of subband decomposition schemes when the encoding bit rate of the audio signal is not lower than the first bit rate threshold and / or the feature analysis result includes an objective signal flag.
[0256] The subband division method for the fourth group is as follows: { {0,1,2,3,4,5,6,7,8,9,10,11,12,13,14,15,16,17,18,19,20,21,22,23,24,26,28,30,32,34,37,40,120}, {0,1,2,3,4,5,6,7,8,9,10,11,12,13,14,15,16,18,20,22,24,26,28,30,32,34,36,38,41,44,47,50,120}, {0,1,2,3,4,5,6,7,8,9,10,11,12,14,16,18,20,22,24,26,28,31,34,37,40,44,48,52,56,60,65,70,120}, {0,1,2,3,4,5,6,7,8,9,10,11,12,13,15,17,19,21,24,27,30,34,38,42,48,54,60,68,76,84,94,106,120}, {0,1,2,3,4,5,6,7,8,10,12,14,16,19,22,25,28,32,36,40,44,49,54,59,64,70,76,82,88,96,104,112,120}, {0,20,23,26,28,30,32,34,36,37,38,39,40,41,42,43,44,45,46,47,48,49,50,52,54,56,58,60,62,64,67,70,120}, {0,50,53,56,58,60,62,64,66,67,68,69,70,71,72,73,74,75,76,77,78,79,80,82,84,86,88,90,92,94,97,100,120}, {0,80,83,86,89,91,93,95,96,97,98,99,100,101,102,103,104,105,106,107,108,109,110,111,112,113,114,115,116,117,118,119,120} }.
[0257] Optionally, the audio signal is a dual channel signal.
[0258] The device 1100 further comprises: a fifth determining module configured to determine a first total scale value based on the scale factor and subband bandwidth of each subband included in the target subband set; a transform module configured to perform mid / side stereo transform coding on the spectrum of the dual-channel signal to obtain a transformed spectrum of the dual-channel signal; a sixth determining module configured to determine a transformed scale factor for each subband in the target subband set based on the transformed spectral values of the dual-channel signal in each subband included in the target subband set; a seventh determination module configured to determine a second total scale value based on the transformed scale factor and the subband bandwidth of each subband included in the target subband set; an eighth determining module configured to determine the dual-channel signal as a signal to be encoded if the first total scale value is not greater than the second total scale value.
[0259] Optionally, the device 1100 further comprises: and determining the converted dual-channel signal as a signal to be encoded if the first total scale value is greater than the second total scale value and the encoding bit rate of the audio signal is not lower than a first bit rate threshold and / or the energy concentration of the audio signal is greater than a concentration threshold.
[0260] Optionally, the scale factors include a left channel scale factor and a right channel scale factor.
[0261] The device 1100 further comprises: When the first total scale value is greater than the second total scale value, the encoding bit rate of the audio signal is lower than a first bit rate threshold, and the energy concentration of the audio signal is not greater than a concentration threshold, a left channel scale factor and a right channel scale factor of each subband included in the target subband set are calculated based on the left channel scale factor and the right channel scale factor of each subband included in the target subband set. channel Scale Factor and Right channel a ninth determination module configured to determine a difference value between the scale factor; a tenth determining module configured to determine a subband center frequency for each subband included in the target subband set based on an initial frequency and a cutoff frequency for each subband included in the target subband set; If there are any subbands in the target subband set that are larger than the difference threshold, channel Scale Factor and Right channel an eleventh determination module configured to determine the dual-channel signal as a signal to be encoded if there is at least one subband having a difference value between the scale factor and the subband center frequency within the first range.
[0262] Optionally, the device 1100 further comprises: If at least one subband does not exist in the target subband set, the transformed dual-channel signal is determined as the signal to be coded.
[0263] In an embodiment of this application, an optimal subband decomposition scheme is selected from multiple subband decomposition schemes based on audio signal characteristics. In other words, the subband decomposition scheme has signal adaptive characteristics and can adapt to the coding bit rate of the audio signal to improve interference resistance. Specifically, the audio signal is separately divided into multiple subband decomposition schemes, and a total scale factor corresponding to each subband decomposition scheme is determined based on the spectral values of the audio signal in the subbands obtained through the division, the bandwidth of each subband, and the coding bit rate of the audio signal. An optimal target subband decomposition scheme is selected based on the total scale factor to obtain an optimal subband set. Then, spectral envelope shaping is performed based on the scale factor of each subband in the optimal subband set, thereby improving coding efficiency and compression efficiency.
[0264] It should be noted that when the audio signal processing device provided in the above embodiments processes an audio signal, the division of the above functional modules is only used as an example for explanation. In actual applications, the above functions may be allocated to different functional modules for implementation based on requirements. In other words, the internal structure of the device is divided into different functional modules to implement all or part of the above functions. In addition, the audio signal processing device and the audio signal processing method embodiments provided in the above embodiments belong to the same concept. For specific implementation processes, please refer to the method embodiments. Details will not be described again here.
[0265] All or part of the above-described embodiments may be implemented using software, hardware, firmware, or any combination thereof. When software is used for implementation, all or part of the embodiments may be implemented in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer instructions are loaded and executed on a computer, the procedures or functions according to the embodiments of this application are generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from a website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, optical fiber, or digital subscriber line (DSL)) or wireless (e.g., infrared, radio waves, or microwave) transmission. The computer-readable storage medium may be any available medium that can be accessed by a computer, or may be a data storage device that integrates one or more available media, such as a server or a data center. The available medium may be a magnetic medium (e.g., a floppy disk, a hard disk, or a magnetic tape), an optical medium (e.g., a digital versatile disc (DVD)), a semiconductor medium (e.g., a solid-state disk (SSD)), etc. Note that the computer-readable storage medium referred to in the embodiments of this application may be a non-volatile storage medium, i.e., a non-transitory storage medium.
[0266] It should be understood that in this specification, "at least one" refers to one or more, and "plurality" refers to two or more. In the description of the embodiments of this application, unless otherwise specified, " / " means "or." For example, A / B may represent A or B. In this specification, "and / or" only describes an association relationship between related objects and indicates that three relationships may exist. For example, A and / or B may represent three cases: only A exists, both A and B exist, and only B exists. In addition, to clearly describe the technical solutions in the embodiments of this application, terms such as "first" and "second" are used in the embodiments of this application to distinguish between the same or similar items having basically the same function and purpose. It should be understood by those skilled in the art that terms such as "first" and "second" do not limit the number or execution order, and terms such as "first" and "second" do not indicate a clear distinction.
[0267] It should be noted that information (including, but not limited to, user device information and user personal information), data (including, but not limited to, data used for analysis, stored data, and displayed data), and signals in the embodiments of this application are used with permission from the user or full permission from all parties, and the capture, use, and processing of related data must comply with the relevant laws, regulations, and standards of the relevant countries and regions. For example, audio signals related to the embodiments of this application are acquired with full permission.
[0268] The above description is of the embodiments provided in this application and is not intended to limit this application. Nohara Any modification, equivalent replacement, or improvement made without departing from the principles thereof shall fall within the protection scope of this application.
Claims
1. 1. A method for processing an audio signal, comprising: separately performing subband decomposition on the audio signal based on a plurality of subband decomposition schemes and cutoff subbands corresponding to the plurality of subband decomposition schemes to obtain a plurality of candidate subband sets, the plurality of candidate subband sets corresponding to the plurality of subband decomposition schemes one-to-one, and each candidate subband set including a plurality of subbands; determining a total scale value for each candidate subband set based on spectral values of the audio signal in the subbands included in the candidate subband set, a coding bit rate of the audio signal, and subband bandwidths of the subbands included in the candidate subband set; selecting one candidate subband set from the plurality of candidate subband sets as a target subband set based on the total scale factor of each candidate subband set, wherein each subband included in the target subband set has a scale factor used to shape a spectral envelope of the audio signal; How to have that.
2. selecting a candidate subband set from the plurality of candidate subband sets as a target subband set based on the total scale value of each candidate subband set, determining a candidate subband set having a smallest total scale value among the plurality of candidate subband sets as the target subband set; The method of claim 1 , comprising:
3. determining a total scale value for each candidate subband set based on spectral values of the audio signal in the subbands included in the candidate subband set, a coding bit rate of the audio signal, and subband bandwidths of the subbands included in the candidate subband set, for a first candidate subband set among the plurality of candidate subband sets, determining scale factors for each subband included in the first candidate subband set based on spectral values of the audio signal in subbands included in the first candidate subband set, the first candidate subband set being any one of the plurality of candidate subband sets; determining a total scale factor for the first candidate subband set based on the coding bit rate of the audio signal and the scale factor and subband bandwidth of each subband included in the first candidate subband set; 3. The method of claim 1 or 2, comprising:
4. determining a scale factor for each subband included in the first candidate subband set based on spectral values of the audio signal within the subbands included in the first candidate subband set; obtaining a maximum value among absolute values of all spectral values of the audio signal in a first subband included in the first candidate subband set, the first subband being any subband in the first candidate subband set; determining a scale factor for the first subband based on the maximum value; The method of claim 3, comprising:
5. the coding bit rate of the audio signal is not lower than a first bit rate threshold and / or the energy concentration of the audio signal is greater than a concentration threshold; determining a total scale factor for the first candidate subband set based on the coding bit rate of the audio signal and the scale factor and subband bandwidth of each subband included in the first candidate subband set, comprising: determining an energy smoothing reference value based on the coding bit rate of the audio signal and a second bit rate threshold; determining a total energy value for each subband included in the first candidate subband set based on the energy smoothing reference value and the scale factor and the subband bandwidth of each subband included in the first candidate subband set; summing the total energy values of the subbands included in the first candidate subband set to obtain the total scale value for the first candidate subband set; Having that, The method according to claim 3 or 4.
6. determining a total energy value for each subband included in the first candidate subband set based on the energy smoothing reference value and the scale factor and the subband bandwidth of each subband included in the first candidate subband set, determining a reference scale value for a first subband included in the first candidate subband set to be a larger value of the scale factor of the first subband and the energy smoothing reference value, the first subband being any subband in the first candidate subband set; determining a product of the reference scale value of the first subband and a subband bandwidth of the first subband as a total energy value of the first subband; The method of claim 5 , comprising:
7. the coding bit rate of the audio signal is lower than a first bit rate threshold and the energy concentration of the audio signal is not greater than a concentration threshold; determining a total scale factor for the first candidate subband set based on the coding bit rate of the audio signal and the scale factor and subband bandwidth of each subband included in the first candidate subband set, comprising: determining an energy smoothing reference value based on the coding bit rate of the audio signal and a second bit rate threshold; determining a scale difference value for each subband included in the first candidate subband set based on the energy-smoothed reference value and the scale factor of each subband included in the first candidate subband set, the scale difference value indicating a difference between a scale factor of a corresponding subband and a scale factor of an adjacent subband of the corresponding subband; determining the total scale value for the first candidate subband set based on the scale difference values of the subbands included in the first candidate subband set and the subband bandwidth of each subband; Having that, The method according to claim 3 or 4.
8. determining scale differential values for the subbands included in the first candidate subband set based on the energy smoothing reference value and the scale factors of each subband included in the first candidate subband set, comprising: determining a first smoothed value, a second smoothed value, and a third smoothed value for a first subband included in the first candidate subband set based on the energy smoothing reference value, the scale factor of the first subband, and scale factors of adjacent subbands of the first subband, the first subband being any subband in the first candidate subband set; determining a scaled difference value for the first subband based on the first smoothed value, the second smoothed value, and the third smoothed value for the first subband; The method of claim 7, comprising:
9. determining a first smoothed value, a second smoothed value, and a third smoothed value for the first subband based on the energy smoothed reference value, the scale factor of the first subband, and scale factors of adjacent subbands of the first subband, includes: If the first subband is the first subband in the first candidate subband set, determine the larger of the scale factor of the first subband and the energy smoothed reference value as the first smoothed value of the first subband; if the first subband is not the first subband in the first candidate subband set, determine the larger of the scale factor of a previous subband adjacent to the first subband and the energy smoothed reference value as the first smoothed value of the first subband; determining a larger value of the scale factor of the first subband and the energy smoothed reference value as the second smoothed value of the first subband; If the first subband is the last subband in the first candidate subband set, determine the larger value of the scale factor of the first subband and the energy smoothing reference value as the third smoothed value of the first subband; if the first subband is not the last subband in the first candidate subband set, determine the larger value of the scale factor of a next subband adjacent to the first subband and the energy smoothing reference value as the third smoothed value of the first subband. The method of claim 8, comprising:
10. determining a scaled difference value for the first subband based on the first smoothed value, the second smoothed value, and the third smoothed value for the first subband, includes: determining a first difference value and a second difference value for a first subband included in the first candidate subband set, the first difference value being an absolute value of a difference value between the first smoothed value and the second smoothed value for the first subband, and the second difference value being an absolute value of a difference value between the second smoothed value and the third smoothed value for the first subband, the first subband being any subband in the first candidate subband set; determining the scaled difference value for the first subband based on the first difference value and the second difference value for the first subband; 10. The method of claim 8 or 9, comprising:
11. determining the total scale value of the first candidate subband set based on the scale difference values of the subbands included in the first candidate subband set and the subband bandwidth of each subband, comprising: determining a smoothing weighting factor for each subband included in the first candidate subband set based on the number of subbands included in the first candidate subband set and the subband bandwidth of each subband; summing the smoothed weighting factors of the subbands included in the first candidate subband set to obtain a total smoothed weighting factor for the first candidate subband set; multiplying the scaled difference values of the subbands included in the first candidate subband set by the smoothed weighting factors to obtain weighted scaled difference values of the subbands included in the first candidate subband set; summing the weighted scale difference values of the subbands included in the first candidate subband set to obtain a total scale value for the first candidate subband set; Dividing the total scale value by the total smoothed weighting factor for the first candidate subband set to obtain the total scale value for the first candidate subband set.
11. The method of any one of claims 7 to 10, comprising:
12. The method further comprises: when the coding bit rate of the audio signal is lower than the first bit rate threshold, performing bandwidth detection on the spectrum of the audio signal to obtain a cut-off frequency of the audio signal; determining the cutoff subbands corresponding to the plurality of subband division schemes based on the cutoff frequencies; 12. The method of any one of claims 1 to 11, comprising:
13. The method further comprises: determining a last subband indicated by each of the plurality of subband decomposition schemes as a cutoff subband corresponding to each subband decomposition scheme when the coding bit rate of the audio signal is not lower than the first bit rate threshold; 13. The method of any one of claims 1 to 12, comprising:
14. The method further comprises: performing a feature analysis on the spectrum of the audio signal to obtain a feature analysis result; determining the plurality of subband decomposition schemes from a plurality of candidate subband decomposition schemes based on the feature analysis result and the coding bit rate of the audio signal; 14. The method of any one of claims 1 to 13, comprising:
15. 15. The method of claim 14, wherein the feature analysis result comprises a subjective signal flag or an objective signal flag, the subjective signal flag indicating that the energy concentration of the audio signal is not greater than the concentration threshold, and the objective signal flag indicating that the energy concentration of the audio signal is greater than the concentration threshold.
16. the audio signal has a frame length of 10 milliseconds and a sampling rate of 88.2 kilohertz or 96 kilohertz, or the audio signal has a frame length of 5 milliseconds and a sampling rate of 88.2 kilohertz or 96 kilohertz, or the audio signal has a frame length of 10 milliseconds and a sampling rate of 44.1 kilohertz or 48 kilohertz; determining the plurality of subband decomposition schemes from a plurality of candidate subband decomposition schemes based on the feature analysis result and the coding bit rate of the audio signal, determining a first group of subband decomposition schemes from the plurality of candidate subband decomposition schemes as the plurality of subband decomposition schemes when the encoding bit rate of the audio signal is lower than the first bit rate threshold and the feature analysis result has the subjective signal flag; Having that, The subband division scheme of the first group is as follows: { {0,1,2,3,4,6,8,10,13,16,20,24,28,33,38,45,52,61,70,79,88,100,112,127,142,160,178,196,217,238,259,280,480}, {0,1,2,3,5,7,9,12,15,18,22,26,30,35,41,48,56,65,74,84,94,106,118,134,150,166,184,202,220,240,260,280,480}, {0,1,2,3,4,5,7,9,11,14,17,21,25,29,34,40,46,52,60,68,76,86,98,110,126,144,162,180,200,224,250,280,480}, {0,2,4,6,8,12,16,21,26,31,36,41,46,51,56,61,66,71,77,83,89,95,103,111,121,131,147,163,179,203,240,280,480}, {0,1,2,3,5,7,9,12,15,19,23,27,32,37,43,49,57,66,76,86,98,110,125,140,158,176,194,216,238,264,290,320,480}, {0,1,2,3,5,7,10,13,17,21,25,30,35,41,47,54,62,70,80,90,102,114,130,146,162,180,198,218,240,264,290,320,480}, {0,1,2,4,6,8,11,14,18,22,26,30,36,42,50,58,66,76,88,100,112,128,144,160,182,204,226,256,286,316,352,400,480}, {0,1,2,4,6,8,11,14,18,22,26,30,36,42,50,58,68,78,90,102,116,132,148,166,186,208,234,262,292,324,360,400,480} }、 16. The method of claim 15.
17. the audio signal has a frame length of 10 milliseconds and a sampling rate of 88.2 kilohertz or 96 kilohertz, or the audio signal has a frame length of 5 milliseconds and a sampling rate of 88.2 kilohertz or 96 kilohertz, or the audio signal has a frame length of 10 milliseconds and a sampling rate of 44.1 kilohertz or 48 kilohertz; determining the plurality of subband decomposition schemes from a plurality of candidate subband decomposition schemes based on the feature analysis result and the coding bit rate of the audio signal, determining a second group of subband decomposition schemes from the plurality of candidate subband decomposition schemes as the plurality of subband decomposition schemes when the encoding bit rate of the audio signal is not lower than the first bit rate threshold and / or the feature analysis result has the objective signal flag; Having that, The subband division method of the second group is as follows: { {0,1,2,3,4,5,6,7,8,10,12,14,16,19,22,26,30,35,40,45,50,57,64,73,82,92,102,112,124,136,148,160,480}, {0,1,2,3,4,5,7,9,11,13,15,18,21,24,28,33,38,44,50,57,64,73,82,93,104,116,128,140,155,170,185,200,480}, {0,1,2,3,4,6,8,10,13,16,20,24,28,33,38,45,52,61,70,79,88,100,112,127,142,160,178,196,217,238,259,280,480}, {0,1,2,4,6,10,14,18,22,26,30,34,42,50,58,66,74,84,96,108,120,136,152,168,192,216,240,272,304,336,376,424,480}, {0,1,2,4,6,10,14,18,26,34,42,50,62,74,86,98,112,128,144,160,176,196,216,236,256,280,304,328,352,384,416,448,480}, {0,80,92,104,112,120,128,136,144,148,152,156,160,164,168,172,176,180,184,188,192,196,200,208,216,224,232,240,248,256,268,280,480}, {0,200,212,224,232,240,248,256,264,268,272,276,280,284,288,292,296,300,304,308,312,316,320,328,336,344,352,360,368,376,388,400,480}, {0,320,332,344,356,364,372,380,384,388,392,396,400,404,408,412,416,420,424,428,432,436,440,444,448,452,456,460,464,468,472,476,480} }、 16. The method of claim 15.
18. The frame length of the audio signal is 5 milliseconds, and the sampling rate is 44.1 kHz or 48 kHz; determining the plurality of subband decomposition schemes from a plurality of candidate subband decomposition schemes based on the feature analysis result and the coding bit rate of the audio signal, determining a third group of subband decomposition schemes from among the plurality of candidate subband decomposition schemes as the plurality of subband decomposition schemes when the encoding bit rate of the audio signal is lower than the first bit rate threshold and the feature analysis result has the subjective signal flag; Having that, The subband division method of the third group is as follows: { {0,1,2,3,4,5,6,7,8,9,10,12,14,16,19,22,26,30,35,39,44,50,56,63,71,80,89,98,108,119,129,140,240}, {0,1,2,3,4,5,6,7,8,9,11,13,15,17,20,24,28,32,37,42,47,53,59,67,75,83,92,101,110,120,130,140,240}, {0,1,2,3,4,5,6,7,8,9,10,11,12,14,17,20,23,26,30,34,38,43,49,55,63,72,81,90,100,112,125,140,240}, {0,1,2,3,4,6,8,10,13,15,18,20,23,25,28,30,33,35,38,41,44,47,51,55,60,65,73,81,89,101,120,140,240}, {0,1,2,3,4,5,6,7,9,11,13,14,16,18,21,24,28,33,38,43,49,55,62,70,79,88,97,108,119,132,145,160,240}, {0,1,2,3,4,5,6,7,8,10,12,14,17,20,23,27,31,35,40,45,51,57,65,73,81,90,99,109,120,132,145,160,240}, {0,1,2,3,4,5,6,7,9,11,13,15,18,21,25,29,33,38,44,50,56,64,72,80,91,102,113,128,143,158,176,200,240}, {0,1,2,3,4,5,6,7,9,11,13,15,18,21,25,29,34,39,45,51,58,66,74,83,93,104,117,131,146,162,180,200,240} }、 16. The method of claim 15.
19. The frame length of the audio signal is 5 milliseconds, and the sampling rate is 44.1 kHz or 48 kHz; determining the plurality of subband decomposition schemes from a plurality of candidate subband decomposition schemes based on the feature analysis result and the coding bit rate of the audio signal, determining a fourth group of subband decomposition schemes from the plurality of candidate subband decomposition schemes as the plurality of subband decomposition schemes when the encoding bit rate of the audio signal is not lower than the first bit rate threshold and / or the feature analysis result has the objective signal flag; Having that, The subband division method of the fourth group is as follows: { {0,1,2,3,4,5,6,7,8,9,10,11,12,13,14,15,16,17,18,19,20,21,22,23,24,26,28,30,32,34,37,40,120}, {0,1,2,3,4,5,6,7,8,9,10,11,12,13,14,15,16,18,20,22,24,26,28,30,32,34,36,38,41,44,47,50,120}, {0,1,2,3,4,5,6,7,8,9,10,11,12,14,16,18,20,22,24,26,28,31,34,37,40,44,48,52,56,60,65,70,120}, {0,1,2,3,4,5,6,7,8,9,10,11,12,13,15,17,19,21,24,27,30,34,38,42,48,54,60,68,76,84,94,106,120}, {0,1,2,3,4,5,6,7,8,10,12,14,16,19,22,25,28,32,36,40,44,49,54,59,64,70,76,82,88,96,104,112,120}, {0,20,23,26,28,30,32,34,36,37,38,39,40,41,42,43,44,45,46,47,48,49,50,52,54,56,58,60,62,64,67,70,120}, {0,50,53,56,58,60,62,64,66,67,68,69,70,71,72,73,74,75,76,77,78,79,80,82,84,86,88,90,92,94,97,100,120}, {0,80,83,86,89,91,93,95,96,97,98,99,100,101,102,103,104,105,106,107,108,109,110,111,112,113,114,115,116,117,118,119,120} }、 16. The method of claim 15.
20. the audio signal is a dual-channel signal; The method further comprises: determining a first total scale value based on the scale factor and a subband bandwidth of each subband included in the target subband set; performing mid / side stereo transform coding on the spectrum of the dual-channel signal to obtain a transformed spectrum of the dual-channel signal; determining a transformed scale factor for each subband in the target subband set based on the transformed spectral values of the dual-channel signal in each subband included in the target subband set; determining a second total scale value based on the transformed scale factor and the subband bandwidth of each subband included in the target subband set; determining the dual-channel signal as a signal to be encoded if the first total scale value is not greater than the second total scale value; 20. The method of any one of claims 1 to 19, comprising:
21. The method further comprises: determining the converted dual-channel signal as a signal to be coded if the first total scale value is greater than the second total scale value, and the coding bit rate of the audio signal is not lower than the first bit rate threshold and / or the energy concentration of the audio signal is greater than the concentration threshold; 21. The method of claim 20, comprising:
22. the scale factors include a left channel scale factor and a right channel scale factor; The method further comprises: when the first total scale value is greater than the second total scale value, the coding bit rate of the audio signal is lower than the first bit rate threshold, and the energy concentration of the audio signal is not greater than the concentration threshold, determining a difference value between the left scale factor and the right scale factor of each subband included in the target subband set based on a left channel scale factor and a right channel scale factor of each subband included in the target subband set; determining a subband center frequency of each subband included in the target subband set based on an initial frequency and a cutoff frequency of each subband included in the target subband set; determining the dual-channel signal as a signal to be coded if there is at least one subband in the target subband set that has a difference value between a left scale factor and a right scale factor that is greater than a difference threshold and has a subband center frequency that is within a first range; 22. The method of claim 20 or 21, comprising:
23. The method further comprises: determining the transformed dual-channel signal as a signal to be coded if the at least one subband does not exist in the target subband set; 23. The method of claim 22, comprising:
24. 1. An audio signal processing apparatus, comprising: a subband decomposition module configured to separately perform subband decomposition on the audio signal based on a plurality of subband decomposition schemes and cutoff subbands corresponding to the plurality of subband decomposition schemes to obtain a plurality of candidate subband sets, the plurality of candidate subband sets corresponding to the plurality of subband decomposition schemes one-to-one, and each candidate subband set including a plurality of subbands; a first determination module configured to determine a total scale value for each candidate subband set based on spectral values of the audio signal in the subbands included in the candidate subband set, a coding bit rate of the audio signal, and subband bandwidths of the subbands included in the candidate subband set; a selection module configured to select one candidate subband set from the plurality of candidate subband sets as a target subband set based on the total scale factor of each candidate subband set, wherein each subband included in the target subband set has a scale factor used to shape a spectral envelope of the audio signal; and An apparatus having:
25. The selection module: determining a candidate subband set having a smallest total scale value among the plurality of candidate subband sets as the target subband set; 25. The apparatus of claim 24, configured to:
26. The first determination module: a first determination sub-module configured to determine, for a first candidate subband set among the plurality of candidate subband sets, scale factors for each subband included in the first candidate subband set based on spectral values of the audio signal in subbands included in the first candidate subband set, the first candidate subband set being any one of the plurality of candidate subband sets; a second determination sub-module configured to determine a total scale factor for the first candidate subband set based on the coding bit rate of the audio signal and the scale factor and subband bandwidth of each subband included in the first candidate subband set; 26. The device according to claim 24 or 25, comprising:
27. The second determination sub-module: obtaining a maximum value among absolute values of all spectral values of the audio signal in a first subband included in the first candidate subband set, the first subband being any subband in the first candidate subband set; determining a scale factor for the first subband based on the maximum value; 27. The apparatus of claim 26, configured to:
28. the coding bit rate of the audio signal is not lower than a first bit rate threshold and / or the energy concentration of the audio signal is greater than a concentration threshold; The second determination sub-module: determining an energy smoothing reference value based on the coding bit rate of the audio signal and a second bit rate threshold; determining a total energy value for each subband included in the first candidate subband set based on the energy smoothing reference value and the scale factor and the subband bandwidth of each subband included in the first candidate subband set; summing the total energy values of the subbands included in the first candidate subband set to obtain the total scale value for the first candidate subband set; It is configured as follows:
28. Apparatus according to claim 26 or 27.
29. The second determination sub-module: determining a reference scale value for a first subband included in the first candidate subband set to be a larger value of the scale factor of the first subband and the energy smoothing reference value, the first subband being any subband in the first candidate subband set; determining a product of the reference scale value of the first subband and a subband bandwidth of the first subband as a total energy value of the first subband; 29. The apparatus of claim 28, configured to:
30. the coding bit rate of the audio signal is lower than a first bit rate threshold and the energy concentration of the audio signal is not greater than a concentration threshold; The second determination sub-module: determining an energy smoothing reference value based on the coding bit rate of the audio signal and a second bit rate threshold; determining a scale difference value for each subband included in the first candidate subband set based on the energy-smoothed reference value and the scale factor of each subband included in the first candidate subband set, the scale difference value indicating a difference between a scale factor of a corresponding subband and a scale factor of an adjacent subband of the corresponding subband; determining the total scale value for the first candidate subband set based on the scale difference values of the subbands included in the first candidate subband set and the subband bandwidth of each subband; It is configured as follows:
28. Apparatus according to claim 26 or 27.
31. The second determination sub-module: determining a first smoothed value, a second smoothed value, and a third smoothed value for a first subband included in the first candidate subband set based on the energy smoothing reference value, the scale factor of the first subband, and scale factors of adjacent subbands of the first subband, the first subband being any subband in the first candidate subband set; determining a scaled difference value for the first subband based on the first smoothed value, the second smoothed value, and the third smoothed value for the first subband; 31. The apparatus of claim 30, configured to:
32. The second decision sub-module: If the first subband is the first subband in the first candidate subband set, determine the larger of the scale factor of the first subband and the energy smoothed reference value as the first smoothed value of the first subband; if the first subband is not the first subband in the first candidate subband set, determine the larger of the scale factor of a previous subband adjacent to the first subband and the energy smoothed reference value as the first smoothed value of the first subband; determining a larger value of the scale factor of the first subband and the energy smoothed reference value as the second smoothed value of the first subband; If the first subband is the last subband in the first candidate subband set, determine the larger value of the scale factor of the first subband and the energy smoothing reference value as the third smoothed value of the first subband; if the first subband is not the last subband in the first candidate subband set, determine the larger value of the scale factor of a next subband adjacent to the first subband and the energy smoothing reference value as the third smoothed value of the first subband.
32. The apparatus of claim 31 configured to:
33. The second determination sub-module: determining a first difference value and a second difference value for a first subband included in the first candidate subband set, the first difference value being an absolute value of a difference value between the first smoothed value and the second smoothed value for the first subband, and the second difference value being an absolute value of a difference value between the second smoothed value and the third smoothed value for the first subband, the first subband being any subband in the first candidate subband set; determining the scaled difference value for the first subband based on the first difference value and the second difference value for the first subband; 33. The apparatus of claim 31 or 32, configured to:
34. The second determination sub-module: determining a smoothing weighting factor for each subband included in the first candidate subband set based on the number of subbands included in the first candidate subband set and the subband bandwidth of each subband; summing the smoothed weighting factors of the subbands included in the first candidate subband set to obtain a total smoothed weighting factor for the first candidate subband set; multiplying the scaled difference values of the subbands included in the first candidate subband set by the smoothed weighting factors to obtain weighted scaled difference values of the subbands included in the first candidate subband set; summing the weighted scale difference values of the subbands included in the first candidate subband set to obtain a total scale value for the first candidate subband set; Dividing the total scale value by the total smoothed weighting factor for the first candidate subband set to obtain the total scale value for the first candidate subband set.
34. Apparatus according to any one of claims 30 to 33, configured to:
35. The apparatus further comprises: a bandwidth detection module configured to, when the encoding bit rate of the audio signal is lower than the first bit rate threshold, perform bandwidth detection on the spectrum of the audio signal to obtain a cut-off frequency of the audio signal; a second determination module configured to determine the cutoff subbands corresponding to the plurality of subband division schemes based on the cutoff frequencies; 35. The apparatus of any one of claims 24 to 34, comprising:
36. The apparatus further comprises: a third determination module configured to determine a last subband indicated by each of the plurality of subband decomposition schemes as a cutoff subband corresponding to each subband decomposition scheme when the coding bit rate of the audio signal is not lower than the first bit rate threshold; 36. The apparatus of any one of claims 24 to 35, comprising:
37. The apparatus further comprises: a feature analysis module configured to perform feature analysis on the spectrum of the audio signal to obtain a feature analysis result; a fourth determination module configured to determine the plurality of subband decomposition schemes from a plurality of candidate subband decomposition schemes based on the feature analysis result and the coding bit rate of the audio signal; 37. The apparatus of any one of claims 24 to 36, comprising:
38. 38. The apparatus of claim 37, wherein the feature analysis result comprises a subjective signal flag or an objective signal flag, the subjective signal flag indicating that the energy concentration of the audio signal is not greater than the concentration threshold, and the objective signal flag indicating that the energy concentration of the audio signal is greater than the concentration threshold.
39. the audio signal has a frame length of 10 milliseconds and a sampling rate of 88.2 kilohertz or 96 kilohertz, or the audio signal has a frame length of 5 milliseconds and a sampling rate of 88.2 kilohertz or 96 kilohertz, or the audio signal has a frame length of 10 milliseconds and a sampling rate of 44.1 kilohertz or 48 kilohertz; The fourth determination module: a third determination sub-module configured to determine a first group of subband decomposition schemes from the plurality of candidate subband decomposition schemes as the plurality of subband decomposition schemes when the coding bit rate of the audio signal is lower than the first bit rate threshold and the feature analysis result has the subjective signal flag; and The subband division scheme of the first group is as follows: { {0,1,2,3,4,6,8,10,13,16,20,24,28,33,38,45,52,61,70,79,88,100,112,127,142,160,178,196,217,238,259,280,480}, {0,1,2,3,5,7,9,12,15,18,22,26,30,35,41,48,56,65,74,84,94,106,118,134,150,166,184,202,220,240,260,280,480}, {0,1,2,3,4,5,7,9,11,14,17,21,25,29,34,40,46,52,60,68,76,86,98,110,126,144,162,180,200,224,250,280,480}, {0,2,4,6,8,12,16,21,26,31,36,41,46,51,56,61,66,71,77,83,89,95,103,111,121,131,147,163,179,203,240,280,480}, {0,1,2,3,5,7,9,12,15,19,23,27,32,37,43,49,57,66,76,86,98,110,125,140,158,176,194,216,238,264,290,320,480}, {0,1,2,3,5,7,10,13,17,21,25,30,35,41,47,54,62,70,80,90,102,114,130,146,162,180,198,218,240,264,290,320,480}, {0,1,2,4,6,8,11,14,18,22,26,30,36,42,50,58,66,76,88,100,112,128,144,160,182,204,226,256,286,316,352,400,480}, {0,1,2,4,6,8,11,14,18,22,26,30,36,42,50,58,68,78,90,102,116,132,148,166,186,208,234,262,292,324,360,400,480} }、 39. The apparatus of claim 38.
40. the audio signal has a frame length of 10 milliseconds and a sampling rate of 88.2 kilohertz or 96 kilohertz, or the audio signal has a frame length of 5 milliseconds and a sampling rate of 88.2 kilohertz or 96 kilohertz, or the audio signal has a frame length of 10 milliseconds and a sampling rate of 44.1 kilohertz or 48 kilohertz; The fourth determination module: a fourth determination sub-module configured to determine a second group of subband decomposition schemes from the plurality of candidate subband decomposition schemes as the plurality of subband decomposition schemes when the coding bit rate of the audio signal is not lower than the first bit rate threshold and / or the feature analysis result has the objective signal flag; and The subband division method of the second group is as follows: { {0,1,2,3,4,5,6,7,8,10,12,14,16,19,22,26,30,35,40,45,50,57,64,73,82,92,102,112,124,136,148,160,480}, {0,1,2,3,4,5,7,9,11,13,15,18,21,24,28,33,38,44,50,57,64,73,82,93,104,116,128,140,155,170,185,200,480}, {0,1,2,3,4,6,8,10,13,16,20,24,28,33,38,45,52,61,70,79,88,100,112,127,142,160,178,196,217,238,259,280,480}, {0,1,2,4,6,10,14,18,22,26,30,34,42,50,58,66,74,84,96,108,120,136,152,168,192,216,240,272,304,336,376,424,480}, {0,1,2,4,6,10,14,18,26,34,42,50,62,74,86,98,112,128,144,160,176,196,216,236,256,280,304,328,352,384,416,448,480}, {0,80,92,104,112,120,128,136,144,148,152,156,160,164,168,172,176,180,184,188,192,196,200,208,216,224,232,240,248,256,268,280,480}, {0,200,212,224,232,240,248,256,264,268,272,276,280,284,288,292,296,300,304,308,312,316,320,328,336,344,352,360,368,376,388,400,480}, {0,320,332,344,356,364,372,380,384,388,392,396,400,404,408,412,416,420,424,428,432,436,440,444,448,452,456,460,464,468,472,476,480} }、 39. The apparatus of claim 38.
41. The frame length of the audio signal is 5 milliseconds, and the sampling rate is 44.1 kHz or 48 kHz; The fourth determination module: a fifth determination sub-module configured to determine a third group of subband decomposition schemes from the plurality of candidate subband decomposition schemes as the plurality of subband decomposition schemes when the coding bit rate of the audio signal is lower than the first bit rate threshold and the feature analysis result has the subjective signal flag; and The subband division method of the third group is as follows: { {0,1,2,3,4,5,6,7,8,9,10,12,14,16,19,22,26,30,35,39,44,50,56,63,71,80,89,98,108,119,129,140,240}, {0,1,2,3,4,5,6,7,8,9,11,13,15,17,20,24,28,32,37,42,47,53,59,67,75,83,92,101,110,120,130,140,240}, {0,1,2,3,4,5,6,7,8,9,10,11,12,14,17,20,23,26,30,34,38,43,49,55,63,72,81,90,100,112,125,140,240}, {0,1,2,3,4,6,8,10,13,15,18,20,23,25,28,30,33,35,38,41,44,47,51,55,60,65,73,81,89,101,120,140,240}, {0,1,2,3,4,5,6,7,9,11,13,14,16,18,21,24,28,33,38,43,49,55,62,70,79,88,97,108,119,132,145,160,240}, {0,1,2,3,4,5,6,7,8,10,12,14,17,20,23,27,31,35,40,45,51,57,65,73,81,90,99,109,120,132,145,160,240}, {0,1,2,3,4,5,6,7,9,11,13,15,18,21,25,29,33,38,44,50,56,64,72,80,91,102,113,128,143,158,176,200,240}, {0,1,2,3,4,5,6,7,9,11,13,15,18,21,25,29,34,39,45,51,58,66,74,83,93,104,117,131,146,162,180,200,240} }、 39. The apparatus of claim 38.
42. The frame length of the audio signal is 5 milliseconds, and the sampling rate is 44.1 kHz or 48 kHz; The fourth determination module: a sixth determination sub-module configured to determine a fourth group of subband decomposition schemes from the plurality of candidate subband decomposition schemes as the plurality of subband decomposition schemes when the coding bit rate of the audio signal is not lower than the first bit rate threshold and / or the feature analysis result has the objective signal flag; and The subband division method of the fourth group is as follows: { {0,1,2,3,4,5,6,7,8,9,10,11,12,13,14,15,16,17,18,19,20,21,22,23,24,26,28,30,32,34,37,40,120}, {0,1,2,3,4,5,6,7,8,9,10,11,12,13,14,15,16,18,20,22,24,26,28,30,32,34,36,38,41,44,47,50,120}, {0,1,2,3,4,5,6,7,8,9,10,11,12,14,16,18,20,22,24,26,28,31,34,37,40,44,48,52,56,60,65,70,120}, {0,1,2,3,4,5,6,7,8,9,10,11,12,13,15,17,19,21,24,27,30,34,38,42,48,54,60,68,76,84,94,106,120}, {0,1,2,3,4,5,6,7,8,10,12,14,16,19,22,25,28,32,36,40,44,49,54,59,64,70,76,82,88,96,104,112,120}, {0,20,23,26,28,30,32,34,36,37,38,39,40,41,42,43,44,45,46,47,48,49,50,52,54,56,58,60,62,64,67,70,120}, {0,50,53,56,58,60,62,64,66,67,68,69,70,71,72,73,74,75,76,77,78,79,80,82,84,86,88,90,92,94,97,100,120}, {0,80,83,86,89,91,93,95,96,97,98,99,100,101,102,103,104,105,106,107,108,109,110,111,112,113,114,115,116,117,118,119,120} }、 39. The apparatus of claim 38.
43. the audio signal is a dual-channel signal; The apparatus further comprises: a fifth determination module configured to determine a first total scale value based on the scale factor and a subband bandwidth of each subband included in the target subband set; a transform module configured to perform mid / side stereo transform coding on the spectrum of the dual channel signal to obtain a transformed spectrum of the dual channel signal; a sixth determination module configured to determine a transformed scale factor for each subband in the target subband set based on the transformed spectral values of the dual-channel signal in each subband included in the target subband set; a seventh determination module configured to determine a second total scale value based on the transformed scale factor and the subband bandwidth of each subband included in the target subband set; an eighth determination module configured to determine the dual-channel signal as a signal to be encoded if the first total scale value is not greater than the second total scale value; 43. The apparatus of any one of claims 24 to 42, comprising:
44. The apparatus further comprises: determining the converted dual-channel signal as a signal to be coded if the first total scale value is greater than the second total scale value, and the coding bit rate of the audio signal is not lower than the first bit rate threshold and / or the energy concentration of the audio signal is greater than the concentration threshold; 44. The apparatus of claim 43, configured to:
45. the scale factors include a left channel scale factor and a right channel scale factor; The apparatus further comprises: a ninth determination module configured to determine a difference value between the left and right scale factors of each subband included in the target subband set based on a left channel scale factor and a right channel scale factor of each subband included in the target subband set when the first total scale value is greater than the second total scale value, the coding bit rate of the audio signal is lower than the first bit rate threshold, and the energy concentration of the audio signal is not greater than the concentration threshold; and a tenth determination module configured to determine a subband center frequency of each subband included in the target subband set based on an initial frequency and a cutoff frequency of each subband included in the target subband set; an eleventh determination module configured to determine the dual-channel signal as a signal to be encoded if there is at least one subband in the target subband set that has a difference value between a left scale factor and a right scale factor that is greater than a difference threshold and has a subband center frequency that is within a first range; 45. The apparatus of claim 43 or 44, comprising:
46. The apparatus further comprises: determining the transformed dual-channel signal as a signal to be coded if the at least one subband does not exist in the target subband set; 46. The apparatus of claim 45, configured to:
47. 1. An audio signal processing device having a memory and a processor, the memory is configured to store a computer program, the computer program having program instructions; The processor is configured to call the computer program to perform the method of any one of claims 1 to 23. Audio signal processing device.
48. 24. A computer readable storage medium having stored thereon a computer program, which, when executed by a processor, performs the steps of the method according to any one of claims 1 to 23.
49. A computer program product storing computer instructions which, when executed by a processor, cause the steps of the method of any one of claims 1 to 23 to be performed.
Citation Information
Patent Citations
Voice signal decoder
JP1996016193A
Frequency Segmentation to Obtain Bands for Efficient Coding of Digital Media
JP2009501945A