Methods, apparatus, storage media and computer program products for extending audio bandwidth
By combining super-resolution networks with high-pass filtering, the problem of high-frequency component loss in audio bandwidth expansion is solved, achieving improved sound quality and efficiency, and adapting to audio signal processing with different cutoff frequencies and bit rates.
Patent Information
- Application Number
- CN202210693082.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-17
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2042-06-17
AI Technical Summary
Existing audio bandwidth extension technologies result in the loss of high-frequency components, reduced sound quality, and are also highly complex and inefficient.
A method combining super-resolution network and high-pass filtering is adopted. By determining the high-pass filtering parameters and super-resolution signal processing, the high-frequency components of the audio are expanded, and different cutoff frequencies and bit rates are adaptively applied during the decoding process.
It improves audio quality, making the listening experience more harmonious and natural, reduces complexity, increases bandwidth expansion efficiency, and keeps the low-frequency components essentially unchanged.
Smart Images

Figure CN117292699B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of audio processing technology, and in particular to an audio bandwidth extension method, apparatus, storage medium, and computer program product. Background Technology
[0002] The wider the bandwidth of an audio signal, the richer its signal components, and the better the sound quality. However, due to limitations in data transmission and storage, audio data compression, or encoding, is usually required. Encoding leads to the loss of high-frequency components, resulting in a decrease in sound quality and a poorer listening experience for the decoded audio. Therefore, to improve sound quality, it is necessary to extend the bandwidth of the audio signal to increase its high-frequency components.
[0003] In related technologies, spectral band replication (SBR) is used to extend the bandwidth of audio to increase its high-frequency components. For example, the high-frequency components of the audio can be obtained by translating or flipping the modified discrete cosine transform (MDCT) spectrum and then performing envelope shaping and noise filling on the resulting MDCT spectrum.
[0004] However, the high-frequency components obtained after bandwidth expansion of audio using bandwidth replication technology often bear obvious replication marks, resulting in insufficient sound quality and an unnatural listening experience. Furthermore, envelope shaping of the MDCT spectrum is highly complex, reducing the efficiency of audio bandwidth expansion. Summary of the Invention
[0005] This application provides a method, apparatus, storage medium, and computer program product for extending audio bandwidth. It can improve audio quality through bandwidth extension, resulting in a more natural listening experience, while reducing complexity and increasing efficiency. The technical solution is as follows:
[0006] In a first aspect, an audio bandwidth extension method is provided, the method comprising:
[0007] A first high-pass filter parameter is determined based on a first bandwidth ratio. The first bandwidth ratio indicates the position of the first cutoff frequency in the first frequency band. The first cutoff frequency is the cutoff frequency of the first audio signal to be bandwidth extended. The first audio signal contains a first frequency component, which includes frequency components in the first frequency band that are not greater than the first cutoff frequency. The first frequency band is determined based on the sampling rate of the first audio signal. The first audio signal is input into a first super-resolution network to obtain a first super-resolution signal. The first super-resolution signal contains a portion of the first frequency component and a second frequency component. The second frequency component includes frequency components in the first frequency band that are greater than the first cutoff frequency. The first super-resolution signal is high-pass filtered according to the first high-pass filter parameter to obtain a second audio signal, which contains a second frequency component. The first audio signal and the second audio signal are superimposed to obtain a bandwidth-extended audio signal.
[0008] This scheme utilizes a super-resolution network to extend high-frequency components, resulting in a more harmonious and natural sound in the final audio, with lower complexity and improved bandwidth extension efficiency. Furthermore, this scheme combines the super-resolution network with a high-pass filter. The high-pass filter parameters are determined based on the cutoff frequency of the audio signal to be bandwidth extended, demonstrating that this scheme can adapt to different cutoff frequencies, and the super-resolution network can handle audio signals with various cutoff frequencies. When the audio signal to be bandwidth extended is the audio signal obtained during the decoding process, since the cutoff frequency is related to the bit rate, this scheme can adapt to different bit rates. In addition, this scheme uses high-pass filtering to ensure that the low-frequency components of the final audio signal remain essentially unchanged, meaning that the low-frequency components are not damaged.
[0009] Optionally, determining the first high-pass filter parameter based on the first bandwidth proportion includes: determining the first bandwidth proportion range from multiple bandwidth proportion ranges; and selecting a set of high-pass filter parameters corresponding to the first bandwidth proportion range from multiple sets of high-pass filter parameters corresponding to the multiple bandwidth proportion ranges, as the first high-pass filter parameter. Wherein, the multiple bandwidth proportion ranges correspond one-to-one with the multiple sets of high-pass filter parameters, and the multiple bandwidth proportion ranges do not overlap.
[0010] This scheme can be applied during the decoding process, for example, at the decoding end. Based on this, before determining the first high-pass filter parameters according to the first bandwidth ratio, the method further includes: parsing and inverse quantizing the bitstream to obtain the MDCT spectrum; performing bandwidth detection on the MDCT spectrum to obtain the first cutoff frequency; and dividing the first cutoff frequency by the bandwidth of the first frequency band to obtain the first bandwidth ratio. Correspondingly, before inputting the first audio signal into the first super-resolution network to obtain the first super-resolution signal, the method further includes: determining the first audio signal based on the MDCT spectrum.
[0011] The bitstream is obtained by encoding the audio signal to be compressed at the encoding end, using MDCT. The bandwidth of the audio signal to be compressed is greater than the bandwidth of the first audio signal. For example, if the bandwidth of the audio signal to be compressed is equal to the bandwidth of the first frequency band, then the audio signal to be compressed is a full-bandwidth signal. Furthermore, the bandwidth of the audio signal refers to its actual bandwidth, or effective bandwidth. The bandwidth detection mentioned above refers to effective bandwidth detection, and the bandwidth percentage refers to the proportion of effective bandwidth to the corresponding frequency band; this bandwidth percentage can be called the effective bandwidth percentage.
[0012] In the encoding process of this application embodiment, the encoding end divides the audio signal to be compressed into frames to obtain multiple audio frames of the signal to be compressed. The encoding end then applies windowing to each of these multiple audio frames of the signal to be compressed to obtain multiple frames of windowed signals to be compressed. The encoding end divides the j-th and j+1-th audio frames into an audio segment, and performs MDCT and quantization on the windowed signal to be compressed corresponding to this audio segment to obtain the encoded bitstream corresponding to that audio segment in the bitstream. Here, j is an integer not less than 0.
[0013] Optionally, based on the MDCT spectrum, determining the first audio signal includes: performing an improved inverse cosine transform (IMDCT) on the MDCT spectrum to obtain the first audio signal; wherein the bandwidth-extended audio signal includes the i-th audio frame and the (i+1)-th audio frame, where i is an integer not less than 1; after superimposing the first audio signal with the second audio signal to obtain the bandwidth-extended audio signal, the method further includes: performing overlap and addition (OLA) on the bandwidth-extended audio signal and the bandwidth-extended reference audio signal to obtain the reconstructed signal of the i-th audio frame, wherein the reference audio signal includes the (i-1)-th audio frame and the i-th audio frame, and the reconstructed signal of the i-th audio frame contains a first frequency component and a second frequency component. It should be noted that the first audio signal, the bandwidth-extended audio signal, and the reference audio signal are all windowed signals, and the OLA operation can be used to remove the window so that the reconstructed signal is the windowed signal. Furthermore, performing IMDCT, super-resolution, and OLA sequentially ensures that the decoding process does not introduce additional delay.
[0014] This solution can also be applied to any device to extend the bandwidth of narrow-bandwidth audio signals stored in these devices. Based on this, before inputting the first audio signal into the first super-resolution network to obtain the first super-resolution signal, the method further includes: upsampling the target audio signal according to a first upsampling parameter to obtain the first audio signal. The target audio signal refers to an audio signal that has frequency components in a second frequency band, where the second frequency band is determined based on the sampling rate of the target audio signal. Correspondingly, before determining the first high-pass filter parameter based on the first bandwidth ratio, the method further includes: determining the first bandwidth ratio based on the first upsampling parameter. In simple terms, the method first upsamples the full-bandwidth but low-quality target audio signal to obtain a partial-bandwidth first audio signal, and then expands the first audio signal into a full-bandwidth audio signal, thereby enhancing the sound quality.
[0015] Optionally, the first super-resolution network includes a first processing module, a second processing module, and a third processing module. The second processing module includes a first convolution submodule, a second convolution submodule, and an addition submodule. The dilation rate of the convolutional layers in the first and second convolution submodules is greater than 1. Inputting a first audio signal into the first super-resolution network to obtain a first super-resolution signal includes: inputting the first audio signal into the first processing module to obtain first data, the first data containing a first frequency component and a second frequency component; inputting the first data into the first convolution submodule to obtain second data; inputting the second data into the second convolution submodule to obtain third data; inputting the second and third data into the addition submodule to obtain fourth data, the fourth data containing a second frequency component, or the fourth data containing a portion of the first frequency component and a second frequency component; and inputting the fourth data into the third processing module to obtain the first super-resolution signal.
[0016] Optionally, the first processing module includes a nonlinear activation layer, and the second frequency component is a frequency component extended by the nonlinear activation layer.
[0017] It should be understood that the first processing mode is used to expand high-frequency components (such as the second frequency component). In the embodiments of this application, the first processing module includes a nonlinear activation layer, and the second frequency component is a frequency component expanded by the nonlinear activation layer. The nonlinear activation layer has a nonlinear filtering function, and the nonlinear filtering can generate a frequency doubling effect, enabling the first processing module to expand high-frequency components, and the spectral structure of the expanded high-frequency components has good consistency with the spectral structure of the frequency components contained in the first audio signal. While expanding high-frequency components, the first processing module may also expand some low-frequency components (such as a part of the first frequency component). The second processing module is used to minimize the expanded low-frequency components. Specifically, since the dilation rate of the convolutional layers in the second processing module is greater than 1, these convolutional layers implicitly have a downsampling function, and with the action of the addition submodule, the second processing module can reduce the expanded low-frequency components.
[0018] Optionally, before inputting the first audio signal into the first super-resolution network to obtain the first super-resolution signal, the method further includes: selecting a super-resolution network corresponding to the sampling rate of the first audio signal from a plurality of super-resolution networks corresponding to a plurality of sampling rates, and using this as the first super-resolution network. It should be understood that multiple super-resolution networks are deployed in the electronic device, and these multiple super-resolution networks correspond one-to-one with the plurality of sampling rates; that is, one super-resolution network corresponds to one sampling rate, different super-resolution networks correspond to different sampling rates, and each super-resolution network is used to perform bandwidth expansion on the audio signal at the corresponding sampling rate.
[0019] Optionally, before selecting the super-resolution network corresponding to the sampling rate of the first audio signal from the multiple super-resolution networks corresponding to the multiple sampling rates as the first super-resolution network, the method further includes: acquiring multiple sets of audio sample sets, each set including multiple audio sample signals, wherein the sampling rates of the audio sample signals in different sets are different, and the sampling rates of the audio sample signals in the same set are the same; and determining multiple super-resolution networks based on the multiple sets of audio sample sets. Each set of audio sample sets corresponds one-to-one with a super-resolution network; one set of audio sample sets is used to determine one super-resolution network, and different sets of audio sample sets are used to determine different super-resolution networks.
[0020] Optionally, based on the multiple sets of audio samples, multiple super-resolution networks are determined, including: selecting a set of audio sample signals from the multiple sets of audio sample sets as a first audio sample set, and performing the following operations on the first audio sample set until the following operations are performed on all multiple sets of audio sample sets:
[0021] Noise is added to multiple audio sample signals in the first audio sample set to obtain multiple first sample signals; the multiple first sample signals are low-pass filtered according to one or more sets of low-pass filtering parameters to obtain multiple second sample signals; the multiple second sample signals are used as the input of the initial super-resolution network, and the multiple first sample signals are used as the output of the initial super-resolution network to train the initial super-resolution network to obtain a super-resolution network.
[0022] The electronic device performs low-pass filtering on the multiple first sample signals according to a set of low-pass filtering parameters. Alternatively, the electronic device groups the multiple first sample signals into multiple groups of first sample signals, each group corresponding one-to-one with a set of low-pass filtering parameters, with each set of parameters used to perform low-pass filtering on its corresponding group. The electronic device then performs low-pass filtering on each of the multiple groups of first sample signals according to these multiple sets of low-pass filtering parameters to obtain multiple groups of second sample signals. That is, the electronic device performs low-pass filtering on each group of first sample signals according to the low-pass filtering parameters in the multiple sets of parameters. The cutoff frequencies of the second sample signals in different groups are different, which is more conducive to training a super-resolution network that can adapt to different cutoff frequencies.
[0023] Furthermore, by adding noise to the audio sample signals, the trained super-resolution network becomes more sensitive to noise signals and exhibits high gain for noise-like components in the audio signal. In other words, the trained super-resolution network can also extend the bandwidth for noise-like components in the audio signal. The aforementioned multiple sets of audio samples can include music sample signals and / or speech sample signals of different styles, thus ensuring good robustness of this scheme for both speech and music signals.
[0024] Optionally, this solution can also perform secondary bandwidth expansion to further improve audio quality. In one implementation, after superimposing the first audio signal and the second audio signal to obtain a bandwidth-expanded audio signal, the solution further includes: upsampling the bandwidth-expanded audio signal according to a second upsampling parameter to obtain a third audio signal, the third audio signal being the audio signal to be subjected to secondary bandwidth expansion; inputting the third audio signal into a second super-resolution network to obtain a second super-resolution signal, the second super-resolution signal containing a third frequency component, the third frequency component including frequency components in a third frequency band greater than a second cutoff frequency, the second cutoff frequency being the cutoff frequency of the third audio signal, the third frequency band being determined based on the sampling rate of the third audio signal; determining a second bandwidth ratio based on the second upsampling parameter, the second bandwidth ratio indicating the position of the second cutoff frequency in the third frequency band; determining a second high-pass filter parameter based on the second bandwidth ratio; performing high-pass filtering on the second super-resolution signal according to the second high-pass filter parameter to obtain a fourth audio signal, the fourth audio signal containing the third frequency component; and superimposing the third audio signal and the fourth audio signal to obtain the audio signal with secondary bandwidth expansion.
[0025] It can be seen that the second bandwidth expansion process is similar to the first bandwidth expansion process. The difference lies in the fact that, in the scenario applied to audio decoding, the first bandwidth ratio in the first bandwidth expansion is determined by bandwidth detection, while the second bandwidth ratio in the second bandwidth expansion is determined based on the second upsampling parameter. That is, bandwidth detection is not required in the second bandwidth expansion.
[0026] Secondly, an audio bandwidth extension device is provided, which has the function of implementing the audio bandwidth extension method described in the first aspect above. The audio bandwidth extension device includes one or more modules for implementing the audio bandwidth extension method provided in the first aspect above.
[0027] That is, an audio bandwidth extension device is provided, the device comprising:
[0028] The first determining module is used to determine the first high-pass filter parameters based on the first bandwidth ratio. The first bandwidth ratio is used to indicate the position of the first cutoff frequency in the first frequency band. The first cutoff frequency is the cutoff frequency of the first audio signal to be bandwidth extended. The first audio signal contains a first frequency component. The first frequency component includes frequency components in the first frequency band that are not greater than the first cutoff frequency. The first frequency band is determined based on the sampling rate of the first audio signal.
[0029] The first super-resolution module is used to input the first audio signal into the first super-resolution network to obtain the first super-resolution signal. The first super-resolution signal includes a portion of the frequency components and the second frequency components in the first frequency components. The second frequency components include frequency components in the first frequency band that are greater than the first cutoff frequency.
[0030] The first high-pass filter module is used to perform high-pass filtering on the first super-division signal according to the first high-pass filter parameters to obtain the second audio signal, the second audio signal containing a second frequency component;
[0031] The first overlay module is used to overlay the first audio signal and the second audio signal to obtain a bandwidth-extended audio signal.
[0032] Optionally, the device further includes:
[0033] The decoding module is used to parse and dequantize the bitstream to obtain the MDCT spectrum;
[0034] A bandwidth detection module is used to perform bandwidth detection on the MDCT spectrum to obtain the first cutoff frequency;
[0035] The second determining module is used to divide the first cutoff frequency by the bandwidth of the first frequency band to obtain the first bandwidth ratio;
[0036] The third determining module is used to determine the first audio signal based on the MDCT spectrum.
[0037] Optionally, the third determining module includes:
[0038] The inverse transform submodule is used to perform an improved inverse cosine transform (IMDCT) on the MDCT spectrum to obtain the first audio signal.
[0039] The bandwidth-extended audio signal includes the i-th audio frame and the (i+1)-th audio frame, where i is an integer not less than 1; the device also includes:
[0040] The overlapping addition submodule is used to perform OLA on the bandwidth-extended audio signal and the bandwidth-extended reference audio signal to obtain the reconstructed signal of the i-th audio frame. The reference audio signal includes the (i-1)-th audio frame and the i-th audio frame. The reconstructed signal of the i-th audio frame contains a first frequency component and a second frequency component.
[0041] Optionally, the device further includes:
[0042] The first upsampling module is used to upsample the target audio signal according to the first upsampling parameters to obtain the first audio signal. The target audio signal refers to an audio signal that has frequency components in the second frequency band. The second frequency band is determined based on the sampling rate of the target audio signal.
[0043] The fourth determining module is used to determine the first bandwidth ratio based on the first upsampling parameter.
[0044] Optionally, the device further includes:
[0045] The second upsampling module is used to upsample the bandwidth-extended audio signal according to the second upsampling parameters to obtain a third audio signal, which is the audio signal to be subjected to a second bandwidth extension.
[0046] The second super-resolution module is used to input the third audio signal into the second super-resolution network to obtain the second super-resolution signal. The second super-resolution signal contains a third frequency component. The third frequency component includes frequency components in the third frequency band that are greater than the second cutoff frequency. The second cutoff frequency is the cutoff frequency of the third audio signal. The third frequency band is determined based on the sampling rate of the third audio signal.
[0047] The fifth determining module is used to determine the second bandwidth proportion based on the second upsampling parameters. The second bandwidth proportion is used to indicate the position of the second cutoff frequency in the third frequency band.
[0048] The sixth determining module is used to determine the second high-pass filter parameters based on the second bandwidth ratio;
[0049] The second high-pass filter module is used to perform high-pass filtering on the second super-division signal according to the second high-pass filter parameters to obtain the fourth audio signal, which contains the third frequency component.
[0050] The second overlay module is used to overlay the third audio signal and the fourth audio signal to obtain an audio signal with double bandwidth expansion.
[0051] Optionally, the first determining module includes:
[0052] The first determining submodule is used to determine the first bandwidth percentage range from multiple bandwidth percentage ranges;
[0053] The first selection submodule is used to select a set of high-pass filter parameters corresponding to the first bandwidth ratio range from multiple sets of high-pass filter parameters corresponding to the multiple bandwidth ratio ranges, and use it as the first high-pass filter parameter.
[0054] Optionally, the device further includes:
[0055] The selection module is used to select the super-resolution network corresponding to the sampling rate of the first audio signal from multiple super-resolution networks corresponding to multiple sampling rates, and use it as the first super-resolution network.
[0056] Optionally, the device further includes:
[0057] The acquisition module is used to acquire multiple audio sample sets. Each audio sample set includes multiple audio sample signals. The sampling rate of the audio sample signals in different audio sample sets is different, while the sampling rate of each audio sample signal in the same audio sample set is the same.
[0058] The seventh determination module is used to determine the multiple super-resolution networks based on the multiple sets of audio sample sets.
[0059] Optionally, the seventh determining module includes:
[0060] The training submodule is used to select one set of audio sample signals as the first audio sample set from the multiple sets of audio sample sets, and perform the following operations on the first audio sample set until the following operations are performed on all multiple sets of audio sample sets:
[0061] Noise is added to multiple audio sample signals in the first audio sample set to obtain multiple first sample signals;
[0062] According to one or more sets of low-pass filtering parameters, the multiple first sample signals are low-pass filtered to obtain multiple second sample signals;
[0063] The multiple second sample signals are used as the input to the initial super-resolution network, and the multiple first sample signals are used as the output of the initial super-resolution network. The initial super-resolution network is then trained to obtain a super-resolution network.
[0064] Optionally, the first super-resolution network includes a first processing module, a second processing module, and a third processing module. The second processing module includes a first convolutional submodule, a second convolutional submodule, and an addition submodule. The dilation rate of the convolutional layers in the first and second convolutional submodules is greater than 1.
[0065] The first super-resolution module is specifically used for:
[0066] The first audio signal is input into the first processing module to obtain the first data, which includes a first frequency component and a second frequency component.
[0067] The first data is input into the first convolutional submodule to obtain the second data;
[0068] The second data is input into the second convolutional submodule to obtain the third data;
[0069] The second and third data are input into the addition submodule to obtain the fourth data, which contains the second frequency component, or the fourth data contains a portion of the frequency component in the first frequency component and the second frequency component.
[0070] The fourth data is input into the third processing module to obtain the first super-resolution signal.
[0071] Optionally, the first processing module includes a nonlinear activation layer, and the second frequency component is a frequency component extended by the nonlinear activation layer.
[0072] Thirdly, an electronic device is provided, comprising a processor and a memory, the memory storing a program for executing the audio bandwidth extension method provided in the first aspect, and storing data related to implementing the audio bandwidth extension method provided in the first aspect. The processor is configured to execute the program stored in the memory. The electronic device may further include a communication bus for establishing a connection between the processor and the memory.
[0073] Fourthly, a computer-readable storage medium is provided, wherein instructions are stored therein, which, when executed on a computer, cause the computer to perform the audio bandwidth expansion method described in the first aspect.
[0074] Fifthly, a computer program product containing instructions is provided, which, when run on a computer, causes the computer to perform the audio bandwidth expansion method described in the first aspect above.
[0075] The technical effects achieved by the second, third, fourth, and fifth aspects mentioned above are similar to those achieved by the corresponding technical means in the first aspect, and will not be repeated here.
[0076] The technical solution provided in this application can bring at least the following beneficial effects:
[0077] Compared to bandwidth replication schemes, this scheme, through the expansion of high-frequency components via a super-resolution network, results in a more harmonious and natural sound in the final audio signal. Compared to schemes requiring envelope shaping of the MDCT spectrum, this scheme has lower complexity and improves bandwidth expansion efficiency. Furthermore, this scheme combines a super-resolution network with a high-pass filter; the high-pass filter parameters are determined based on the cutoff frequency of the audio signal to be bandwidth expanded. This demonstrates that the scheme can adapt to different cutoff frequencies, and the super-resolution network in this scheme can handle audio signals with various cutoff frequencies. When the audio signal to be bandwidth expanded is the audio signal obtained during the decoding process, the cutoff frequency is related to the encoding bitrate; the lower the bitrate, the smaller the cutoff frequency. In this case, this scheme can effectively adapt to different bitrates. Moreover, this scheme uses high-pass filtering to ensure that the low-frequency components of the final audio signal remain essentially unchanged, meaning that the low-frequency components are not damaged. Attached Figure Description
[0078] Figure 1 This is a schematic diagram of a Bluetooth interconnection scenario provided in an embodiment of this application;
[0079] Figure 2 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application;
[0080] Figure 3 This is a flowchart of an audio bandwidth extension method provided in an embodiment of this application;
[0081] Figure 4 This is a flowchart of an overlapping addition method provided in an embodiment of this application;
[0082] Figure 5 This is a schematic diagram of the structure of a first super-resolution network provided in an embodiment of this application;
[0083] Figure 6 This is a schematic diagram of the structure of a first processing module provided in an embodiment of this application;
[0084] Figure 7 This is a schematic diagram of the structure of a second processing module provided in an embodiment of this application;
[0085] Figure 8 This is a schematic diagram of another first super-resolution network provided in the embodiments of this application;
[0086] Figure 9 This is a flowchart of another audio bandwidth extension method provided in an embodiment of this application;
[0087] Figure 10 This is a flowchart of another method for extending audio bandwidth provided in an embodiment of this application;
[0088] Figure 11 This is a schematic diagram of the structure of an audio bandwidth expansion device provided in an embodiment of this application. Detailed Implementation
[0089] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the implementation methods of this application will be further described in detail below with reference to the accompanying drawings.
[0090] First, the implementation environment and background knowledge involved in the embodiments of this application will be introduced.
[0091] With the widespread adoption of true wireless stereo (TWS) earbuds, smart speakers, and smartwatches in daily life, the demand for high-quality audio playback experiences has become increasingly urgent, especially in environments where Bluetooth signals are susceptible to interference, such as subways, airports, and train stations. In Bluetooth interconnection scenarios, due to the limitations on data transmission size imposed by the Bluetooth channel connecting the audio transmitting and receiving devices, the audio signal must be compressed by the audio encoder in the transmitting device before being transmitted to the receiving device. The compressed audio signal is then decoded by the audio decoder in the receiving device before playback. Therefore, the proliferation of wireless Bluetooth devices has spurred the rapid development of various Bluetooth audio codecs.
[0092] Currently, Bluetooth audio codecs include sub-band coding (SBC), advanced audio coding (AAC), aptX series encoders, low-latency high-definition audio codec (LHDC), low-power low-latency LC3 audio codec, and LCSplus, etc.
[0093] Encoding can lead to the loss of high-frequency components in audio signals, resulting in reduced sound quality and a poorer listening experience for the decoded audio signal. Especially in low-bitrate scenarios, Bluetooth audio codecs reduce bandwidth to save bitrate, causing a significant loss of high-frequency components. Therefore, to improve sound quality, audio receiving devices need to extend the bandwidth of the audio signal to increase its high-frequency components. It is understood that the audio bandwidth extension method provided in this application can be applied to audio receiving devices in Bluetooth interconnect scenarios, i.e., the decoding end in Bluetooth interconnect scenarios.
[0094] Figure 1 This is a schematic diagram illustrating a Bluetooth interconnection scenario provided in an embodiment of this application. See also... Figure 1 In Bluetooth interconnection scenarios, audio transmitting devices can be mobile phones, computers, tablets, etc. Computers can be laptops, desktop computers, etc., and tablets can be handheld tablets, in-vehicle tablets, etc. Audio receiving devices in Bluetooth interconnection scenarios can be TWS earphones, smart speakers, wireless headphones, wireless neckband headphones, smartwatches, smart glasses, smart in-vehicle devices, etc. In other embodiments, the audio receiving device in Bluetooth interconnection scenarios can also be a mobile phone, computer, tablet, etc.
[0095] It should be noted that, in addition to its application in Bluetooth interconnection scenarios, this solution can also be applied to decoding in other device interconnection scenarios. In other words, this solution can be applied to the audio decoding process.
[0096] In addition to the narrow bandwidth of audio signals caused by audio encoding as mentioned above, some audio sources, after being sampled and processed by audio acquisition devices, inherently have narrow bandwidth. For example, using digital technology to transcribe old vinyl records into digital audio, or using low-sensitivity recording equipment to record music or speech. In such scenarios, the acquired raw audio signal lacks high-frequency components, resulting in poor sound quality. To improve sound quality, the audio bandwidth expansion method provided in this application embodiment can be used to increase high-frequency components. It is understood that the audio bandwidth expansion method provided in this application embodiment can be applied to any device to expand the bandwidth of narrow-bandwidth audio signals stored in these devices, thereby enhancing sound quality and improving the listening experience.
[0097] It should also be noted that the audio bandwidth expansion method provided in this application embodiment can expand the bandwidth of both speech signals and music signals.
[0098] Please refer to Figure 2 , Figure 2 This is a schematic diagram illustrating the structure of an electronic device according to an embodiment of this application. Optionally, the electronic device is... Figure 1 Any of the devices shown. The electronic device includes one or more processors 201, a communication bus 202, a memory 203, and one or more communication interfaces 204.
[0099] The processor 201 is a general-purpose central processing unit (CPU), a network processing unit (NP), a microprocessor, or one or more integrated circuits for implementing the solutions of this application, such as an application-specific integrated circuit (ASIC), a programmable logic device (PLD), or a combination thereof. Optionally, the PLD is a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.
[0100] The communication bus 202 is used to transmit information between the aforementioned components. Optionally, the communication bus 202 may be divided into an address bus, a data bus, a control bus, etc. For ease of illustration, only one thick line is used to represent it in the figure, but this does not mean that there is only one bus or one type of bus.
[0101] Optionally, the memory 203 may be a read-only memory (ROM), random access memory (RAM), electrically erasable programmable read-only memory (EEPROM), optical disc (including compact disc read-only memory (CD-ROM), compressed optical disc, laser disc, digital versatile optical disc, Blu-ray disc, etc.), magnetic disk storage medium, or other magnetic storage device, or any other medium capable of carrying or storing desired program code in the form of instructions or data structures that can be accessed by a computer, but not limited thereto. The memory 203 exists independently and is connected to the processor 201 via the communication bus 202, or the memory 203 is integrated with the processor 201.
[0102] Communication interface 204 uses any transceiver-like device for communicating with other devices or communication networks. Communication interface 204 includes a wired communication interface, and optionally, also includes a wireless communication interface. The wired communication interface is, for example, an Ethernet interface. Optionally, the Ethernet interface is an optical interface, an electrical interface, or a combination thereof. The wireless communication interface is a wireless local area network (WLAN) interface, a cellular network communication interface, or a combination thereof.
[0103] Optionally, in some embodiments, the electronic device includes multiple processors, such as Figure 2 The processors 201 and 205 shown are illustrated. Each of these processors is a single-core processor or a multi-core processor. Optionally, a processor here refers to one or more devices, circuits, and / or processing cores used for processing data (such as computer program instructions).
[0104] In a specific implementation, as one embodiment, the electronic device also includes an output device 206 and an input device 207. The output device 206 communicates with the processor 201 and can display information in various ways. For example, the output device 206 can be a liquid crystal display (LCD), a light-emitting diode (LED) display device, a cathode ray tube (CRT) display device, or a projector. The input device 207 communicates with the processor 201 and can receive user input in various ways. For example, the input device 207 can be a mouse, keyboard, touchscreen device, or sensing device.
[0105] In some embodiments, memory 203 stores program code 210 for executing the scheme of this application, and processor 201 is capable of executing the program code 210 stored in memory 203. The program code includes one or more software modules, and the electronic device can implement the following by means of processor 201 and program code 210 in memory 203. Figure 3 The embodiment provides a method for extending the bandwidth of audio.
[0106] Figure 3 This is a flowchart of an audio bandwidth extension method provided in an embodiment of this application. The method is applied to, for example... Figure 2 The electronic device shown. Please refer to... Figure 3 The method includes the following steps.
[0107] Step 301: Determine the first high-pass filter parameter based on the first bandwidth ratio. The first bandwidth ratio is used to indicate the position of the first cutoff frequency in the first frequency band. The first cutoff frequency is the cutoff frequency of the first audio signal whose bandwidth is to be extended. The first audio signal contains a first frequency component. The first frequency component includes frequency components in the first frequency band that are not greater than the first cutoff frequency. The first frequency band is determined based on the sampling rate of the first audio signal.
[0108] In order for this scheme to combine the super-resolution network in step 302 and the high-pass filter in step 303 to expand the bandwidth and adapt to different cutoff frequencies, the electronic device must determine the first high-pass filter parameters based on the first bandwidth ratio before executing step 303.
[0109] The first bandwidth percentage indicates the position of the cutoff frequency of the first audio signal within its frequency band. The cutoff frequency of the first audio signal is the first cutoff frequency, and the frequency band of the first audio signal is the first frequency band. The first frequency band is determined based on the sampling rate of the first audio signal. For example, the first frequency band and the sampling rate of the first audio signal satisfy a sampling theorem. This theorem can be the Nyquist sampling theorem, which states that the sampling rate of a signal should be no less than twice the bandwidth. In this embodiment, an example is given where the sampling rate of the signal is equal to twice the bandwidth; the bandwidth of the first frequency band is half the sampling rate of the first audio signal.
[0110] For example, suppose the sampling rate of the first audio signal is 44.1 kHz, the first frequency band is 0-22.05 kHz, and the bandwidth is 22.05 kHz. If the first cutoff frequency is 16 kHz, 16 divided by 22.05 is approximately 0.73, then the first bandwidth ratio is 0.73.
[0111] It should be noted that the audio signal in the embodiments of this application is a time-domain signal, such as a pulse code modulation (PCM) signal or other modulation type time-domain signal.
[0112] Optionally, one implementation of the electronic device determining the first high-pass filter parameter based on the first bandwidth proportion is as follows: The first bandwidth proportion range is determined from multiple bandwidth proportion ranges; from multiple sets of high-pass filter parameters corresponding to the multiple bandwidth proportion ranges, a set of high-pass filter parameters corresponding to the first bandwidth proportion range is selected as the first high-pass filter parameter. The electronic device stores the correspondence between the multiple bandwidth proportion ranges and the multiple sets of high-pass filter parameters, with each bandwidth proportion range and each set of high-pass filter parameters corresponding one-to-one, and the multiple bandwidth proportion ranges do not overlap.
[0113] For example, the multiple bandwidth percentage ranges include six ranges: [0, 0.4], (0.4, 0.6], (0.6, 0.7], (0.7, 0.85], (0.85, 0.95], and (0.95, 1). The multiple sets of high-pass filter parameters include six sets of parameters that correspond one-to-one with these six ranges. Assuming the first bandwidth percentage is 0.73, then the first bandwidth percentage range is the range (0.7, 0.85], and the first high-pass filter parameters are the set of parameters corresponding to the range (0.7, 0.85).
[0114] As mentioned above, this scheme can be applied to the decoding process, meaning the aforementioned electronic device can act as a decoding end. In this case, before determining the first high-pass filter parameters based on the first bandwidth ratio, the electronic device parses and dequantizes the bitstream to obtain the MDCT spectrum. The electronic device then performs bandwidth detection on the MDCT spectrum to obtain the first bandwidth ratio. Optionally, the electronic device performs bandwidth detection on the MDCT spectrum to obtain the first cutoff frequency, and divides the first cutoff frequency by the bandwidth of the first frequency band to obtain the first bandwidth ratio.
[0115] The bitstream is obtained by encoding the audio signal to be compressed at the encoding end, and the encoding process uses MDCT. The bandwidth of the audio signal to be compressed is greater than the bandwidth of the first audio signal. For example, the bandwidth of the audio signal to be compressed is equal to the bandwidth of the first frequency band, that is, the audio signal to be compressed is a full-bandwidth signal.
[0116] Furthermore, the bandwidth of an audio signal refers to its actual bandwidth, or effective bandwidth. The bandwidth detection mentioned above refers to effective bandwidth detection, and the bandwidth ratio refers to the proportion of the effective bandwidth to the corresponding frequency band. The bandwidth ratio can be called the effective bandwidth ratio. For example, the effective bandwidth of the first audio signal is the bandwidth between the minimum frequency and the first cutoff frequency of the first frequency band. By performing effective bandwidth detection on the MDCT spectrum, the first cutoff frequency can be obtained. The value of the first cutoff frequency is equal to the effective bandwidth of the first audio signal. Dividing the effective bandwidth of the first audio signal by the bandwidth of the first frequency band yields the first effective bandwidth ratio.
[0117] In the encoding process of this application embodiment, the encoding end divides the audio signal to be compressed into frames to obtain multiple audio frames of the signal to be compressed. The encoding end then applies windowing to each of these multiple audio frames of the signal to be compressed to obtain multiple frames of windowed signals to be compressed. The encoding end divides the j-th and j+1-th audio frames into an audio segment, and performs MDCT and quantization on the windowed signal to be compressed corresponding to this audio segment to obtain the encoded bitstream corresponding to that audio segment in the bitstream. Here, j is an integer not less than 0.
[0118] The frame length for framing the audio signal can be long or short; that is, the frame length of the audio frame is not limited in this embodiment. Optionally, the quantization parameters used for quantizing the audio segment are determined using the psychoacoustic masking effect, which can perceptually distinguish between important and secondary frequency information in the audio signal. Combining MDCT with the psychological masking effect can improve coding performance.
[0119] For example, the encoding end uses the 0th and 1st audio frames as the 0th audio segment, the 1st and 2nd audio frames as the 1st audio segment, the 2nd and 3rd audio frames as the 2nd audio segment, and so on, to obtain multiple audio segments. Assuming there are a total of M audio frames, the total number of these multiple audio segments is M-1.
[0120] Optionally, the 0th audio frame is a preset frame, which can be an all-zero signal or an all-one signal. The preset frame is not the actual audio signal; the audio frames starting from the 1st audio frame are the actual audio signals. The preset frame is set so that the reconstructed signal of the 1st audio frame can be obtained by overlapping and adding (OLA) the audio signals of the 0th and 1st audio segments during the subsequent decoding process.
[0121] After obtaining multiple audio segments, the encoder performs MDCT on the windowed signals to be compressed corresponding to the multiple audio segments according to the following formula (1) to obtain the MDCT spectrum to be quantized for each audio segment.
[0122]
[0123] Referring to formula (1), an audio frame includes N sampling points, an audio segment includes 2N sampling points, and the MDCT spectrum to be quantized corresponding to an audio segment includes N frequency points. n This refers to the windowed signal to be compressed, i.e., the windowed signal to be compressed. n X corresponds to the nth sampling point in an audio segment. k This represents the value of the k-th frequency point in the MDCT spectrum to be quantized corresponding to the audio segment. As can be seen from formula (1), MDCT will cause temporal aliasing in the audio signal.
[0124] The encoding end quantizes the MDCT spectra corresponding to the multiple audio segments according to the quantization parameters to obtain the quantized MDCT spectra corresponding to the multiple audio segments. Since MDCT spectra often have a cutoff frequency, i.e., the values of frequency points in the MDCT spectrum after the cutoff frequency are zero, the electronic device, after parsing and dequantizing the bitstream to obtain the MDCT spectrum, can sequentially traverse the frequency points in the MDCT spectrum from high frequency to low frequency. The value of the first non-zero frequency point encountered is the first cutoff frequency. The electronic device divides the first cutoff frequency by the bandwidth of the first frequency band to obtain the first bandwidth ratio. That is, in this embodiment, the electronic device can perform bandwidth detection on the MDCT spectrum by traversing from high frequency to low frequency. In other embodiments, the electronic device can also perform bandwidth detection on the MDCT spectrum in other ways, and this solution does not limit this method.
[0125] In addition, after parsing and dequantizing the bitstream to obtain the MDCT spectrum, the electronic device also determines the first audio signal based on the MDCT spectrum.
[0126] In one implementation, the first audio signal corresponds to a first audio segment, which is one of multiple audio segments obtained based on the audio signal to be compressed. The first audio segment corresponds to the i-th and (i+1)-th audio frames, where i is an integer not less than 1. The electronic device performs an improved inverse cosine transform (IMDCT) on the MDCT spectrum to obtain the first audio signal. The following formula (2) is the IMDCT formula, and the electronic device can perform IMDCT on the MDCT spectrum according to formula (2).
[0127]
[0128] Referring to formula (2), an audio frame includes N sampling points, and an audio segment includes 2N sampling points. k This represents the value of the k-th frequency point in the dequantized MDCT spectrum corresponding to an audio segment. The dequantized MDCT spectrum corresponding to an audio segment includes N frequency points. n y represents the value of the nth sample point in the first audio signal. n It is also a windowed signal.
[0129] It should be noted that the first audio signal obtained by the electronic device through IMDCT is a windowed signal.
[0130] In another implementation, the first audio signal corresponds to the i-th audio frame. The electronic device performs IMDCT on the MDCT spectrum to obtain a first windowed signal. The first windowed signal corresponds to a first audio segment. The electronic device performs OLA on the first windowed signal and the reference windowed signal to obtain the first reconstructed signal of the i-th audio frame, which is the first audio signal. The reference windowed signal corresponds to a reference audio segment, which corresponds to the (i-1)-th and i-th audio frames. The reference windowed signal is the signal obtained by the electronic device after performing IMDCT on the MDCT spectrum corresponding to the reference audio segment.
[0131] It should be noted that in this implementation, the first audio signal obtained by the electronic device through IMDCT and OLA is the windowed signal, i.e., the signal without a window. The electronic device can then reconstruct the final audio signal through step 304. That is, the bandwidth-extended audio signal obtained in step 304 is the reconstructed signal with improved sound quality, and there is no need to perform OLA again. In the previous implementation, the bandwidth-extended audio signal obtained by the electronic device through step 304 is not the final reconstructed audio signal, and OLA still needs to be performed to obtain the reconstructed audio signal.
[0132] Figure 4 This is a flowchart illustrating an overlapping addition method provided in an embodiment of this application. See also... Figure 4 During the decoding process, the electronic device sequentially obtains the first windowed signals corresponding to multiple audio segments using IMDCT. These audio segments include the 0th, 1st, 2nd, and so on. Specifically, the 0th audio segment corresponds to the 0th and 1st audio frames, the 1st audio segment corresponds to the 1st and 2nd audio frames, and so on, with the i-th audio segment corresponding to the i-th and (i+1)-th audio frames. After obtaining the first windowed signal corresponding to the 1st audio segment, the electronic device adds the first windowed signal corresponding to the 1st audio frame in the 0th audio segment to the first windowed signal corresponding to the 1st audio frame in the 1st audio segment, to obtain the first reconstructed signal of the 1st audio frame. Similarly, after obtaining the first windowed signal corresponding to the i-th audio segment, the electronic device adds the first windowed signal corresponding to the i-th audio frame in the (i-1)-th audio segment to the first windowed signal corresponding to the i-th audio frame in the i-th audio segment, to obtain the first reconstructed signal of the i-th audio frame. Figure 4 The windowed signal in the text refers to the first windowed signal, and the reconstructed signal refers to the first reconstructed signal.
[0133] It should be noted that OLA introduces a one-frame delay during the decoding process. Figure 4 It can also be seen that, when the first windowed signal of the (i+1)th audio frame is obtained based on the bitstream, the electronic device can reconstruct the reconstructed signal of the ith audio frame, but cannot reconstruct the reconstructed signal of the (i+1)th audio frame. In other words, when the first windowed signal of the current frame is obtained, the electronic device can reconstruct the reconstructed signal of the previous frame.
[0134] As mentioned above, this solution can also be applied to any device to extend the bandwidth of audio signals stored in these devices that have narrow bandwidth. The following section will describe how to determine the first bandwidth ratio and how to obtain the first audio signal in this scenario.
[0135] In this embodiment, the electronic device upsamples the target audio signal according to a first upsampling parameter to obtain a first audio signal. The target audio signal refers to an audio signal that has frequency components throughout a second frequency band, which is determined based on the sampling rate of the target audio signal. For example, the sampling rate of the second frequency band and the target audio signal satisfy a sampling theorem, such as the Nyquist sampling theorem. Furthermore, the electronic device determines a first bandwidth proportion based on the first upsampling parameter. In this embodiment, the electronic device determines the first bandwidth proportion as the reciprocal of the first upsampling parameter.
[0136] It should be noted that the upsampling in this paper refers to time-domain upsampling. The target audio signal is a full-bandwidth signal whose frequency components fill the second frequency band, but its bandwidth is narrow and its sound quality is poor. The bandwidth of the first audio signal obtained through upsampling is basically the same as that of the target audio signal. However, the frequency components of the first audio signal do not fill the first frequency band, meaning it is not a full-bandwidth signal, and its bandwidth is also narrow. The first audio signal can be used as the audio signal to be bandwidth extended. This scheme is used to extend the bandwidth of the first audio signal to obtain a full-bandwidth signal with higher sound quality whose frequency components fill the first frequency band.
[0137] For example, the target audio signal has a sampling rate of 2F kHz, a second frequency band of 0-F kHz, and a bandwidth of F kHz, meaning the target audio signal is a full-bandwidth signal. Assuming the first upsampling parameter is 'a', the sampling rate of the first audio signal obtained by upsampling the target audio signal by the electronic device is a*2F kHz, the bandwidth of the first audio signal is a*F kHz, and the first bandwidth percentage is 1 / a. For example, if F = 22.05, the target audio signal has a sampling rate of 44.1kHz and a bandwidth of 22.05kHz, corresponding to a second frequency band bandwidth of 22.05kHz. If a = 2, indicating a doubling of the upsampling value, the first audio signal has a sampling rate of 88.2kHz and a bandwidth of 22.05kHz, corresponding to a first frequency band bandwidth of 44.1kHz, and a first bandwidth percentage of 0.5. This example also shows that the cutoff frequency of the first audio signal is equal to its bandwidth.
[0138] It should be noted that the first audio signal in this scenario is a windowless signal.
[0139] Step 302: Input the first audio signal into the first super-resolution network to obtain the first super-resolution signal. The first super-resolution signal includes a portion of the frequency components in the first frequency component and the second frequency component. The second frequency component includes the frequency components in the first frequency band that are greater than the first cutoff frequency.
[0140] In this embodiment, the electronic device inputs a first audio signal to be bandwidth extended into a first super-resolution network to obtain a first super-resolution signal. If the first frequency component contained in the first audio signal is considered a low-frequency component, then the second frequency component contained in the first super-resolution signal is the extended high-frequency component. That is, the first super-resolution network is used to extend the high-frequency component.
[0141] The first super-resolution network is a neural network, or a deep learning network. This application does not limit the network structure, network parameters, training method, etc., of the first super-resolution network. The following is an exemplary description of the network structure and data processing procedure of a first super-resolution network provided in this application embodiment.
[0142] In one implementation, the first super-resolution network includes a first processing module, a second processing module, and a third processing module. The second processing module includes a first convolutional submodule, a second convolutional submodule, and an addition submodule. The dilation rate of the convolutional layers in the first and second convolutional submodules is greater than 1.
[0143] The data processing procedure for an electronic device to input a first audio signal into a first super-resolution network to obtain a first super-resolution signal includes: inputting the first audio signal into a first processing module to obtain first data, the first data containing a first frequency component and a second frequency component; inputting the first data into a first convolution submodule to obtain second data; inputting the second data into a second convolution submodule to obtain third data; inputting the second data and the third data into an addition submodule to obtain fourth data, the fourth data containing a second frequency component, or the fourth data containing a portion of the first frequency component and a second frequency component; and inputting the fourth data into a third processing module to obtain the first super-resolution signal.
[0144] Based on the above discussion, we can conclude that... Figure 5 The diagram shows a first super-resolution network. It should be noted that the third processing module may include an output layer. Alternatively, the third processing module may also include other processing layers. For example, the third processing module may have the same structure as the first processing module, or the third processing module may have the same structure as the second processing module, or the third processing module may include parts with the same structure as the first processing module and parts with the same structure as the second processing module, or the third processing module may be a module with other structures.
[0145] It should be understood that the first processing mode is used to expand high-frequency components (such as the second frequency component). In the embodiments of this application, the first processing module includes a nonlinear activation layer, and the second frequency component is a frequency component expanded by the nonlinear activation layer. The nonlinear activation layer has a nonlinear filtering function, and the nonlinear filtering can produce a frequency doubling effect, enabling the first processing module to expand high-frequency components. While expanding high-frequency components, the first processing module may also expand some low-frequency components (such as a portion of the first frequency component). The second processing module is used to minimize the expanded low-frequency components. Specifically, since the dilation rate of the convolutional layers in the second processing module is greater than 1, these convolutional layers implicitly have a downsampling function. Combined with the function of the summing submodule, the second processing module can reduce the expanded low-frequency components.
[0146] Optionally, the nonlinear activation layer in the first processing module may employ a sinusoidal activation function or other nonlinear activation functions.
[0147] In this embodiment, the first processing module is a convolutional module. Besides a non-linear activation function, the first processing module also includes convolutional layers, and the dilation rate of the convolutional layers in the first processing module is 1. Optionally, the convolutional kernel of the convolutional layer in the first processing module is 3 or other values, and the number of channels of the convolutional layer in the first processing module is 2 or other values. In this embodiment, the maximum number of channels in the convolutional layer of the first processing module is 2.
[0148] Figure 6 This is a schematic diagram of the structure of a first processing module provided in an embodiment of this application. Figure 6 In this diagram, the WaveConv module represents the first processing module, Conv1d represents a one-dimensional convolutional layer, d=1 indicates that the dilation rate of this convolutional layer is 1, and Sin represents a non-linear activation layer using a sinusoidal activation function. The input data of this WaveConv module is the first audio signal, and the output data is the first data.
[0149] Optionally, the dilation rate of the convolutional layers in the first and second convolutional submodules is 2 or 3, the convolutional kernel is 3 or other values, and the number of channels is 2 or other values. Optionally, the first and second convolutional submodules also include activation layers, such as nonlinear activation layers or other types of activation layers. In this embodiment, the example of a first super-resolution network where all activation layers are nonlinear activation layers is used for illustration.
[0150] Figure 7 This is a schematic diagram of the structure of a second processing module provided in an embodiment of this application. Figure 7 In the diagram, ResConv represents the second processing module, WaveConv21 represents the first convolutional submodule, and WaveConv22 represents the second convolutional submodule. This represents the addition submodule. The specific structures of WaveConv21 and WaveConv22 are similar to... Figure 6 The WaveConv modules shown have the same structure, but the dilation rate of the convolutional layers in WaveConv21 and WaveConv22 is 1, i.e., d = 1. The input data of this ResConv module is the first data, and the output data is the fourth data.
[0151] Figure 8 This is a schematic diagram of another first super-resolution network provided in an embodiment of this application. See also... Figure 8 The first super-resolution network includes WaveConv1, ResConv1, ResConv2, and WaveConv2. WaveConv1 is the first processing module, ResConv1 is the second processing module, and ResConv2 and WaveConv2 are sub-modules within the third processing module. The specific structures of ResConv1 and ResConv2 are similar to those of... Figure 7 The ResConv module shown has the same structure, and the specific structures of WaveConv1 and WaveConv2 are also the same. Figure 6 The WaveConv modules shown have the same structure. Figure 8 In this context, d=1 indicates that the dilation rate of the corresponding convolutional layer is 1, and d=2 indicates that the dilation rate of the corresponding convolutional layer is 2.
[0152] It can be seen that, Figure 8 The first super-resolution network shown includes some convolutional layers and nonlinear activation layers. Its network structure is relatively simple, with few parameters and low computational complexity. This scheme uses... Figure 8 The first super-resolution network shown can be used to expand bandwidth quickly and efficiently to expand high-frequency components.
[0153] In this embodiment, the first super-resolution network can be used to extend the bandwidth of audio signals with multiple sampling rates, including a first sampling rate equal to the sampling rate of the first audio signal. For example, the first super-resolution network is a network trained with audio sample signals of multiple sampling rates. Alternatively, the first super-resolution network is dedicated to extending the bandwidth of audio signals with the first sampling rate. For example, the first super-resolution network is a network trained with audio sample signals of the first sampling rate. When the first super-resolution network is dedicated to processing audio signals with the first sampling rate, the performance of the first super-resolution network is better. The following will describe the first super-resolution network dedicated to processing audio signals with the first sampling rate as an example.
[0154] Before inputting the first audio signal into the first super-resolution network to obtain the first super-resolution signal, the electronic device selects the super-resolution network corresponding to the sampling rate of the first audio signal from among multiple super-resolution networks corresponding to multiple sampling rates, and uses it as the first super-resolution network. It should be understood that the electronic device deploys multiple super-resolution networks, each corresponding one-to-one with a sampling rate; that is, one super-resolution network corresponds to one sampling rate, and different super-resolution networks correspond to different sampling rates. Each super-resolution network is used to extend the bandwidth of the audio signal at its corresponding sampling rate.
[0155] It should be noted that these multiple super-resolution networks are pre-trained networks. The training process can be performed on a server or on other devices, and this application embodiment does not limit this. The following describes the training process of these multiple super-resolution networks by taking the training process performed on an electronic device as an example.
[0156] An electronic device acquires multiple audio sample sets, each containing multiple audio sample signals. The sampling rates of the audio sample signals in different audio sample sets are different, while the sampling rates of all audio sample signals within the same audio sample set are the same. Based on these multiple audio sample sets, the electronic device determines multiple super-resolution networks. Specifically, each audio sample set corresponds one-to-one with a super-resolution network; one audio sample set is used to determine one super-resolution network, and different audio sample sets are used to determine different super-resolution networks.
[0157] Optionally, the electronic device selects one set of audio sample signals from the plurality of audio sample sets as the first audio sample set, and performs the following operation on the first audio sample set until the following operation is performed on all the plurality of audio sample sets:
[0158] Noise is added to multiple audio sample signals in the first audio sample set to obtain multiple first sample signals; the multiple first sample signals are low-pass filtered according to one or more sets of low-pass filtering parameters to obtain multiple second sample signals; the multiple second sample signals are used as the input of the initial super-resolution network, and the multiple first sample signals are used as the output of the initial super-resolution network to train the initial super-resolution network to obtain a super-resolution network.
[0159] Adding noise to an audio sample signal yields at least one first sample signal. For example, adding M noises to an audio sample signal yields M first sample signals, where M is an integer greater than or equal to 1. Optionally, when M is greater than 1, the amplitudes and / or types of these M noise signals are different. It should be noted that when M equals 1, the number of the multiple audio sample signals can be the same as the number of the multiple first sample signals. When M is greater than 1, the number of the multiple audio sample signals is greater than the number of the multiple first sample signals. It should also be noted that the number of noises added to different audio sample signals can be the same or different. For example, adding 2 noises to one audio sample signal yields 2 first sample signals, and adding 3 noises to another audio signal yields another 3 first sample signals. It should be noted that by adding noise to the audio sample signals, the trained super-resolution network becomes sensitive to noise signals and exhibits high gain for noise-like components in the audio signal; that is, the trained super-resolution network can also have a bandwidth expansion effect for noise-like components in the audio signal.
[0160] Furthermore, the electronic device performs low-pass filtering on the multiple first sample signals according to a set of low-pass filtering parameters. Alternatively, the electronic device groups the multiple first sample signals into multiple sets of first sample signals, each set corresponding one-to-one with a set of low-pass filtering parameters, with each set of parameters used to perform low-pass filtering on its corresponding set. The electronic device then performs low-pass filtering on each of the multiple sets of first sample signals according to these multiple sets of low-pass filtering parameters to obtain multiple sets of second sample signals. That is, the electronic device performs low-pass filtering on each set of first sample signals according to the low-pass filtering parameters in the multiple sets of parameters. The cutoff frequencies of the second sample signals in different sets of these multiple sets of second sample signals are different, which is more conducive to training a super-resolution network that can adapt to different cutoff frequencies.
[0161] The aforementioned multiple audio sample sets can include music sample signals and / or speech sample signals of different styles. For example, each audio sample set includes music sample signals of one or more styles, as well as speech signals of one or more individuals. This facilitates the training of a super-resolution network capable of processing both music and speech signals. Different music styles can refer to different instruments, different musical genres, etc. Different instruments have different frequency ranges, and different musical genres also have different frequency ranges. Different speech styles can refer to different timbre, different pitches, etc. Generally speaking, music signals have a larger frequency range and may have greater frequency fluctuations, while speech signals have a relatively smaller frequency range and relatively smaller frequency fluctuations. The characteristics of speech signals are relatively easier to learn through artificial intelligence (AI) networks. Research and applications of AI audio super-resolution in related technologies mainly focus on speech signals, with little research on music signals. Therefore, the embodiments of this application provide a method for bandwidth expansion of music and speech signals, which has good robustness for both speech and music signals.
[0162] As described above, the input and output data used to train the initial super-resolution network are the second sample signal and the first sample signal, respectively. Since the second sample signal is obtained by low-pass filtering the first sample signal, the spectral structure of the low-frequency components in the second sample signal is consistent with the spectral structure of the high-frequency components in the first sample signal. After training the initial super-resolution network using the second and first sample signals with good spectral structure consistency, the resulting super-resolution network possesses the property of expanding to produce high-frequency components consistent with the spectral structure of the low-frequency components of the input signal. In other words, after inputting the first audio signal into the trained first super-resolution network, the spectral structure of the high-frequency components expanded by the super-resolution network has good structural consistency with the spectral structure of the frequency components of the first audio signal; that is, the high-frequency components contained in the first super-resolution signal and the low-frequency components contained in the first audio signal have a high degree of consistency in spectral structure.
[0163] Step 303: Perform high-pass filtering on the first super-division signal according to the first high-pass filtering parameters to obtain the second audio signal, which contains the second frequency component.
[0164] In this embodiment, after obtaining the first super-resolution signal through the first super-resolution network, the electronic device performs high-pass filtering on the first super-resolution signal according to the first high-pass filtering parameters to minimize the extended low-frequency components. In other words, high-pass filtering is used to minimize the number of first frequency components in the second audio signal.
[0165] It should be noted that when adopting, such as Figure 8 In the case of the first super-resolution network shown, the first super-resolution network has reduced the extended low-frequency components as much as possible. In order to further reduce the extended low-frequency components, the electronic device can also perform high-pass filtering on the first super-resolution signal to further reduce the extended low-frequency components.
[0166] Step 304: Superimpose the first audio signal and the second audio signal to obtain a bandwidth-extended audio signal.
[0167] In this embodiment, after receiving a second audio signal containing a second frequency component, the electronic device superimposes the first audio signal and the second audio signal to obtain a bandwidth-extended audio signal. The bandwidth-extended audio signal contains both the first and second frequency components; that is, the bandwidth-extended audio signal is a full-bandwidth signal.
[0168] It should be noted that since the extended high-frequency components and the low-frequency components contained in the first audio signal have a high degree of consistency in spectral structure, the spectral structure of the bandwidth-extended audio signal obtained after superimposing the first audio signal and the second audio signal also has a high degree of consistency with the spectral structure of the first audio signal. As a result, the final audio sounds very natural and the quality is well improved.
[0169] As mentioned above, in one implementation, the first audio signal is obtained by the electronic device through IMDCT during the decoding process, and the first audio signal is a windowed signal. In this implementation, the bandwidth-extended audio signal obtained by the electronic device in step 304 is also a windowed signal. The electronic device needs to process the bandwidth-extended audio signal to obtain the reconstructed audio signal.
[0170] In this implementation, the electronic device superimposes the first audio signal and the second audio signal to obtain a bandwidth-extended audio signal. Then, it interleaves and adds the bandwidth-extended audio signal and the bandwidth-extended reference audio signal to obtain the reconstructed signal of the i-th audio frame. The bandwidth-extended audio signal includes the i-th and (i+1)-th audio frames, and the reference audio signal includes the (i-1)-th and i-th audio frames. The reconstructed signal of the i-th audio frame contains a first frequency component and a second frequency component, where i is an integer not less than 1. If the time when the bandwidth-extended audio signal is obtained is the current time, then the time when the bandwidth-extended reference audio signal is obtained is a historical time. Therefore, the reference audio signal can also be called a historical audio signal.
[0171] It should be noted that the OLA in the decoding process can not only remove windows, but also align inter-frame samples to make the inter-frame transition smooth. OLA will generate a system delay of one frame. After OLA, the reconstructed signal is obtained and no additional delay is introduced.
[0172] In another implementation, the first audio signal is obtained by the electronic device during decoding using IMDCT and OLA, and this first audio signal is a windowless signal. In this implementation, the bandwidth-extended audio signal obtained by the electronic device in step 304 is also a windowless signal. However, since bandwidth extension occurs after OLA, OLA has already introduced a one-frame system delay. Furthermore, the super-resolution network's processing of the first audio signal causes misalignment of samples between frames, introducing perceived noise. To eliminate this perceived noise, the electronic device needs to perform inter-frame smoothing on the bandwidth-extended audio signal, which introduces additional delay. Simply put, in this implementation, OLA introduces a one-frame delay, and subsequent bandwidth extension and inter-frame smoothing introduce additional delay.
[0173] In another implementation, the first audio signal is obtained by upsampling the full-band target audio signal by the electronic device. The first audio signal is a windowless signal. The bandwidth-extended audio signal obtained by the electronic device in step 304 is a playable, higher-quality audio signal.
[0174] Figure 9 This is a flowchart of another audio bandwidth extension method provided in an embodiment of this application. See also... Figure 9 This method, applied in the audio decoding process, is a bandwidth-adaptive super-resolution method, which includes the following steps 901 to 907.
[0175] Step 901: Unpacking and Dequantization. That is, during the decoding process, the electronic device first parses and dequantizes the bitstream to obtain the MDCT spectrum of each audio segment in multiple audio segments.
[0176] Step 902: Bandwidth detection to determine high-pass filter parameters. That is, the electronic device performs bandwidth detection on the MDCT spectrum of each audio segment to determine the first bandwidth proportion corresponding to each audio segment. Based on the first bandwidth proportion corresponding to each audio segment, the electronic device determines the first high-pass filter parameters corresponding to each audio segment.
[0177] Step 903: IMDCT. That is, the electronic device performs IMDCT on the MDCT spectrum of each audio segment to obtain the first audio signal corresponding to each audio segment to be bandwidth-extended. The first audio signal contains low-frequency components and is a windowed signal.
[0178] Step 904: AI Audio Super-Resolution. That is, the electronic device inputs the first audio signal corresponding to each audio segment into the first super-resolution network to obtain the first super-resolution signal corresponding to each audio segment. The first super-resolution signal contains the extended high-frequency components and also contains some extended low-frequency components.
[0179] Step 905: High-pass filtering. That is, the electronic device performs high-pass filtering on the first super-resolution signal corresponding to each audio segment according to the first high-pass filtering parameters corresponding to each audio segment, to obtain the second audio signal corresponding to each audio segment. The second audio signal contains the extended high-frequency components.
[0180] Step 906: High and low frequency signal superposition. That is, the electronic device superimposes the first audio signal and the second audio signal corresponding to each audio segment to obtain the bandwidth-expanded audio signal corresponding to each audio segment. The bandwidth-expanded audio signal is a full-bandwidth windowed signal.
[0181] Step 907: Frame Overlapping and Addition. That is, the electronic device performs OLA on the bandwidth-extended audio signals corresponding to the multiple audio segments to obtain the reconstructed signals of the multiple audio frames, and outputs the waveforms of the reconstructed signals of the multiple audio frames.
[0182] Optionally, the electronic device can further improve audio quality through secondary bandwidth expansion. In one implementation, the electronic device superimposes a first audio signal and a second audio signal to obtain a bandwidth-expanded audio signal, and then upsamples the bandwidth-expanded audio signal according to a second upsampling parameter to obtain a third audio signal, which is the audio signal to be subjected to secondary bandwidth expansion. The electronic device inputs the third audio signal into a second super-resolution network to obtain a second super-resolution signal, which contains a third frequency component. The third frequency component includes frequency components in a third frequency band that are higher than a second cutoff frequency. The second cutoff frequency is the cutoff frequency of the third audio signal, and the third frequency band is determined based on the sampling rate of the third audio signal. For example, the sampling rate of the third frequency band and the third audio signal satisfy the sampling theorem. The electronic device determines a second bandwidth ratio based on the second upsampling parameter, which indicates the position of the second cutoff frequency in the third frequency band. The electronic device determines a second high-pass filter parameter based on the second bandwidth ratio and performs high-pass filtering on the second super-resolution signal according to the second high-pass filter parameter to obtain a fourth audio signal, which contains the third frequency component. The electronic device superimposes the third audio signal and the fourth audio signal to obtain an audio signal with double bandwidth expansion.
[0183] The network structure of the second super-resolution network can be the same as or different from that of the first super-resolution network. The network parameters of the second super-resolution network can also be the same as or different from those of the first super-resolution network. Optionally, the second super-resolution network is one of the aforementioned super-resolution networks that differs from the first super-resolution network, and the sampling rate of the second super-resolution network corresponds to that of the third audio signal.
[0184] As can be seen, the second bandwidth expansion process is similar to the first bandwidth expansion process. The difference lies in the fact that, in the scenario where bandwidth expansion is performed during decoding, the first bandwidth proportion in the first bandwidth expansion is determined by bandwidth detection, while the second bandwidth proportion in the second bandwidth expansion is determined based on the second upsampling parameter; that is, bandwidth detection is not required in the second bandwidth expansion. The specific implementation method of the second bandwidth expansion can be referred to the relevant description in the aforementioned embodiments, and will not be repeated here.
[0185] It should be noted that if the first audio signal is obtained by the electronic device through IMDCT during the decoding process, then after obtaining the double-bandwidth-extended audio signal, the electronic device needs to obtain the reconstructed audio signal through OLA based on the double-bandwidth-extended audio signal. The specific implementation is similar to the above description of the electronic device overlapping and adding the bandwidth-extended audio signal and the bandwidth-extended reference audio signal to obtain the reconstructed signal of the i-th audio frame, and will not be repeated here. If the first audio signal is obtained by the electronic device through IMDCT and OLA during the decoding process, then the electronic device also needs to obtain the reconstructed audio signal through inter-frame smoothing based on the double-bandwidth-extended audio signal. If the first audio signal is obtained by the electronic device after upsampling the full-bandwidth target audio signal, then the double-bandwidth-extended audio signal obtained by the electronic device is a higher-quality audio signal that can be played.
[0186] In an exemplary embodiment, the original audio signal has a sampling rate of 48kHz, corresponding to a first frequency band of 24kHz and a bandwidth of 24kHz. After encoding the original audio signal, the bandwidth is reduced to 16kHz. The decoding end parses, dequantizes, and performs IMDCT on the bitstream to obtain the audio signal to be bandwidth-extended. Bandwidth detection determines that the bandwidth of the audio signal to be bandwidth-extended is 16kHz. After the decoding end performs the first bandwidth extension on the audio signal using this scheme, the bandwidth of the audio signal can be increased to 24kHz; that is, the first bandwidth extension can increase the incomplete frequency band audio signal to a full frequency band signal. The decoding end then performs a second bandwidth extension on the bandwidth-extended audio signal using this scheme, increasing the bandwidth of the audio signal to 48kHz and the sampling rate to 96kHz.
[0187] Figure 10 This is a flowchart illustrating yet another method for extending audio bandwidth according to an embodiment of this application. See also... Figure 10 This method, applied in the audio decoding process, is a sampling rate extended super-resolution method, comprising steps 1001 to 1007. It should be noted that... Figure 10 The process shown includes two bandwidth expansion processes. The first bandwidth expansion process includes steps 1001 to 1006, and the second bandwidth expansion process includes steps 1007 to 1012.
[0188] Steps 1001 to 1006 and Figure 9 Steps 901 to 906 in the process shown are the same, and will not be repeated here.
[0189] Step 1007: Upsampling. That is, after obtaining the bandwidth-expanded audio signals corresponding to each audio segment through the first bandwidth expansion, the electronic device upsamples the bandwidth-expanded audio signals corresponding to each audio segment according to the second upsampling parameters to obtain the third audio signal corresponding to each audio segment to be subjected to a second bandwidth expansion. The third audio signal is the bandwidth signal.
[0190] Step 1008: The electronic device determines the parameters of the second high-pass filter based on the first upsampling parameters.
[0191] Step 1009: AI Audio Super-Resolution. That is, the electronic device inputs the third audio signal corresponding to each audio segment into the second super-resolution network to obtain the second super-resolution signal corresponding to each audio segment.
[0192] Step 1010: High-pass filtering. That is, the electronic device performs high-pass filtering on the second super-resolution signal corresponding to each audio segment according to the second high-pass filtering parameters corresponding to each audio segment, so as to obtain the fourth audio signal corresponding to each audio segment.
[0193] Step 1011: High and low frequency signal superposition. The electronic device superimposes the third and fourth audio signals corresponding to each audio segment to obtain the audio signal corresponding to each audio segment after double bandwidth expansion. The audio signal after double bandwidth expansion is a windowed signal.
[0194] Step 1012: Frame Overlapping and Addition. That is, the electronic device performs OLA on the audio signals corresponding to the multiple audio segments that have undergone secondary bandwidth expansion to obtain the reconstructed signals of multiple audio frames, and outputs the waveforms of the reconstructed signals of the multiple audio frames.
[0195] As can be seen from the above, electronic devices first extend the bandwidth to super-resolution the audio signal to the full frequency band, and then further enhance the audio signal to an ultra-high-definition signal through upsampling and super-resolution networks. After the second bandwidth extension, the sound field of the audio signal will be further extended, and the listening experience will be greatly improved.
[0196] Alternatively, electronic devices can further enhance audio quality to meet demand through multiple bandwidth expansions.
[0197] Optionally, the electronic device can perform bandwidth expansion in some cases and not in others to reduce the waste of computing power and memory resources. For example, after determining the first bandwidth percentage, if the first bandwidth percentage is less than a threshold, it indicates that the bandwidth of the first audio signal is narrow and bandwidth expansion is needed; in this case, the electronic device executes steps 301 to 304. If the first bandwidth percentage is greater than or equal to the threshold, it indicates that the bandwidth of the first audio signal is wide and bandwidth expansion is not needed; in this case, the electronic device does not execute steps 301 to 304. Alternatively, if the first audio signal is a first-type signal, such as a voice signal or music signal, the electronic device executes steps 301 to 304. If the first audio signal is a second-type signal, such as street sound, the electronic device does not execute steps 301 to 304. Here, the first-type signal can refer to a subjective signal, and the second-type signal can refer to an objective signal. Alternatively, the electronic device can also determine whether bandwidth expansion is needed based on other conditions. The above-mentioned judgment conditions can be used individually or in combination.
[0198] In summary, compared to bandwidth replication schemes, this scheme, through the expansion of high-frequency components via a super-resolution network, results in a more harmonious and natural sound for the final audio signal. Compared to schemes requiring envelope shaping of the MDCT spectrum, this scheme has lower complexity, improves bandwidth expansion efficiency, and can be applied to devices with lower computing power. Furthermore, this scheme combines the super-resolution network with a high-pass filter; the high-pass filter parameters are determined based on the cutoff frequency of the audio signal to be bandwidth expanded. Therefore, this scheme can adapt to different cutoff frequencies, and the super-resolution network in this scheme can handle audio signals with various cutoff frequencies.
[0199] Furthermore, this solution can be applied to audio codecs, such as in the decoding process of Bluetooth audio codecs, to receive the audio bitstream in real time and improve audio quality. When the audio signal to be bandwidth extended is the audio signal obtained during the decoding process, since the cutoff frequency is related to the encoding bitrate (the lower the bitrate, the smaller the cutoff frequency), this solution can essentially adapt to different bitrates. In addition, this solution uses high-pass filtering to ensure that the low-frequency components of the final audio signal remain essentially unchanged, i.e., the low-frequency components are not damaged. This solution can also be designed as a standalone program for audio source enhancement and applied to any device.
[0200] Figure 11 This is a schematic diagram of the structure of an audio bandwidth extension device 1100 provided in an embodiment of this application. The audio bandwidth extension device 1100 can be implemented by software, hardware, or a combination of both as part or all of an electronic device. This electronic device can be... Figure 2 The electronic device shown. See also Figure 11 The device 1100 includes: a first determining module 1101, a first super-resolution module 1102, a first high-pass filtering module 1103, and a first superposition module 1104.
[0201] The first determining module 1101 is used to determine the first high-pass filter parameters according to the first bandwidth ratio. The first bandwidth ratio is used to indicate the position of the first cutoff frequency in the first frequency band. The first cutoff frequency is the cutoff frequency of the first audio signal to be bandwidth extended. The first audio signal contains a first frequency component. The first frequency component includes frequency components in the first frequency band that are not greater than the first cutoff frequency. The first frequency band is determined based on the sampling rate of the first audio signal.
[0202] The first super-resolution module 1102 is used to input the first audio signal into the first super-resolution network to obtain the first super-resolution signal. The first super-resolution signal includes a portion of the frequency components and the second frequency components in the first frequency components. The second frequency components include frequency components in the first frequency band that are greater than the first cutoff frequency.
[0203] The first high-pass filter module 1103 is used to perform high-pass filtering on the first super-division signal according to the first high-pass filter parameters to obtain the second audio signal, the second audio signal containing the second frequency component;
[0204] The first superposition module 1104 is used to superimpose the first audio signal and the second audio signal to obtain a bandwidth-extended audio signal.
[0205] Optionally, the device 1100 further includes:
[0206] The decoding module is used to parse and dequantize the bitstream to obtain the MDCT spectrum;
[0207] A bandwidth detection module is used to perform bandwidth detection on the MDCT spectrum to obtain the first cutoff frequency;
[0208] The second determining module is used to divide the first cutoff frequency by the bandwidth of the first frequency band to obtain the first bandwidth ratio;
[0209] The third determining module is used to determine the first audio signal based on the MDCT spectrum.
[0210] Optionally, the third determining module includes:
[0211] The inverse transform submodule is used to perform an improved inverse cosine transform (IMDCT) on the MDCT spectrum to obtain the first audio signal.
[0212] The bandwidth-extended audio signal includes the i-th audio frame and the (i+1)-th audio frame, where i is an integer not less than 1;
[0213] The device 1100 also includes:
[0214] The overlapping addition submodule is used to perform OLA on the bandwidth-extended audio signal and the bandwidth-extended reference audio signal to obtain the reconstructed signal of the i-th audio frame. The reference audio signal includes the (i-1)-th audio frame and the i-th audio frame. The reconstructed signal of the i-th audio frame contains a first frequency component and a second frequency component.
[0215] Optionally, the device 1100 further includes:
[0216] The first upsampling module is used to upsample the target audio signal according to the first upsampling parameters to obtain the first audio signal. The target audio signal refers to an audio signal that has frequency components in the second frequency band. The second frequency band is determined based on the sampling rate of the target audio signal.
[0217] The fourth determining module is used to determine the first bandwidth ratio based on the first upsampling parameter.
[0218] Optionally, the device 1100 further includes:
[0219] The second upsampling module is used to upsample the bandwidth-extended audio signal according to the second upsampling parameters to obtain a third audio signal, which is the audio signal to be subjected to a second bandwidth extension.
[0220] The second super-resolution module is used to input the third audio signal into the second super-resolution network to obtain the second super-resolution signal. The second super-resolution signal contains a third frequency component. The third frequency component includes frequency components in the third frequency band that are greater than the second cutoff frequency. The second cutoff frequency is the cutoff frequency of the third audio signal. The third frequency band is determined based on the sampling rate of the third audio signal.
[0221] The fifth determining module is used to determine the second bandwidth proportion based on the second upsampling parameters. The second bandwidth proportion is used to indicate the position of the second cutoff frequency in the third frequency band.
[0222] The sixth determining module is used to determine the second high-pass filter parameters based on the second bandwidth ratio;
[0223] The second high-pass filter module is used to perform high-pass filtering on the second super-division signal according to the second high-pass filter parameters to obtain the fourth audio signal, which contains the third frequency component.
[0224] The second overlay module is used to overlay the third audio signal and the fourth audio signal to obtain an audio signal with double bandwidth expansion.
[0225] Optionally, the first determining module includes:
[0226] The first determining submodule is used to determine the first bandwidth percentage range from multiple bandwidth percentage ranges;
[0227] The first selection submodule is used to select a set of high-pass filter parameters corresponding to the first bandwidth ratio range from multiple sets of high-pass filter parameters corresponding to the multiple bandwidth ratio ranges, and use it as the first high-pass filter parameter.
[0228] Optionally, the device 1100 further includes:
[0229] The selection module is used to select the super-resolution network corresponding to the sampling rate of the first audio signal from multiple super-resolution networks corresponding to multiple sampling rates, and use it as the first super-resolution network.
[0230] Optionally, the device 1100 further includes:
[0231] The acquisition module is used to acquire multiple audio sample sets. Each audio sample set includes multiple audio sample signals. The sampling rate of the audio sample signals in different audio sample sets is different, while the sampling rate of each audio sample signal in the same audio sample set is the same.
[0232] The seventh determination module is used to determine the multiple super-resolution networks based on the multiple sets of audio sample sets.
[0233] Optionally, the seventh determining module includes:
[0234] The training submodule is used to select one set of audio sample signals as the first audio sample set from the multiple sets of audio sample sets, and perform the following operations on the first audio sample set until the following operations are performed on all multiple sets of audio sample sets:
[0235] Noise is added to multiple audio sample signals in the first audio sample set to obtain multiple first sample signals;
[0236] According to one or more sets of low-pass filtering parameters, the multiple first sample signals are low-pass filtered to obtain multiple second sample signals;
[0237] The multiple second sample signals are used as the input to the initial super-resolution network, and the multiple first sample signals are used as the output of the initial super-resolution network. The initial super-resolution network is then trained to obtain a super-resolution network.
[0238] Optionally, the first super-resolution network includes a first processing module, a second processing module, and a third processing module. The second processing module includes a first convolutional submodule, a second convolutional submodule, and an addition submodule. The dilation rate of the convolutional layers in the first and second convolutional submodules is greater than 1.
[0239] The first super-resolution module is specifically used for:
[0240] The first audio signal is input into the first processing module to obtain the first data, which includes a first frequency component and a second frequency component.
[0241] The first data is input into the first convolutional submodule to obtain the second data;
[0242] The second data is input into the second convolutional submodule to obtain the third data;
[0243] The second and third data are input into the addition submodule to obtain the fourth data, which contains the second frequency component, or the fourth data contains a portion of the frequency component in the first frequency component and the second frequency component.
[0244] The fourth data is input into the third processing module to obtain the first super-resolution signal.
[0245] Optionally, the first processing module includes a nonlinear activation layer, and the second frequency component is a frequency component extended by the nonlinear activation layer.
[0246] In summary, compared to bandwidth replication schemes, this scheme, through the expansion of high-frequency components via a super-resolution network, results in a more harmonious and natural sound in the final audio signal. Compared to schemes requiring envelope shaping of the MDCT spectrum, this scheme has lower complexity and improves bandwidth expansion efficiency. Furthermore, this scheme combines a super-resolution network with a high-pass filter; the high-pass filter parameters are determined based on the cutoff frequency of the audio signal to be bandwidth expanded. This demonstrates that this scheme can adapt to different cutoff frequencies, and the super-resolution network in this scheme can handle audio signals with various cutoff frequencies. When the audio signal to be bandwidth expanded is the audio signal obtained during the decoding process, the cutoff frequency is related to the encoding bitrate; the lower the bitrate, the smaller the cutoff frequency. In this case, this scheme can effectively adapt to different bitrates. Moreover, this scheme uses high-pass filtering to ensure that the low-frequency components of the final audio signal remain essentially unchanged, meaning that the low-frequency components are not damaged.
[0247] It should be noted that the audio bandwidth expansion device provided in the above embodiments is only illustrated by the division of the above functional modules when expanding the bandwidth of audio signals. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the audio bandwidth expansion device and the audio bandwidth expansion method embodiments provided in the above embodiments belong to the same concept, and the specific implementation process can be found in the method embodiments, which will not be repeated here.
[0248] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer instructions. When the computer instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium accessible to a computer, or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., digital versatile disc (DVD)), or a semiconductor medium (e.g., solid-state disk (SSD)). It is worth noting that the computer-readable storage medium mentioned in the embodiments of this application can be a non-volatile storage medium; in other words, it can be a non-transient storage medium.
[0249] It should be understood that "at least one" as mentioned herein refers to one or more, and "multiple" refers to two or more. In the description of the embodiments of this application, unless otherwise stated, " / " means "or," for example, A / B can mean A or B; "and / or" in this document is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. In addition, in order to clearly describe the technical solutions of the embodiments of this application, the terms "first," "second," etc., are used in the embodiments of this application to distinguish identical or similar items with substantially the same function and effect. Those skilled in the art will understand that the terms "first," "second," etc., do not limit the quantity or execution order, and the terms "first," "second," etc., are not necessarily different.
[0250] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, stored data, displayed data, etc.), and signals involved in the embodiments of this application are all authorized by the user or fully authorized by all parties, and the collection, use, and processing of related data must comply with the relevant laws, regulations, and standards of the relevant countries and regions. For example, the audio signals involved in the embodiments of this application were all obtained under full authorization.
[0251] The above descriptions are embodiments provided in this application and are not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.
Claims
1. A method for extending the bandwidth of audio signals, characterized in that, The method includes: The first high-pass filter parameter is determined based on the first bandwidth ratio. The first bandwidth ratio is used to indicate the position of the first cutoff frequency in the first frequency band. The first cutoff frequency is the cutoff frequency of the first audio signal to be bandwidth extended. The first audio signal contains a first frequency component. The first frequency component includes frequency components in the first frequency band that are not greater than the first cutoff frequency. The first frequency band is determined based on the sampling rate of the first audio signal. The first audio signal is input into the first super-resolution network to obtain the first super-resolution signal. The first super-resolution signal includes a portion of the frequency components and the second frequency components in the first frequency components. The second frequency components include frequency components in the first frequency band that are greater than the first cutoff frequency. According to the first high-pass filtering parameters, the first super-division signal is subjected to high-pass filtering to obtain a second audio signal, the second audio signal containing the second frequency component; The first audio signal and the second audio signal are superimposed to obtain a bandwidth-extended audio signal.
2. The method as described in claim 1, characterized in that, Before determining the first high-pass filter parameter based on the first bandwidth ratio, the method further includes: The bitstream is parsed and dequantized to obtain the improved cosine transform MDCT spectrum; The bandwidth of the MDCT spectrum is detected to obtain the first cutoff frequency; Divide the first cutoff frequency by the bandwidth of the first frequency band to obtain the first bandwidth ratio; Before inputting the first audio signal into the first super-resolution network to obtain the first super-resolution signal, the method further includes: The first audio signal is determined based on the MDCT spectrum.
3. The method as described in claim 2, characterized in that, Determining the first audio signal based on the MDCT spectrum includes: The MDCT spectrum is subjected to an improved inverse cosine transform (IMDCT) to obtain the first audio signal; The bandwidth-extended audio signal includes the i-th audio frame and the (i+1)-th audio frame, where i is an integer not less than 1; after superimposing the first audio signal and the second audio signal to obtain the bandwidth-extended audio signal, the method further includes: The bandwidth-extended audio signal and the bandwidth-extended reference audio signal are overlapped and added together using OLA to obtain the reconstructed signal of the i-th audio frame. The reference audio signal includes the (i-1)-th audio frame and the i-th audio frame. The reconstructed signal of the i-th audio frame contains the first frequency component and the second frequency component.
4. The method as described in claim 1, characterized in that, Before inputting the first audio signal into the first super-resolution network to obtain the first super-resolution signal, the method further includes: The target audio signal is upsampled according to the first upsampling parameter to obtain the first audio signal. The target audio signal refers to an audio signal that has frequency components in the second frequency band. The second frequency band is determined based on the sampling rate of the target audio signal. Before determining the first high-pass filter parameter based on the first bandwidth ratio, the method further includes: The first bandwidth percentage is determined based on the first upsampling parameter.
5. The method according to any one of claims 1-4, characterized in that, After superimposing the first audio signal and the second audio signal to obtain a bandwidth-extended audio signal, the method further includes: According to the second upsampling parameter, the bandwidth-extended audio signal is upsampled to obtain a third audio signal, which is the audio signal to be subjected to a second bandwidth extension. The third audio signal is input into the second super-resolution network to obtain a second super-resolution signal. The second super-resolution signal contains a third frequency component, which includes frequency components in the third frequency band that are greater than the second cutoff frequency. The second cutoff frequency is the cutoff frequency of the third audio signal. The third frequency band is determined based on the sampling rate of the third audio signal. The second bandwidth percentage is determined based on the second upsampling parameter, and the second bandwidth percentage is used to indicate the position of the second cutoff frequency in the third frequency band. The second high-pass filter parameters are determined based on the second bandwidth ratio. The second super-resolution signal is high-pass filtered according to the second high-pass filtering parameters to obtain a fourth audio signal, the fourth audio signal containing the third frequency component; The third audio signal is superimposed on the fourth audio signal to obtain an audio signal with double bandwidth expansion.
6. The method according to any one of claims 1-4, characterized in that, The step of determining the first high-pass filter parameter based on the first bandwidth ratio includes: Determine the first bandwidth percentage range within which the first bandwidth percentage falls from multiple bandwidth percentage ranges; From the multiple sets of high-pass filter parameters corresponding to the multiple bandwidth ratio ranges, select a set of high-pass filter parameters that corresponds to the first bandwidth ratio range as the first high-pass filter parameter.
7. The method according to any one of claims 1-4, characterized in that, Before inputting the first audio signal into the first super-resolution network to obtain the first super-resolution signal, the method further includes: From multiple super-resolution networks corresponding to multiple sampling rates, the super-resolution network corresponding to the sampling rate of the first audio signal is selected as the first super-resolution network.
8. The method as described in claim 7, characterized in that, Before selecting the super-resolution network corresponding to the sampling rate of the first audio signal from multiple super-resolution networks corresponding to multiple sampling rates as the first super-resolution network, the method further includes: Multiple audio sample sets are acquired, each set including multiple audio sample signals. The sampling rates of the audio sample signals in different audio sample sets are different, while the sampling rates of the audio sample signals in the same audio sample set are the same. Based on the multiple sets of audio samples, the multiple super-resolution networks are determined respectively.
9. The method as described in claim 8, characterized in that, The step of determining the multiple super-resolution networks based on the multiple sets of audio samples includes: From the plurality of audio sample sets, select one set of audio sample signals as the first audio sample set, and perform the following operation on the first audio sample set until the following operation is performed on all the plurality of audio sample sets: Noise is added to multiple audio sample signals in the first audio sample set to obtain multiple first sample signals; According to one or more sets of low-pass filtering parameters, the plurality of first sample signals are low-pass filtered to obtain a plurality of second sample signals; The multiple second sample signals are used as inputs to the initial super-resolution network, and the multiple first sample signals are used as outputs to train the initial super-resolution network to obtain a super-resolution network.
10. The method according to any one of claims 1-4, 8, and 9, characterized in that, The first super-resolution network includes a first processing module, a second processing module, and a third processing module. The second processing module includes a first convolution submodule, a second convolution submodule, and an addition submodule. The dilation rate of the convolutional layers in the first convolution submodule and the second convolution submodule is greater than 1. The step of inputting the first audio signal into the first super-resolution network to obtain the first super-resolution signal includes: The first audio signal is input into the first processing module to obtain first data, the first data including the first frequency component and the second frequency component; The first data is input into the first convolutional submodule to obtain the second data; The second data is input into the second convolutional submodule to obtain the third data; The second data and the third data are input into the addition submodule to obtain the fourth data, which includes the second frequency component, or the fourth data includes a portion of the frequency component in the first frequency component and the second frequency component. The fourth data is input into the third processing module to obtain the first super-resolution signal.
11. The method as described in claim 10, characterized in that, The first processing module includes a nonlinear activation layer, and the second frequency component is a frequency component extended through the nonlinear activation layer.
12. An audio bandwidth extension device, characterized in that, The device includes: The first determining module is used to determine the first high-pass filter parameter based on the first bandwidth ratio. The first bandwidth ratio is used to indicate the position of the first cutoff frequency in the first frequency band. The first cutoff frequency is the cutoff frequency of the first audio signal to be bandwidth extended. The first audio signal contains a first frequency component. The first frequency component includes frequency components in the first frequency band that are not greater than the first cutoff frequency. The first frequency band is determined based on the sampling rate of the first audio signal. The first super-resolution module is used to input the first audio signal into the first super-resolution network to obtain the first super-resolution signal. The first super-resolution signal includes a portion of the frequency components and the second frequency components in the first frequency components. The second frequency components include frequency components in the first frequency band that are greater than the first cutoff frequency. The first high-pass filter module is used to perform high-pass filtering on the first super-division signal according to the first high-pass filter parameters to obtain a second audio signal, wherein the second audio signal contains the second frequency component; The first superposition module is used to superimpose the first audio signal and the second audio signal to obtain a bandwidth-extended audio signal.
13. The apparatus as claimed in claim 12, characterized in that, The device further includes: The decoding module is used to parse and dequantize the bitstream to obtain the improved cosine transform MDCT spectrum. A bandwidth detection module is used to perform bandwidth detection on the MDCT spectrum to obtain the first cutoff frequency; The second determining module is used to divide the first cutoff frequency by the bandwidth of the first frequency band to obtain the first bandwidth ratio; The third determining module is used to determine the first audio signal based on the MDCT spectrum.
14. The apparatus as claimed in claim 13, characterized in that, The third determining module includes: The inverse transform submodule is used to perform an improved inverse cosine transform (IMDCT) on the MDCT spectrum to obtain the first audio signal. The bandwidth-extended audio signal includes the i-th audio frame and the (i+1)-th audio frame, where i is an integer not less than 1; the device further includes: An overlapping addition submodule is used to perform OLA on the bandwidth-extended audio signal and the bandwidth-extended reference audio signal to obtain the reconstructed signal of the i-th audio frame. The reference audio signal includes the (i-1)-th audio frame and the i-th audio frame. The reconstructed signal of the i-th audio frame contains the first frequency component and the second frequency component.
15. The apparatus as claimed in claim 12, characterized in that, The device further includes: The first upsampling module is used to upsample the target audio signal according to the first upsampling parameters to obtain the first audio signal. The target audio signal refers to an audio signal that has frequency components in the second frequency band, and the second frequency band is determined based on the sampling rate of the target audio signal. The fourth determining module is used to determine the first bandwidth ratio based on the first upsampling parameters.
16. The apparatus according to any one of claims 12-15, characterized in that, The device further includes: The second upsampling module is used to upsample the bandwidth-extended audio signal according to the second upsampling parameters to obtain a third audio signal, wherein the third audio signal is the audio signal to be subjected to a second bandwidth extension. The second super-resolution module is used to input the third audio signal into the second super-resolution network to obtain a second super-resolution signal. The second super-resolution signal contains a third frequency component. The third frequency component includes frequency components in the third frequency band that are greater than the second cutoff frequency. The second cutoff frequency is the cutoff frequency of the third audio signal. The third frequency band is determined based on the sampling rate of the third audio signal. The fifth determining module is used to determine the second bandwidth ratio based on the second upsampling parameter, wherein the second bandwidth ratio is used to indicate the position of the second cutoff frequency in the third frequency band; The sixth determining module is used to determine the second high-pass filter parameters based on the second bandwidth ratio; The second high-pass filter module is used to perform high-pass filtering on the second super-division signal according to the second high-pass filter parameters to obtain a fourth audio signal, the fourth audio signal containing the third frequency component; The second overlay module is used to overlay the third audio signal and the fourth audio signal to obtain an audio signal with double bandwidth expansion.
17. The apparatus according to any one of claims 12-15, characterized in that, The first determining module includes: The first determining submodule is used to determine the first bandwidth percentage range in which the first bandwidth percentage is located from multiple bandwidth percentage ranges; The first selection submodule is used to select a set of high-pass filter parameters corresponding to the first bandwidth ratio range from multiple sets of high-pass filter parameters corresponding to the multiple bandwidth ratio ranges, and use it as the first high-pass filter parameter.
18. The apparatus according to any one of claims 12-15, characterized in that, The device further includes: The selection module is used to select, from multiple super-resolution networks corresponding to multiple sampling rates, the super-resolution network corresponding to the sampling rate of the first audio signal, and use it as the first super-resolution network.
19. The apparatus as claimed in claim 18, characterized in that, The device further includes: The acquisition module is used to acquire multiple audio sample sets. Each audio sample set includes multiple audio sample signals. The sampling rate of the audio sample signals in different audio sample sets is different, while the sampling rate of each audio sample signal in the same audio sample set is the same. The seventh determining module is used to determine the multiple super-resolution networks based on the multiple sets of audio sample sets.
20. The apparatus as claimed in claim 19, characterized in that, The seventh determining module includes: The training submodule is used to select one set of audio sample signals as the first audio sample set from the multiple sets of audio sample sets, and perform the following operations on the first audio sample set until the following operations are performed on all the multiple sets of audio sample sets: Noise is added to multiple audio sample signals in the first audio sample set to obtain multiple first sample signals; According to one or more sets of low-pass filtering parameters, the plurality of first sample signals are low-pass filtered to obtain a plurality of second sample signals; The multiple second sample signals are used as inputs to the initial super-resolution network, and the multiple first sample signals are used as outputs to train the initial super-resolution network to obtain a super-resolution network.
21. The apparatus as described in any one of claims 12-15, 19, and 20, characterized in that, The first super-resolution network includes a first processing module, a second processing module, and a third processing module. The second processing module includes a first convolution submodule, a second convolution submodule, and an addition submodule. The dilation rate of the convolutional layers in the first convolution submodule and the second convolution submodule is greater than 1. The first super-resolution module is specifically used for: The first audio signal is input into the first processing module to obtain first data, the first data including the first frequency component and the second frequency component; The first data is input into the first convolutional submodule to obtain the second data; The second data is input into the second convolutional submodule to obtain the third data; The second data and the third data are input into the addition submodule to obtain the fourth data, which includes the second frequency component, or the fourth data includes a portion of the frequency component in the first frequency component and the second frequency component. The fourth data is input into the third processing module to obtain the first super-resolution signal.
22. The apparatus as claimed in claim 21, characterized in that, The first processing module includes a nonlinear activation layer, and the second frequency component is a frequency component extended through the nonlinear activation layer.
23. A computer-readable storage medium, characterized in that, The storage medium stores a computer program, which, when executed by a processor, implements the steps of the method described in any one of claims 1-11.
24. A computer program product, characterized in that, The computer program product stores computer instructions, which, when executed by a processor, implement the steps of the method described in any one of claims 1-11.
Citation Information
Patent Citations
Enhanced audio decoder
CN102598121A
Reduced-bandwidth speech enhancement with bandwidth extension
WO2021207131A1