A Bluetooth audio pitch shifting method, apparatus, medium, and Bluetooth device
By utilizing the time-frequency conversion data during the Bluetooth encoding and decoding process, and employing discrete cosine transform to calculate the amplitude pseudospectrum and fundamental frequency of the audio spectral coefficients, phase and frequency processing is performed. This solves the problems of excessive computation and latency in existing technologies, achieving low-complexity and low-latency audio pitch shifting and improving the user experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-17
- Publication Date
- 2026-03-06
AI Technical Summary
Existing Bluetooth audio pitch shifting technology increases the computational load and system latency at the Bluetooth transmitter when using the frequency domain method, affecting the user experience. Latency is a key technical indicator, especially for users such as broadcasters.
By utilizing the time-frequency conversion data during the Bluetooth encoding and decoding process, the amplitude pseudospectrum of the audio spectral coefficients is calculated using discrete cosine transform to obtain the fundamental frequency, and phase and frequency phase shifting is performed to reduce system complexity and computational load.
It reduces system latency, improves user experience, and allows for flexible setting of pitch parameters to meet the needs of different users.
Smart Images

Figure CN115641859B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of audio encoding and decoding technology, and in particular to a Bluetooth audio pitch shifting method, apparatus, medium and Bluetooth device. Background Technology
[0002] Audio pitch shifting has applications in many scenarios. For example, one party in a call may not want the other party to hear their voice clearly, and in the currently popular live streaming, broadcasters may want to add sound effects to their voices, such as changing a female voice to a male voice or vice versa. Pitch shifting algorithms can be inserted into the Bluetooth call link to achieve this. Two typical algorithms can be used for pitch shifting: one is the time-domain method, which uses synchronous superposition and fixed synthesis. Its advantage is its simplicity and good effect when the pitch shifting range is small, but distortion may occur when the pitch shifting range is large. The other is the frequency-domain method, i.e., the phase vocoder. It uses Fourier transform to convert the signal to the frequency domain, extracts the amplitude and frequency, calculates the updated amplitude and frequency, and uses inverse Fourier transform to convert it back to the time domain. Although it can support larger-scale pitch shifting and better ensure the scale and effect of pitch shifting, current technology using the frequency-domain method increases the computational load on the Bluetooth transmitter and increases the end-to-end audio latency of the entire system. For some users, such as broadcasters, latency is a critical technical indicator that directly affects the user experience. Summary of the Invention
[0003] To address the problems existing in the prior art, this application mainly provides a Bluetooth audio pitch shifting method, apparatus, medium, and Bluetooth device. By utilizing the time-frequency conversion data already available during the Bluetooth encoding and decoding process, the system complexity can be reduced, the amount of computation can be decreased, the system latency can be reduced, and the user experience can be improved.
[0004] To achieve the above objectives, one technical solution adopted in this application is: providing a Bluetooth audio pitch shifting method, which includes: during the Bluetooth audio encoding and / or decoding process, calculating the amplitude pseudo-spectrum corresponding to each spectral coefficient using each spectral coefficient in the current frame audio spectral coefficient obtained by discrete cosine transform, a predetermined number of adjacent spectral coefficients before each spectral coefficient, and a predetermined number of adjacent spectral coefficients after each spectral coefficient; calculating the fundamental frequency of the current frame audio signal corresponding to the current frame audio spectral coefficient using the amplitude pseudo-spectrum of the current frame audio spectral coefficient to obtain the current frame audio fundamental frequency; performing phase shift processing on the phase of the current frame audio fundamental frequency and performing frequency up / down processing on the current frame audio fundamental frequency to obtain the pitch-shifted current frame audio fundamental frequency; generating the pitch-shifted current frame audio spectral coefficient using the pitch-shifted current frame audio fundamental frequency, and completing the remaining Bluetooth audio encoding and / or decoding steps using the pitch-shifted current frame audio spectral coefficient.
[0005] Another technical solution adopted in this application is: providing a Bluetooth audio pitch-shifting device, comprising: an amplitude pseudo-spectrum calculation module, used to calculate the amplitude pseudo-spectrum corresponding to each spectral coefficient in the current frame audio spectral coefficient obtained by discrete cosine transform, a predetermined number of adjacent spectral coefficients before each spectral coefficient, and a predetermined number of adjacent spectral coefficients after each spectral coefficient during Bluetooth audio encoding and / or decoding; a fundamental frequency acquisition module, used to calculate the fundamental frequency of the current frame audio signal corresponding to the current frame audio spectral coefficient using the amplitude pseudo-spectrum of the current frame audio spectral coefficient to obtain the current frame audio fundamental frequency; a pitch-shifting processing module, used to perform phase shifting processing on the phase of the current frame audio fundamental frequency and frequency up / down processing on the current frame audio fundamental frequency to obtain the pitch-shifted current frame audio fundamental frequency; and a post-processing module, used to generate the pitch-shifted current frame audio spectral coefficient using the pitch-shifted current frame audio fundamental frequency, and use the pitch-shifted current frame audio spectral coefficient to complete the remaining Bluetooth audio encoding and / or decoding steps.
[0006] Another technical solution adopted in this application is to provide a Bluetooth device, which includes an encoder and a decoder, wherein the encoder and / or decoder are provided with the above-mentioned Bluetooth audio pitch shifting device.
[0007] Another technical solution adopted in this application is to provide a computer-readable storage medium storing computer instructions that are operated to execute a Bluetooth audio pitch shifting method as described above.
[0008] The beneficial effects that the technical solution of this application can achieve are: providing a Bluetooth audio pitch shifting method, device, medium and Bluetooth device. By utilizing the time-frequency conversion data already available in the Bluetooth encoding and decoding process, this application can reduce system complexity, reduce computation, reduce system latency, improve user experience, and also flexibly set pitch shifting parameters to meet the needs of different users. Attached Figure Description
[0009] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0010] Figure 1 This is a schematic diagram of the existing Bluetooth audio pitch shifting process;
[0011] Figure 2 This is a flowchart illustrating a specific implementation of a Bluetooth audio pitch shifting method according to this application;
[0012] Figure 3This is a schematic diagram of the Bluetooth audio pitch shifting process in a specific embodiment of a Bluetooth audio pitch shifting method according to this application;
[0013] Figure 4 This is a schematic diagram of the discrete cosine transform spectral coefficients and the fundamental frequency and harmonics in the corresponding pseudospectral amplitude in a specific embodiment of a Bluetooth audio pitch shifting method of this application;
[0014] Figure 5 This is a schematic diagram of a specific embodiment of a Bluetooth audio pitch-shifting device according to this application;
[0015] The accompanying drawings illustrate specific embodiments of this application, which will be described in more detail below. These drawings and descriptions are not intended to limit the scope of the concept in any way, but rather to illustrate the concept of this application to those skilled in the art through reference to particular embodiments. Detailed Implementation
[0016] The preferred embodiments of this application will now be described in detail with reference to the accompanying drawings, so that the advantages and features of this application can be more easily understood by those skilled in the art, thereby providing a clearer and more definite definition of the scope of protection of this application.
[0017] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising..." does not exclude the presence of additional identical elements in the process, method, article, or apparatus that includes the element.
[0018] Audio pitch shifting has applications in many scenarios. For example, one party in a call may not want the other party to recognize their voice, and in the currently popular live streaming, broadcasters may want to add sound effects to their voices, such as changing a female voice to a male voice or vice versa. Pitch shifting algorithms can be inserted into the Bluetooth call link to achieve this. Figure 1 As shown.
[0019] Regarding pitch parameters, the most commonly used law method is the 12-tone equal temperament, which divides a pure octave into 12 equal semitones, with the physical vibration frequency of each tone differing by a factor of 1 / 2. Two adjacent pure octave frequencies differ by a factor of two. If the vibration frequencies of each frequency component of a sound are increased... A doubling of the pitch is equivalent to raising the tone by a semitone, which is the same as lowering the vibration frequency of each frequency component of the sound. A multiple of 100 is equivalent to lowering the pitch by a semitone.
[0020] The pitch shift parameter (semitone) can be preset by the system or selected by the user (via an app). Assuming the original audio signal has a frequency of f1 and the shifted frequency is f2, the following relationship holds:
[0021] f2 = f1x2 semitone / 12 Semitone = ±1, ±2, ±3, ...
[0022] In the above formula, semitone is the pitch shift parameter that needs to be input. A value greater than 0 indicates a rising pitch, and a value less than 0 indicates a falling pitch.
[0023] Existing technologies mostly employ frequency domain methods for pitch shifting, namely phase vocoders. These methods use Fourier transform to convert the signal to the frequency domain, extract amplitude and frequency, calculate updated amplitude and frequency, and use inverse Fourier transform to convert it back to the time domain. While this can support larger-scale pitch shifting and better ensure the scale and effect of pitch shifting, existing technologies using frequency domain methods increase the computational load on the Bluetooth transmitter and increase the end-to-end audio latency of the entire system. For some users, such as broadcasters, latency is a critical technical indicator that directly affects the user experience.
[0024] This application utilizes existing time-frequency conversion data from the Bluetooth encoding / decoding process to reduce system complexity, computational load, system latency, and improve user experience.
[0025] The technical solution of this application and how the technical solution of this application solves the above-mentioned technical problems are described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments. The embodiments of this application will now be described with reference to the accompanying drawings.
[0026] Figure 2 This illustration shows a specific implementation of a Bluetooth audio pitch shifting method according to this application.
[0027] exist Figure 2The Bluetooth audio howling detection and suppression method of this application, as shown, includes: process S201, in the Bluetooth audio encoding and / or decoding process, calculating the amplitude pseudo-spectrum corresponding to each spectral coefficient using each spectral coefficient in the current frame audio spectral coefficient obtained by discrete cosine transform, a predetermined number of adjacent spectral coefficients before each spectral coefficient, and a predetermined number of adjacent spectral coefficients after each spectral coefficient; process S202, calculating the fundamental frequency of the current frame audio signal corresponding to the current frame audio spectral coefficient using the amplitude pseudo-spectrum of the current frame audio spectral coefficient to obtain the current frame audio fundamental frequency; process S203, performing phase shift processing on the phase of the current frame audio fundamental frequency and performing up / down processing on the frequency of the current frame audio fundamental frequency to obtain the current frame audio fundamental frequency after pitch shifting; and process S204, generating the current frame audio spectral coefficient after pitch shifting using the current frame audio fundamental frequency after pitch shifting, and completing the remaining Bluetooth audio encoding and / or decoding steps using the current frame audio spectral coefficient after pitch shifting.
[0028] This application utilizes existing time-frequency conversion data from the Bluetooth encoding / decoding process to reduce system complexity, computational load, system latency, and improve user experience.
[0029] Process S201 represents the process of calculating the amplitude pseudospectrum of each spectral coefficient in the current frame audio spectral coefficient obtained by discrete cosine transform, a predetermined number of adjacent spectral coefficients before each spectral coefficient, and a predetermined number of adjacent spectral coefficients after each spectral coefficient during Bluetooth audio encoding and / or decoding. This process can facilitate the calculation of the fundamental frequency of the current frame audio using the amplitude pseudospectrum of the current frame audio spectral coefficient.
[0030] Specifically, unlike the spectral coefficients obtained through Fourier transform, which have a good correspondence with the spectral components of actual audio data, the audio spectral coefficients obtained through discrete cosine transform have a deviance in their correspondence with the spectral components of actual audio data. This application uses the pseudo-spectral amplitude, calculated from the audio spectral coefficients obtained through discrete cosine transform, which has a good correspondence with the spectral components of actual audio data, to calculate the fundamental frequency, thus enabling a more accurate calculation of the fundamental frequency.
[0031] In one specific embodiment of this application, the audio spectrum coefficients obtained by the discrete cosine transform are the audio spectrum coefficients X(k) obtained by the low-latency improved discrete cosine transform during the LC3 encoding process.
[0032]
[0033] In one specific embodiment of this application, the audio spectrum coefficients obtained by discrete cosine transform are the audio spectrum coefficients obtained after arithmetic and residual decoding, noise filling, global gain, and time-domain noise shaping steps during the LC3 decoding process.
[0034] In one specific embodiment of this application, the audio spectrum coefficients obtained by discrete cosine transform are the audio spectrum coefficients obtained by discrete cosine transform when using an AAC Bluetooth audio device for encoding and decoding.
[0035] In one specific embodiment of this application, before performing the Bluetooth audio encoding and / or decoding process, ultra-low frequency tone signals and near-Nyquist tone signals in the current frame audio signal are filtered out, such as... Figure 3 As shown.
[0036] Specifically, taking 8K sampling rate as an example, frequency components below 150Hz and above 3850Hz are filtered out to avoid ultra-low frequency tone signals and near-Nyquist tone signals affecting the robustness of the pitch shifting algorithm. Signal processing for other sampling rates is similar.
[0037] In one specific embodiment of this application, the predetermined number of adjacent spectral coefficients are multiple spectral coefficients, that is, the amplitude pseudo-spectrum corresponding to the spectral coefficient is calculated by using the same number of spectral coefficients before and after each spectral coefficient of the current frame audio spectral coefficient and the spectral coefficient itself.
[0038] In one specific embodiment of this application, the predetermined number of adjacent spectral coefficients constitute one spectral coefficient. That is, using the preceding spectral coefficient, the following spectral coefficient, and the spectral coefficient itself, the amplitude pseudo-spectrum corresponding to that spectral coefficient is obtained. Specifically, the calculation process can be as follows:
[0039]
[0040] Where X(-1)=X(N) F ) = 0
[0041] Process S202 represents the process of calculating the fundamental frequency of the current frame audio signal corresponding to the current frame audio spectral coefficient using the amplitude pseudospectral of the current frame audio spectral coefficient, which can facilitate the modulation processing of the obtained fundamental frequency and further perform the remaining encoding and decoding process to obtain the modulation-modulated Bluetooth audio.
[0042] Specifically, because the pseudo-spectral amplitude has a good correspondence with the spectral components of the actual audio data, the fundamental frequency of the pseudo-spectral amplitude is the fundamental frequency of the actual audio signal.
[0043] In one specific embodiment of this application, the process of calculating the fundamental frequency of the current frame audio signal corresponding to the current frame audio spectral coefficients using the amplitude pseudospectral of the current frame audio spectral coefficients to obtain the fundamental frequency of the current frame audio includes: determining the sub-band index of the frequency sub-band to which the maximum amplitude pseudospectral value belongs in all amplitude pseudospectral values of the current frame audio spectral coefficients as the integer part of the index of the current frame audio fundamental frequency; calculating the fractional part of the index of the current frame audio fundamental frequency using the ratio of the amplitude value of the previous sub-band to the amplitude value of the next sub-band of the frequency sub-band to which the maximum amplitude pseudospectral value belongs, or the ratio of the amplitude value of the second-to-last sub-band to the amplitude value of the second-to-last sub-band of the frequency sub-band to which the maximum amplitude pseudospectral value belongs; and obtaining the current frame audio based on the integer part of the index of the current frame audio fundamental frequency and the fractional part of the index of the previous frame audio fundamental frequency.
[0044] In one specific embodiment of this application, the process of calculating the fundamental frequency of the current frame audio signal corresponding to the current frame audio spectral coefficients using the amplitude pseudospectral of the current frame audio spectral coefficients to obtain the fundamental frequency of the current frame audio includes: determining the sub-band index of the frequency sub-band to which the maximum amplitude pseudospectral value belongs among all the amplitude pseudospectral values of the current frame audio spectral coefficients as the integer part of the index of the current frame audio fundamental frequency; calculating the fractional part of the index of the current frame audio fundamental frequency using the ratio of the amplitude value of the previous sub-band to the amplitude value of the next sub-band; and obtaining the current frame audio based on the integer part of the index of the current frame audio fundamental frequency and the fractional part of the index of the previous frame audio fundamental frequency.
[0045] In one specific embodiment of this application, the process of calculating the fundamental frequency of the current frame audio signal corresponding to the current frame audio spectral coefficients using the amplitude pseudospectral of the current frame audio spectral coefficients to obtain the fundamental frequency of the current frame audio includes: determining the sub-band index of the frequency sub-band to which the maximum amplitude pseudospectral value belongs in all amplitude pseudospectral values of the current frame audio spectral coefficients as the integer part of the index of the current frame audio fundamental frequency; calculating the fractional part of the index of the current frame audio fundamental frequency using the ratio of the amplitude value of the second-to-last sub-band of the frequency sub-band to which the maximum amplitude pseudospectral value belongs to the frequency sub-band to which the maximum amplitude pseudospectral value belongs to the frequency sub-band to the second-to-last sub-band; and obtaining the current frame audio based on the integer part of the index of the current frame audio fundamental frequency and the fractional part of the index of the previous frame audio fundamental frequency.
[0046] Specifically, the fundamental frequency corresponds to the fundamental tone frequency in speech. The fundamental tone frequency of a person is typically in the range of 50–400 Hz. When we convert speech to the frequency domain (performing a discrete cosine transform), the fundamental frequency is usually represented by a frequency index. For example, if the sampling rate of the input speech is 8000 Hz and its Nyquist bandwidth is 4000 Hz, after performing a low-latency improved discrete cosine transform, we can obtain X(k), k = 0…N_F-1, where N_F = 80, meaning the distance between two adjacent frequency indices is 50 Hz. Assuming someone's fundamental frequency is 240 Hz, then ideally, the integer part of the fundamental frequency is 240 / 50 = 4, and the fractional part is 40 / 50 = 0.8. Therefore, the fundamental frequency here is a combination of the integer and fractional parts of the frequency index.
[0047] In a specific example of this application, the calculation process of determining the sub-band index of the sub-band to which the frequency corresponding to the maximum amplitude pseudo-spectral value belongs among all the amplitude pseudo-spectral values of the current frame audio spectral coefficients is the integer part of the index of the current frame audio fundamental frequency can be as follows:
[0048] Pitch int =k, when X Est (k)=Max(X Est (k)),
[0049] In a specific example of this application, the calculation process for obtaining the fractional part of the index of the fundamental frequency of the current frame audio using the ratio of the amplitude value of the sub-band preceding the frequency corresponding to the maximum amplitude pseudo-spectral value to the amplitude value of the sub-band following the frequency corresponding to the maximum amplitude pseudo-spectral value, or the ratio of the amplitude value of the second-to-last sub-band preceding the frequency corresponding to the frequency corresponding to the maximum amplitude pseudo-spectral value to the amplitude value of the second-to-last sub-band following the frequency corresponding to the frequency corresponding to the maximum amplitude pseudo-spectral value, is as follows:
[0050]
[0051]
[0052]
[0053]
[0054]
[0055] If ratio peak Pitch less than the preset ratio threshold threshold Then Pitch frac =Pitch frac1 Otherwise, Pitch frac =Pitch frac2 .
[0056] The aforementioned ratio threshold can be taken as an empirical value of 0.96.
[0057] Process S203 represents the process of performing phase shift processing on the phase of the current frame audio base frequency and performing up / down processing on the frequency of the current frame audio base frequency to obtain the pitch-shifted current frame audio base frequency. This process facilitates the adjustment of the frequency and phase of the current frame audio base frequency to obtain the pitch-shifted current frame audio base frequency, so as to further perform the remaining encoding and decoding processes to obtain the pitch-shifted Bluetooth audio.
[0058] In one specific embodiment of this application, before performing phase shifting processing on the phase of the current frame audio base frequency, amplitude extraction and phase extraction are performed on the current frame audio base frequency.
[0059] Specifically, the calculation process for amplitude extraction can be as follows:
[0060]
[0061] The phase extraction process can be as follows:
[0062]
[0063] In an optional embodiment of this application, the calculation process for phase-shifting the phase of the current frame audio fundamental frequency can be as follows:
[0064]
[0065] In an optional specific example of this application, the calculation process for adjusting the frequency of the current frame audio fundamental frequency can be as follows:
[0066]
[0067] Semitone indicates the magnitude of pitch shifting, i.e., the pitch parameter.
[0068] This example allows for flexible setting of pitch shift parameters to meet the needs of different customers.
[0069] In one specific embodiment of this application, the Bluetooth pitch shifting method further includes obtaining the current frame audio harmonics of the current frame audio signal based on the current frame audio fundamental frequency and the amplitude pseudospectral of the current frame audio spectral coefficients; and performing phase shifting processing on the phase of the current frame audio harmonics and raising / lowering processing on the frequency of the current frame audio harmonics to obtain the pitch-shifted current frame audio harmonics.
[0070] In one specific embodiment of this application, the aforementioned current frame audio harmonics include the second harmonic of the current frame audio fundamental frequency. The process of obtaining the current frame audio harmonics of the current frame audio signal based on the current frame audio fundamental frequency and the amplitude pseudospectral of the current frame audio spectral coefficients includes: determining the sub-band index of the sub-band corresponding to the frequency of the maximum amplitude pseudospectral value within a width range equal to twice the index of the current frame audio fundamental frequency, centered at twice the index of the current frame audio fundamental frequency, as the integer part of the index of the second harmonic of the current frame audio fundamental frequency.
[0071] Specifically, for example Figure 4 The audio signal in the example shown Figure 4 (A) represents the discrete cosine transform spectral coefficients. Figure 4 (B) is the corresponding pseudospectral amplitude, which shows that... Figure 4 (B) indicates that the index of the fundamental frequency is 27. By searching with twice the fundamental frequency (54) as the center and once the fundamental frequency (27) as the center, the index of the second harmonic can be obtained as 52.
[0072] In one specific embodiment of this application, the current frame audio harmonics include the second, third, fourth, and fifth harmonics of the current frame audio fundamental frequency.
[0073] Specifically, in tone signals, besides the most important fundamental frequency, its harmonics are also crucial. Their contribution to sound quality is typically ranked in order: second harmonic, third harmonic, fourth harmonic, etc. To balance sound quality and computational complexity, the second to fifth harmonics are usually processed.
[0074] Specifically, following the method for calculating the second harmonic described above, the process for calculating the third harmonic involves determining the sub-band index of the frequency corresponding to the maximum amplitude pseudospectral value within a width range equal to one index of the current frame's audio fundamental frequency, centered at three times the index of the current frame's audio fundamental frequency. The calculation methods for the integer parts of the fourth and fifth harmonic indices are similar.
[0075] In an optional specific example of this application, the fractional part of the audio harmonics of the current frame is calculated in the same way as the fundamental frequency of the audio of the current frame.
[0076] In an optional specific example of this application, the method for phase-shifting the phase of the current frame audio harmonics and the method for frequency adjustment are the same as the method for processing the current frame audio fundamental frequency.
[0077] Process S204 represents the process of generating the audio spectrum coefficients of the current frame after pitch shifting using the audio base frequency of the current frame after pitch shifting, and using the audio spectrum coefficients of the current frame after pitch shifting to complete the remaining Bluetooth audio encoding and / or decoding steps. This process can obtain the audio spectrum coefficients after pitch shifting and complete the pitch shifting of Bluetooth audio under the premise of low system complexity, low computational load, low system latency, and good user experience.
[0078] In one specific embodiment of this application, the calculation process of generating the spectral coefficients of the current frame audio after pitch shifting using the fundamental frequency of the current frame audio after pitch shifting can be as follows:
[0079]
[0080] Pitch represents the fundamental frequency component. int Pitch represents its integer part. frac It represents its decimal part.
[0081] In one specific embodiment of this application, the process of generating the audio spectrum coefficients of the current frame after pitch shifting using the fundamental frequency of the current frame audio after pitch shifting includes: generating the audio spectrum coefficients of the current frame audio after pitch shifting using the fundamental frequency of the current frame audio after pitch shifting and the harmonics of the current frame audio.
[0082] In one specific embodiment of this application, such as Figure 3 As shown, when pitch shifting is performed during LC3 encoding, the process of using the audio spectrum coefficients of the current frame after pitch shifting to complete the remaining Bluetooth audio encoding and / or decoding steps includes: transform domain noise shaping step, time domain noise shaping step, quantization step, noise level estimation step, arithmetic coding, and residual coding steps.
[0083] In this specific embodiment, attack detection and bandwidth detection are performed according to the LC3 standard specification, but resampling and long-term post-filter encoding steps are omitted. Specifically, resampling and long-term post-filter encoding steps are used to enhance the fundamental frequency components in the sound during decoding. After pitch shifting, the fundamental frequency changes, and continuing to execute the long-term post-filter would lead to a decrease in sound quality.
[0084] In one specific embodiment of this application, when pitch shifting is performed during LC3 decoding, the long-term post-filter decoding step is not performed.
[0085] Figure 5 This application illustrates a Bluetooth audio pitch shifting device.
[0086] exist Figure 5The Bluetooth audio howling detection and suppression device shown includes an amplitude pseudo-spectrum calculation module 501, used to calculate the amplitude pseudo-spectrum corresponding to each spectral coefficient in the current frame audio spectral coefficients obtained by discrete cosine transform, a predetermined number of adjacent spectral coefficients before each spectral coefficient, and a predetermined number of adjacent spectral coefficients after each spectral coefficient during Bluetooth audio encoding and / or decoding; a fundamental frequency acquisition module 502, used to calculate the fundamental frequency of the current frame audio signal corresponding to the current frame audio spectral coefficient using the amplitude pseudo-spectrum of the current frame audio spectral coefficients to obtain the current frame audio fundamental frequency; a pitch shifting processing module 503, used to perform phase shifting processing on the phase of the current frame audio fundamental frequency and frequency shifting processing on the current frame audio fundamental frequency to obtain the pitch-shifted current frame audio fundamental frequency; and a post-processing module 504, used to generate the pitch-shifted current frame audio spectral coefficients using the pitch-shifted current frame audio fundamental frequency, and use the pitch-shifted current frame audio spectral coefficients to complete the remaining Bluetooth audio encoding and / or decoding steps.
[0087] This application utilizes existing time-frequency conversion data from the Bluetooth encoding / decoding process to reduce system complexity, computational load, system latency, and improve user experience. It also allows for flexible setting of pitch shift parameters to meet the needs of different users.
[0088] An amplitude pseudospectral calculation module 501 is used to calculate the amplitude pseudospectral of each spectral coefficient in the current frame audio spectral coefficient obtained by discrete cosine transform, a predetermined number of adjacent spectral coefficients before each spectral coefficient, and a predetermined number of adjacent spectral coefficients after each spectral coefficient during Bluetooth audio encoding and / or decoding. This module is useful for calculating the fundamental frequency of the current frame audio using the amplitude pseudospectral of the current frame audio spectral coefficient.
[0089] The fundamental frequency acquisition module 502 is used to calculate the fundamental frequency of the current frame audio signal corresponding to the current frame audio spectrum coefficient using the amplitude pseudospectral of the current frame audio spectrum coefficient. It can facilitate the modulation processing of the obtained fundamental frequency and further perform the remaining encoding and decoding process to obtain the modulation Bluetooth audio.
[0090] The pitch-shifting module 503 is used to perform phase-shifting processing on the phase of the current frame audio base frequency and to perform pitch-shifting processing on the frequency of the current frame audio base frequency to obtain the pitch-shifted current frame audio base frequency, so as to further perform the remaining encoding and decoding process to obtain the channel-shifted Bluetooth audio.
[0091] The post-processing module 504 is used to generate the audio spectrum coefficients of the current frame after pitch shifting using the audio base frequency of the current frame after pitch shifting, and to complete the remaining Bluetooth audio encoding and / or decoding steps using the audio spectrum coefficients of the current frame after pitch shifting. It can obtain the audio spectrum coefficients after pitch shifting and complete the pitch shifting of Bluetooth audio under the premise of low system complexity, low computational load, low system latency, and good user experience.
[0092] The Bluetooth audio pitch shifting device provided in this application can be used to execute the Bluetooth audio pitch shifting method in any of the above embodiments. Its principle and technical effect are similar, and will not be described again here.
[0093] In one specific embodiment of this application, the functional modules of the Bluetooth audio pitch-shifting device of this application may be directly in hardware, in software modules executed by a processor, or in a combination of both.
[0094] Software modules may reside in RAM memory, flash memory, ROM memory, EPROM memory, EEPROM memory, registers, hard disks, removable disks, CD-ROMs, or any other form of storage medium known in this art. An exemplary storage medium is coupled to the processor, enabling the processor to read information from and write information to the storage medium.
[0095] The processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic, discrete hardware components, or any combination thereof. A general-purpose processor can be a microprocessor, but alternatively, it can be any conventional processor, controller, microcontroller, or state machine. The processor can also be implemented as a combination of computing devices, such as a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors incorporating a DSP core, or any other such configuration. Alternatively, the storage medium can be integrated with the processor. The processor and storage medium can reside in an ASIC. The ASIC can reside in the user terminal. Alternatively, the processor and storage medium can reside as discrete components in the user terminal.
[0096] In one specific embodiment of this application, a Bluetooth device includes an encoder and a decoder. The encoder and / or decoder are provided with the Bluetooth audio pitch shifting device described in any of the above embodiments.
[0097] In another specific embodiment of this application, a computer-readable storage medium stores computer instructions that are operated to perform the Bluetooth audio pitch shifting method in any of the above embodiments.
[0098] In another specific embodiment of this application, a computer device includes a processor and a memory, the memory storing computer instructions that are operated to perform the Bluetooth audio pitch shifting method in any of the above embodiments.
[0099] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0100] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0101] The above are merely embodiments of this application and do not limit the scope of this patent application. Any equivalent structural transformations made using the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the scope of patent protection of this application.
Claims
1. A Bluetooth audio pitch shifting method, characterized by, The method comprises: In the process of Bluetooth audio encoding and / or decoding, each spectral coefficient in the current frame audio spectral coefficient obtained by discrete cosine transform, the predetermined number of adjacent spectral coefficients before the each spectral coefficient, and the predetermined number of adjacent spectral coefficients after the each spectral coefficient are used to calculate the corresponding amplitude pseudo-spectrum of the each spectral coefficient; The fundamental frequency of the current frame audio signal corresponding to the current frame audio spectral coefficient is calculated by using the amplitude pseudo-spectrum of the current frame audio spectral coefficient to obtain the current frame audio fundamental frequency; The phase of the current frame audio fundamental frequency is subjected to phase shift processing, and the frequency of the current frame audio fundamental frequency is subjected to frequency raising and lowering processing to obtain the pitch-modulated current frame audio fundamental frequency; And The pitch-modulated current frame audio spectral coefficient is generated by using the pitch-modulated current frame audio fundamental frequency, and the remaining steps of Bluetooth audio encoding and / or decoding are completed by using the pitch-modulated current frame audio spectral coefficient; The process of calculating the fundamental frequency of the current frame audio signal corresponding to the current frame audio spectral coefficient by using the amplitude pseudo-spectrum of the current frame audio spectral coefficient to obtain the current frame audio fundamental frequency comprises: The sub-band index of the sub-band to which the frequency corresponding to the maximum amplitude pseudo-spectrum value in all the amplitude pseudo-spectrum of the current frame audio spectral coefficient belongs is determined as the index integer part of the current frame audio fundamental frequency; The index decimal part of the current frame audio fundamental frequency is calculated by using the ratio of the amplitude value of the sub-band before the sub-band to which the frequency corresponding to the maximum amplitude pseudo-spectrum value belongs to the amplitude value of the sub-band after the sub-band to which the frequency corresponding to the maximum amplitude pseudo-spectrum value belongs, or the ratio of the amplitude value of the second sub-band before the sub-band to which the frequency corresponding to the maximum amplitude pseudo-spectrum value belongs to the amplitude value of the second sub-band after the sub-band to which the frequency corresponding to the maximum amplitude pseudo-spectrum value belongs; and The current frame audio is obtained according to the index integer part of the current frame audio fundamental frequency and the index decimal part of the previous frame audio fundamental frequency.
2. The Bluetooth audio transposing method according to claim 1, wherein, Further comprising: The current frame audio harmonic of the current frame audio signal is obtained according to the current frame audio fundamental frequency and the amplitude pseudo-spectrum of the current frame audio spectral coefficient; And The phase of the current frame audio harmonic is subjected to phase shift processing, and the frequency of the current frame audio harmonic is subjected to frequency raising and lowering processing to obtain the pitch-modulated current frame audio harmonic; The process of generating the pitch-modulated current frame audio spectral coefficient by using the pitch-modulated current frame audio fundamental frequency comprises: The pitch-modulated current frame audio spectral coefficient is generated by using the pitch-modulated current frame audio fundamental frequency and the current frame audio harmonic.
3. The Bluetooth audio transposing method of claim 1, wherein, Further comprising: Before the process of Bluetooth audio encoding and / or decoding is performed, the ultra-low frequency tone signal and the near-Nyquist tone signal in the current frame audio signal are filtered out.
4. The Bluetooth audio pitch-modulating method according to claim 1, wherein The predetermined number of spectral coefficients is one spectral coefficient.
5. The Bluetooth audio transposing method of claim 2, wherein, The current frame audio harmonic comprises the second harmonic of the current frame audio fundamental frequency; The process of obtaining the current frame audio harmonic of the current frame audio signal according to the current frame audio fundamental frequency and the amplitude pseudo-spectrum of the current frame audio spectral coefficient comprises: The subband index of the subband to which the frequency corresponding to the maximum amplitude pseudo-spectrum value in the amplitude pseudo-spectrum centered on the index of the current frame audio fundamental frequency with a width of one time the index of the current frame audio fundamental frequency is determined as the index integer part of the second harmonic of the current frame audio fundamental frequency.
6. The Bluetooth audio transposing method of claim 2, wherein the current frame audio harmonics include the second, third, fourth and fifth harmonics of the current frame audio fundamental frequency. The current frame audio fundamental frequency is obtained by using the amplitude pseudo-spectrum of the current frame audio spectrum coefficients.
7. A Bluetooth audio transposing apparatus, characterized by, The amplitude pseudo-spectrum calculation module is configured to calculate the amplitude pseudo-spectrum corresponding to each of the current frame audio spectrum coefficients by using the each of the current frame audio spectrum coefficients, the predetermined number of adjacent spectrum coefficients before the each of the current frame audio spectrum coefficients, and the predetermined number of adjacent spectrum coefficients after the each of the current frame audio spectrum coefficients in the process of Bluetooth audio encoding and / or decoding. The fundamental frequency obtaining module is configured to obtain the current frame audio fundamental frequency by using the amplitude pseudo-spectrum of the current frame audio spectrum coefficients to calculate the fundamental frequency of the current frame audio signal corresponding to the current frame audio spectrum coefficients. The transposing processing module is configured to perform phase shifting processing on the phase of the current frame audio fundamental frequency, and perform ascending / descending processing on the frequency of the current frame audio fundamental frequency to obtain the transposed current frame audio fundamental frequency. The post-processing module is configured to generate the transposed current frame audio spectrum coefficients by using the transposed current frame audio fundamental frequency, and complete the remaining steps of Bluetooth audio encoding and / or decoding by using the transposed current frame audio spectrum coefficients. The post-processing module is further configured to generate the transposed current frame audio spectrum coefficients by using the transposed current frame audio fundamental frequency and the current frame audio harmonics. The fundamental frequency obtaining module includes: The subband index of the subband to which the frequency corresponding to the maximum amplitude pseudo-spectrum value in the amplitude pseudo-spectrum centered on the index of the current frame audio fundamental frequency with a width of one time the index of the current frame audio fundamental frequency is determined as the index integer part of the second harmonic of the current frame audio fundamental frequency. The index fractional part of the current frame audio fundamental frequency is calculated by using the ratio of the amplitude value of the subband before the subband to which the frequency corresponding to the maximum amplitude pseudo-spectrum value belongs, and the amplitude value of the subband after the subband to which the frequency corresponding to the maximum amplitude pseudo-spectrum value belongs, or the ratio of the amplitude value of the second subband before the subband to which the frequency corresponding to the maximum amplitude pseudo-spectrum value belongs, and the amplitude value of the second subband after the subband to which the frequency corresponding to the maximum amplitude pseudo-spectrum value belongs. The computer instructions are operated to perform the Bluetooth audio transposing method of any one of claims 1-6.
9. A Bluetooth device comprising an encoder and a decoder, wherein the encoder and / or the decoder is provided with the Bluetooth audio transposing apparatus of claim 7. 8. A computer-readable storage medium storing computer instructions, wherein,
Citation Information
Patent Citations
Audio error hiding method and device
CN112992160A
Fundamental frequency extraction model training method and device and fundamental frequency extraction method and device
CN114067784A