Decoding method and apparatus, encoding method and apparatus, electronic device, medium, product and system

By analyzing the high-frequency harmonic level identification and fundamental pitch period, combining the first frequency band signal, the second frequency band signal is determined, and the problem of inaccurate recovery of high-frequency signals in low-code rate encoding is solved, and the accurate recovery of high-frequency signals and the improvement of audio quality is achieved.

WO2025162259A1PCT designated stage Publication Date: 2025-08-07BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/074704
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-02-02
Filing Date
2025-01-24
Publication Date
2025-08-07

AI Technical Summary

Technical Problem

In low-code encoding, high-frequency signals cause spectrum holes due to insufficient bit allocation. The prior art restored signal accuracy through filling noise or spectrum folding is poor, affecting the hearing feeling.

Method used

By analyzing the high-frequency harmonic level identification and the pitch period, combining the first frequency band signal, the second frequency band signal is determined, and the matching signal is used to copy the matching signal to the second frequency band to restore the high-frequency signal and avoid spectral holes.

Benefits of technology

It improves the recovery accuracy of high-frequency signals, improves the overall quality of audio decoding, reduces noise interference, and improves the listening experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025074704_07082025_PF_FP_ABST
    Figure CN2025074704_07082025_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure relates to a decoding method and apparatus, an encoding method and apparatus, an electronic device, a medium, a product and a system. The decoding method of the present disclosure comprises: parsing a code stream to obtain a high-frequency harmonic wave level identifier and a signal of a first frequency band of an audio frame, wherein the high-frequency harmonic wave level identifier is used for indicating the harmonic wave level of a second frequency band of the audio frame, and the frequency of the second frequency band is higher than that of the first frequency band; determining a signal of the second frequency band on the basis of a pitch period of the audio frame, the high-frequency harmonic wave level identifier, and the signal of the first frequency band; and obtaining a decoded signal of the audio frame on the basis of the signal of the first frequency band and the signal of the second frequency band.
Need to check novelty before this filing date? Find Prior Art

Description

Decoding and encoding methods, devices, electronic devices, media, products and systems

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application is based on and claims priority to an application filed in China with application number 202410154599.8 and filing date February 2, 2024. The disclosure of the application in China is hereby incorporated as a whole into this application. Technical Field

[0003] The present disclosure relates to the field of audio processing technology, and in particular to a decoding and encoding method, device, electronic device, medium, product and system. Background Art

[0004] In audio coding, encoders typically use a higher bitrate to encode low-frequency signals, to which the human ear is more sensitive, and a lower bitrate to encode high-frequency signals, to which the human ear is less sensitive. When encoding at low bitrates, the bit allocation strategy often doesn't allocate enough bits to the high-frequency band, making high-frequency encoding impossible. This results in spectral holes in the decoded high frequencies, affecting the subjective listening experience.

[0005] In related technologies, spectrum holes can be filled in the following two ways: (1) filling with random noise: using a random number generator to generate noise and normalizing it as the spectrum of the high-frequency hole; (2) spectrum folding: copying the normalized spectrum of the low-frequency band to the high-frequency hole. Summary of the Invention

[0006] According to some embodiments of the present disclosure, a decoding method is provided, comprising: parsing a high-frequency harmonic level identifier and a signal of a first frequency band of an audio frame from a bit stream, wherein the high-frequency harmonic level identifier is used to indicate the harmonic level of a second frequency band of the audio frame, and the frequency of the second frequency band is higher than that of the first frequency band; determining the signal of the second frequency band based on the fundamental frequency period of the audio frame, the high-frequency harmonic level identifier and the signal of the first frequency band; and obtaining a decoded signal of the audio frame based on the signal of the first frequency band and the signal of the second frequency band.

[0007] According to other embodiments of the present disclosure, an encoding method is provided, comprising: obtaining a high-frequency harmonic level identifier and a signal of a first frequency band of an audio frame based on an input audio frame, wherein the high-frequency harmonic level identifier is used to indicate the harmonic level of a second frequency band of the audio frame, and the high-frequency harmonic level identifier is used together with the fundamental frequency period of the audio frame and the signal of the first frequency band to determine the signal of the second frequency band, the frequency of the second frequency band being higher than that of the first frequency band; and generating a code stream based on the high-frequency harmonic level identifier and the signal of the first frequency band.

[0008] According to some further embodiments of the present disclosure, a decoding device is provided, including: a first decoding module, configured to parse out a high-frequency harmonic level identifier and a signal of a first frequency band of an audio frame from a bit stream, wherein the high-frequency harmonic level identifier is used to indicate the harmonic level of a second frequency band of the audio frame, and the frequency of the second frequency band is higher than that of the first frequency band; a second decoding module, configured to determine the signal of the second frequency band based on the fundamental period of the audio frame, the high-frequency harmonic level identifier and the signal of the first frequency band; and an audio recovery module, configured to obtain a decoded signal of the audio frame based on the signal of the first frequency band and the signal of the second frequency band.

[0009] According to some further embodiments of the present disclosure, an encoding device is provided, including: an acquisition module, configured to acquire a high-frequency harmonic level identifier and a signal of a first frequency band of an audio frame based on an input audio frame, wherein the high-frequency harmonic level identifier is used to indicate the harmonic level of a second frequency band of the audio frame, and the high-frequency harmonic level identifier is used together with the fundamental frequency period of the audio frame and the signal of the first frequency band to determine the signal of the second frequency band, and the frequency of the second frequency band is higher than that of the first frequency band; and a generation module, configured to generate a code stream based on the high-frequency harmonic level identifier and the signal of the first frequency band.

[0010] According to some further embodiments of the present disclosure, an electronic device is provided, comprising: a memory; and a processor coupled to the memory, the processor being configured to execute the decoding method of any embodiment of the present disclosure and / or the encoding method of any embodiment of the present disclosure based on instructions stored in the memory.

[0011] According to some further embodiments of the present disclosure, a computer-readable storage medium is provided, on which a computer program is stored. When the program is executed by a processor, the decoding method of any embodiment of the present disclosure and / or the encoding method of any embodiment of the present disclosure is performed.

[0012] According to some further embodiments of the present disclosure, a computer program product is provided, comprising: instructions, which, when executed by a processor, implement the decoding method of any embodiment of the present disclosure and / or the encoding method of any embodiment of the present disclosure.

[0013] According to some further embodiments of the present disclosure, there is provided an audio processing system, comprising: a decoding device according to any embodiment of the present disclosure and an encoding device according to any embodiment of the present disclosure.

[0014] According to some further embodiments of the present disclosure, a computer program is provided, comprising: instructions, which, when executed by a processor, implement the decoding method of any embodiment of the present disclosure and / or the encoding method of any embodiment of the present disclosure.

[0015] Other features, aspects and advantages of the present disclosure will become apparent from the following detailed description of exemplary embodiments of the present disclosure with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] The following describes embodiments of the present disclosure with reference to the accompanying drawings. It should be understood that the drawings described below only relate to some embodiments of the present disclosure and do not constitute a limitation to the present disclosure. In the accompanying drawings:

[0017] FIG1 is a schematic diagram showing a flow chart of a decoding method according to some embodiments of the present disclosure;

[0018] FIG2 is a schematic diagram showing a sub-band duplication start frequency according to some embodiments of the present disclosure;

[0019] FIG3 is a schematic diagram showing a code stream format according to some embodiments of the present disclosure;

[0020] FIG4 is a schematic diagram showing a decoding process according to some embodiments of the present disclosure;

[0021] FIG5 is a schematic diagram showing a flow chart of an encoding method according to some embodiments of the present disclosure;

[0022] FIG6 is a schematic diagram showing frequency band division according to some embodiments of the present disclosure;

[0023] FIG7 is a schematic diagram showing an encoding process according to some embodiments of the present disclosure;

[0024] FIG8 is a schematic structural diagram of a decoding device according to some embodiments of the present disclosure;

[0025] FIG9 is a schematic structural diagram of an encoding device according to some embodiments of the present disclosure;

[0026] FIG10 is a schematic structural diagram of an electronic device according to some embodiments of the present disclosure;

[0027] FIG11 is a schematic structural diagram of an electronic device according to some other embodiments of the present disclosure;

[0028] FIG12 shows a schematic structural diagram of an audio processing system according to some embodiments of the present disclosure. DETAILED DESCRIPTION

[0029] The following will be combined with the accompanying drawings in the embodiments of the present disclosure to clearly and completely describe the technical solutions in the embodiments of the present disclosure. It should be understood that the present disclosure can be implemented in various forms and should not be interpreted as being limited to the embodiments described herein.

[0030] It should be understood that the various steps described in the method embodiments of the present disclosure can be performed in different orders and / or performed in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this respect. Unless otherwise specifically stated, the relative arrangement, numerical expressions, and numerical values ​​of the steps set forth in these embodiments should be interpreted as being merely exemplary and do not limit the scope of the present disclosure.

[0031] The term “including” and its variations used in the present disclosure are open terms that include at least the following elements / features but do not exclude other elements / features, that is, “including but not limited to.” The term “based on” means “at least in part based on.”

[0032] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are only used to distinguish different devices, modules, or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules, or units. Unless otherwise specified, concepts such as "first" and "second" are not intended to imply that the objects described in such a manner must be in a given order in time, space, ranking, or any other manner.

[0033] It should be noted that the modifications of "one" and "multiple" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, they should be understood as "one or more".

[0034] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.

[0035] The following detailed description of the embodiments of the present disclosure is provided in conjunction with the accompanying drawings, but the present disclosure is not limited to these specific embodiments. The following specific embodiments may be combined with each other, and the same or similar concepts or processes may not be described in detail in some embodiments. In addition, in one or more embodiments, specific features, structures, or characteristics may be combined in any suitable manner that will be apparent to those skilled in the art from this disclosure.

[0036] The inventors found that in the case of low-bit-rate encoding, the decoding end fills the spectrum holes caused by insufficient bit allocation by filling noise or spectrum folding. Not only can obvious high-frequency noise be heard in the recovered signal, but the filled noise and the folded high-frequency part are often too different from the original signal. Especially when the harmonic level of the high-frequency part is high, this difference will be more obvious, resulting in poor accuracy of the recovered audio signal and a poor listening experience.

[0037] In order to improve the accuracy of audio signal restoration at the decoding end, the present disclosure proposes a decoding method, which is described below with reference to FIG. 1 to FIG. 4 .

[0038] Figure 1 is a flowchart of some embodiments of the decoding method disclosed herein. As shown in Figure 1, the method of this embodiment includes steps S102 to S106. The decoding method disclosed herein can be performed by a decoding device (decoding end), which can be a decoder or a device with wireless transceiver capabilities, such as a mobile phone, a tablet computer, a computer with wireless transceiver capabilities, a virtual reality (VR) device, an augmented reality (AR) device, etc., without limitation to the examples given.

[0039] In step S102, a high-frequency harmonic level indicator and a signal of a first frequency band of an audio frame are parsed from the bit stream.

[0040] The decoding device (decoding end) can receive a code stream obtained by encoding the audio frame. For example, the code stream includes code stream data for a first frequency band, code stream data for a second frequency band, and a high-frequency harmonic level indicator. The frequency of the second frequency band is higher than that of the first frequency band, that is, the lowest frequency of the second frequency band is higher than the highest frequency of the first frequency band.

[0041] In some embodiments, the bitstream data of the first frequency band is obtained by encoding the signal of the first frequency band of the audio frame, and the bitstream data of the second frequency band is obtained by encoding the signal of the second frequency band of the audio frame. For example, the bitstream data of the second frequency band is side information of the signal of the second frequency band.

[0042] In some embodiments, the code stream data of the first frequency band is decoded to obtain a signal of the first frequency band.

[0043] The high-frequency harmonic level identifier is used to indicate the harmonic level of the second frequency band of the audio frame. For example, the high-frequency harmonic level identifier is the harmonic amplitude of the second frequency band; or the harmonic level is divided into different levels, and the high-frequency harmonic level identifier is the level corresponding to the harmonic of the second frequency band; or the high-frequency harmonic level identifier is a parameter in the existing coding and decoding method that can reflect the harmonic level, etc., and is not limited to the examples given. For example, when the CELT (Constrain Edenergy Lapped Transform) coding and decoding mode is adopted, the high-frequency harmonic level identifier is the tapset (filter group) parameter value, and the tapset parameter value is used to represent the 5-tap filter coefficient, which can reflect the high-frequency harmonic level.

[0044] In step S104, a signal in the second frequency band is determined according to the pitch period of the audio frame, the high frequency harmonic level indicator, and the signal in the first frequency band.

[0045] In some embodiments, the bitstream includes a pitch period, and / or the pitch period is obtained from a signal in a first frequency band. The pitch period can be obtained by analyzing the signal in the first frequency band using an existing algorithm.

[0046] The pitch period can be directly carried in the bitstream, or it can be obtained by analyzing the signal in the first frequency band. When the pitch period is carried in the bitstream, the pitch period can be represented by a parameter that can reflect the pitch period in existing coding methods, without limitation to the examples given. For example, when the CELT coding mode is used, the pitch period is represented by the period parameter value.

[0047] In some embodiments, when the harmonic level of the second frequency band indicated by the high-frequency harmonic level indicator is higher than a threshold, the signal of the second frequency band is determined according to the pitch period and the signal of the first frequency band.

[0048] The high-frequency harmonic level indicator can be directly compared with the threshold to determine whether it is above the threshold. For example, if the high-frequency harmonic level indicator is a tapset parameter value, if the tapset parameter value is higher than its corresponding threshold, it means that the harmonic level of the second frequency band is high. In this case, the signal of the second frequency band is determined based on the fundamental pitch period and the signal of the first frequency band.

[0049] In some embodiments, the signal of the first frequency band is copied to the second frequency band according to the pitch period, and the signal of the second frequency band is restored according to the copied signal.

[0050] If the harmonic level of the second frequency band is higher than the threshold, it indicates that the harmonic level of the second frequency band is relatively high. If a segment of the signal in the first frequency band is arbitrarily copied to the second frequency band, it is likely that the copied signal will be significantly different from the harmonics in the second frequency band, resulting in inaccurate audio frame recovery. The pitch period can represent the pitch frequency. According to the characteristics of the audio signal, the waveform corresponding to the pitch frequency or an integer multiple of the pitch frequency is closer to the waveform of the harmonics in the high frequency band. Therefore, based on the pitch period and the signal in the first frequency band, a portion that better matches the original signal in the second frequency band can be selected and copied to the second frequency band to recover the signal in the second frequency band, thereby improving the accuracy of audio frame decoding.

[0051] In step S106 , a decoded signal of the audio frame is obtained according to the signal of the first frequency band and the signal of the second frequency band.

[0052] For example, the signal of the first frequency band and the signal of the second frequency band are spliced ​​to obtain a restored audio frame.

[0053] According to the method disclosed herein, a high-frequency harmonic level indicator and a low-frequency first-band signal of an audio frame can be parsed from a bitstream. Furthermore, based on the audio frame's pitch period, the high-frequency harmonic level indicator, and the signal of the first-band, a high-frequency second-band signal can be recovered, thereby recovering an audio frame based on the signal of the first-band and the signal of the second-band. By recovering the signal of the second-band based on the high-frequency harmonic level indicator, the pitch period, and the signal of the first-band, and taking into account the characteristics of the audio signal, the recovered signal of the second-band can be made more accurate, thereby improving the accuracy of audio frame decoding.

[0054] The following describes in detail some embodiments of how to determine the signal of the second frequency band based on the pitch period and the signal of the first frequency band.

[0055] In some embodiments, for each sub-band, the signal of the second frequency band is determined according to the pitch period, the start frequency of the sub-band, the frequency of the first harmonic component after the start frequency of the sub-band, and the signal of the first frequency band.

[0056] In some embodiments, for each subband, the signal of the second frequency band is determined based on the pitch period, the start frequency of the subband, the frequency of the first harmonic component after the start frequency of the subband, the width of the subband, and the signal of the first frequency band.

[0057] In some embodiments, a subband copy starting frequency of each subband in one or more subbands in the second frequency band is determined based on the fundamental period; a signal in the first frequency band corresponding to the subband copy starting frequency of each subband is copied to each subband; and a signal of the second frequency band is determined based on side information of the signal of each subband and the signal of the second frequency band.

[0058] The second frequency band can be divided into one or more subbands, and the division of the subbands in the second frequency band can also be determined based on the coding scheme of the second frequency band. The subband division method can refer to the existing technology and is not further described here. For each subband, based on the subband replication start frequency and the width of the subband, the signal to be replicated corresponding to the subband is determined in the first frequency band (i.e., the signal in the first frequency band corresponding to the subband replication start frequency of the subband), and then the signal to be replicated corresponding to the subband is copied to the subband.

[0059] For example, the side information of the signal of the second frequency band includes information such as envelope energy and harmonic level of the signal of the second frequency band, but is not limited to the examples given. The signal of each sub-band can be adjusted (shaped) according to the side information to determine the signal of the second frequency band.

[0060] In some embodiments, for each subband, the frequency of the first harmonic component after the start frequency of the subband is determined according to the pitch period; and the subband copy start frequency of the subband is determined according to the pitch period, the start frequency of the subband and the frequency of the first harmonic component.

[0061] According to the characteristics of the audio signal, for each sub-band, if the signal corresponding to the fundamental frequency or an integer multiple of the fundamental frequency in the first frequency band can be copied to the frequency of the first harmonic component after the sub-band, the signal of the second frequency band can be restored more accurately.

[0062] In some embodiments, for each sub-band, a reference frequency corresponding to the sub-band is determined based on the fundamental period, the starting frequency of the sub-band and the frequency of the first harmonic component, and a bandwidth between the reference frequency corresponding to the sub-band and the starting frequency of the first sub-band in the second frequency band is determined; when the bandwidth is not less than the width of the sub-band, the reference frequency corresponding to the sub-band is determined as the sub-band copy starting frequency of the sub-band.

[0063] In some embodiments, for each sub-band, a first frequency difference between the frequency of the first harmonic component and the starting frequency of the sub-band is determined; based on the fundamental frequency period, the fundamental frequency is determined; a second frequency difference between a preset multiple value of the fundamental frequency and the first frequency difference is determined as the reference frequency corresponding to the sub-band; a bandwidth between the reference frequency corresponding to the sub-band and the starting frequency of the first sub-band in the second frequency band is determined; and when the bandwidth is not less than the width of the sub-band, the reference frequency corresponding to the sub-band is determined as the sub-band copy starting frequency of the sub-band.

[0064] In some embodiments, the preset multiple is two.

[0065] In some embodiments, it is determined whether the first frequency difference is greater than the fundamental frequency, and if so, the preset multiple is greater than 1. Preferably, the preset multiple is 2. If the first frequency difference is less than or equal to the fundamental frequency, the preset multiple may be 1.

[0066] In some embodiments, the reference frequency corresponding to each sub-band may be determined using the following formula: Among them, freq pitch represents the fundamental frequency determined by the fundamental pitch period, Indicates the starting frequency of the i-th subband, i is an integer greater than 0, k i *freq pitch Indicates the frequency of the first harmonic component after the starting frequency of the i-th sub-band, k i is an integer greater than 0, a is a preset value, and offset is an offset value; if the bandwidth between the reference frequency corresponding to the sub-band and the starting frequency of the first sub-band in the second frequency band is not less than the width of the sub-band, the reference frequency corresponding to the sub-band is determined as the sub-band copy starting frequency of the sub-band. The high-frequency first harmonic component and k1*freq pitchThere may be a deviation between them, and offset adjusts the deviation to improve the accuracy of the sub-band copy starting frequency.

[0067] For example, a = 2. According to the characteristics of the audio signal, the harmonic component of the second frequency band is an integer k of the fundamental frequency. i For example, for each sub-band, the starting frequency of the sub-band can be divided by the fundamental frequency, and the obtained value can be rounded up to get k i Other methods can also be used to determine k i .

[0068] As shown in Figure 2, the second frequency band is divided into four sub-bands: sub1, sub2, sub3, and sub4. The starting frequency of each sub-band is freq_start1, freq_start2, freq_start3, and freq_start4. The doubled fundamental frequency is 2*freq_pitch. The sub-band copy starting frequency of the first sub-band is freq_cpy1. The bandwidth from freq_cpy1 to 2*freq_pitch is equal to the bandwidth from freq_start1 to the frequency of the first harmonic component after freq_start1. In this way, when copying, the waveform of 2*freq_pitch can be copied to the frequency of the first harmonic component after freq_start1, making the copied waveform closer to the original waveform of the second frequency band, thereby improving the accuracy of signal recovery in the second frequency band.

[0069] To prevent the copied signal from creating a band hole in the second frequency band, it's necessary to determine whether the bandwidth between the reference frequency corresponding to each subband and the start frequency of the first subband in the second frequency band is less than the width of each subband. For example, as shown in Figure 2, the width between freq_cpy1 and freq_start1 is less than the width of sub1. If so, using freq_cpy1 as the subband copy start frequency will result in a band hole in sub1 after the copy. Therefore, if the width between freq_cpy1 and freq_start1 is not less than the width of sub1, the reference frequency corresponding to subband sub1, as determined above, is determined as the subband copy start frequency for subband sub1.

[0070] In some embodiments, for each subband, if the bandwidth between the reference frequency corresponding to the subband and the starting frequency of the first subband in the second frequency band is less than the width of the subband, a subband copy starting frequency for the subband is determined based on the coding scheme of the second frequency band. Furthermore, the signal corresponding to the subband copy starting frequency of each subband in the first frequency band is copied to each subband; and the signal of the second frequency band is determined based on side information between the signal of each subband and the signal of the second frequency band.

[0071] For example, the second frequency band adopts the BWE (Bandwidth Extension) coding mode, which can correspond to multiple specific coding schemes. For example, IGF (Intelligent Gap Filling) is a scheme in the BWE coding mode, and SBR (Spectral Band Replication) is another scheme in the BWE coding mode, which is not limited to the examples given. Different coding schemes correspond to different methods for determining the sub-band replication starting frequency (or sub-band replication range). For example, the sub-band replication starting frequency of each sub-band in IGF adopts a fixed preset value, and the sub-band replication starting frequency of each sub-band in SBR is obtained by subtracting the width of the sub-band from the starting frequency of the sub-band. The method for determining the sub-band replication starting frequency of other coding schemes can be specifically referred to the existing technology and will not be repeated here.

[0072] If, for different sub-bands, the bandwidth between the reference frequency corresponding to some sub-bands and the starting frequency of the first sub-band in the second frequency band is smaller than the width of the sub-band, and the bandwidth between the reference frequency corresponding to some sub-bands and the starting frequency of the first sub-band in the second frequency band is not smaller than the width of the sub-band, different methods can be used to determine the sub-band copy starting frequency of each sub-band, and the details will not be repeated here.

[0073] The method of the above embodiment determines the subband replication start frequency for each subband based on the pitch period, the start frequency of each subband, and the frequency of the first harmonic component following the start frequency of each subband. This allows accurate selection of the low-frequency signal to be replicated in each subband, improving the accuracy of high-frequency (second frequency band) signal recovery. Furthermore, by considering the bandwidth of each subband, high-frequency holes are avoided, further improving the accuracy of high-frequency signal recovery, thereby improving the overall accuracy of audio frame decoding.

[0074] In some embodiments, step S104 may be replaced by using the high-frequency harmonic level identifier and the signal of the first frequency band to determine the signal of the second frequency band.

[0075] In some embodiments, when the harmonic level of the second frequency band indicated by the high-frequency harmonic level indicator is higher than a threshold, the signal of the second frequency band is determined according to the pitch period and the signal of the first frequency band.

[0076] In some embodiments, when the harmonic level of the second frequency band indicated by the high-frequency harmonic level identifier is not higher than a threshold, the signal of the second frequency band is determined according to the coding scheme of the second frequency band and the signal of the first frequency band.

[0077] If the harmonic level of the second frequency band is not higher than the threshold, the signal of the second frequency band can be directly restored based on the coding scheme of the second frequency band and the signal of the first frequency band. For example, the second frequency band adopts the BWE coding mode, which can correspond to multiple specific coding schemes. For details on how to determine the sub-band replication start frequency (or sub-band replication range) corresponding to different coding schemes, please refer to the existing technology.

[0078] For details on how to determine the signal of the second frequency band based on the pitch period and the signal of the first frequency band, and how to determine the signal of the second frequency band based on the coding scheme of the second frequency band and the signal of the first frequency band, please refer to the above embodiments and will not be repeated here.

[0079] The signals of the first frequency band and the second frequency band may use different encoding and decoding methods. In some embodiments, the code stream data of the first frequency band in the code stream is decoded using a first decoding method to obtain the signal of the first frequency band, and the code stream data of the second frequency band is decoded using a second decoding method to obtain side information of the signal of the second frequency band.

[0080] In some embodiments, the signal of the first frequency band is obtained by decoding the code stream data of the first frequency band in the code stream using a CELT decoding mode.

[0081] In some embodiments, the side information of the second frequency band is obtained by decoding the code stream data of the second frequency band in the code stream using a BWE decoding mode.

[0082] As shown in Figure 3, the payload of the data packet in the code stream consists of two parts. On the basis of the original CELT payload format (opus (CELT) payload format), a supplementary payload (padding payload) is added to carry the BWE payload (BWE payload), and it can also carry the length information of the BWE payload (BWE payload len). The CELT payload part includes the code stream data of the first frequency band, and the BWE payload part includes the code stream data of the second frequency band.

[0083] During decoding, the code stream data for the first and second frequency bands can be obtained from the payload portion of the data packet in the code stream and decoded separately. The CELT payload or BWE payload can carry high-frequency harmonic level identifiers, pitch period, and range information for the first frequency band to facilitate decoding.

[0084] In some embodiments, the envelope of the signal of each sub-band is adjusted based on the side information of the signal of the second frequency band; the signal in the adjusted second frequency band is post-processed and inverse MDCT (Modified Discrete Cosine Transform) is performed to obtain a restored signal of the second frequency band; the signal of the first frequency band and the restored signal of the second frequency band are spliced ​​to obtain a restored audio frame.

[0085] For example, as shown in FIG4 , the bitstream data of the first frequency band and the bitstream data of the second frequency band in the bitstream can be decoded separately. The bitstream data of the first frequency band is decoded using a CELT decoder, while the bitstream data of the second frequency band is decoded using a BWE decoder. The BWE decoding process includes decoding the side information, determining the subband replication start frequency based on the pitch period, copying the low-frequency (first frequency band) signal to the high-frequency (second frequency band), calculating the high-frequency energy loss, adjusting the high-frequency envelope based on the side information, performing post-processing such as time domain inverse filtering and inverse MDCT, and obtaining the high-frequency (second frequency band) signal. The high-frequency (second frequency band) signal is then concatenated with the low-frequency (first frequency band) signal to output the audio signal.

[0086] The present disclosure further provides an encoding method. Some embodiments of the encoding method of the present disclosure are described below in conjunction with FIG5 .

[0087] FIG5 is a flowchart of some embodiments of the encoding method disclosed herein. As shown in FIG4 , the method of this embodiment includes steps S502 to S504. The encoding method disclosed herein can be performed by an encoding device (encoding end), which can be an encoder or a device with wireless transceiver capabilities, such as a mobile phone, a tablet computer, a computer with wireless transceiver capabilities, a virtual reality (VR) device, an augmented reality (AR) device, etc., without limitation to the examples given.

[0088] In step S502 , a high-frequency harmonic level indicator and a signal of a first frequency band of the audio frame are obtained according to the input audio frame.

[0089] For example, the high-frequency harmonic level identifier is used to indicate the harmonic level of the second frequency band of the audio frame. For example, the high-frequency harmonic level identifier can be the harmonic amplitude of the second frequency band; or the harmonic levels can be divided into different levels, and the high-frequency harmonic level identifier can be the level corresponding to the harmonics of the second frequency band; or the high-frequency harmonic level identifier can be a parameter in an existing coding method that can reflect the harmonic level, etc., without limitation to the examples given.

[0090] The high frequency harmonic level identifier may be used together with the pitch period of the audio frame and the signal of the first frequency band to determine the signal of the second frequency band.

[0091] The second frequency band has a higher frequency than the first frequency band. For example, the division of the first frequency band and the second frequency band is determined according to a target bit rate for encoding audio frames, wherein the lower the target bit rate, the smaller the range of the first frequency band and the larger the range of the second frequency band.

[0092] For example, the target bit rate is divided into different ranges, each range corresponding to a division method of the first frequency band and the second frequency band. At the codec end, the ranges of the first frequency band and the second frequency band can be determined according to the target bit rate.

[0093] As shown in Figure 6, when the target bit rate is 16-20 kbps, the first frequency band is 0-6.8 kHz and the second frequency band is 6.8 kHz-20 kHz; when the target bit rate is 20-24 kbps, the first frequency band is 0-8 kHz and the second frequency band is 8 kHz-20 kHz; when the target bit rate is 24-28 kbps, the first frequency band is 0-9.6 kHz and the second frequency band is 9.6 kHz-20 kHz; when the target bit rate is 28-32 kbps, the first frequency band is 0-12 kHz and the second frequency band is 12 kHz-20 kHz.

[0094] In some embodiments, the code stream includes a pitch period.

[0095] In step S504, a code stream is generated according to the high-frequency harmonic level indicator and the signal of the first frequency band.

[0096] In some embodiments, a signal of a first frequency band is encoded to obtain code stream data of the first frequency band; side information of a second frequency band is determined as code stream data of the second frequency band; and a high-frequency harmonic level identifier, the code stream data of the first frequency band, and the code stream data of the second frequency band are written into the code stream.

[0097] In some embodiments, the code stream data of the first frequency band is obtained by encoding the signal of the first frequency band in the audio frame using the CELT coding mode; and the side information of the second frequency band is obtained by encoding the signal of the second frequency band using the BWE coding mode.

[0098] The CELT encoder and the BWE encoder can be used to encode the signal of the first frequency band and the signal of the second frequency band respectively. The encoded code stream data of the first frequency band and the code stream data of the second frequency band can be transmitted in the data packet format shown in Figure 3, which will not be repeated here.

[0099] The CELT payload also carries CELT codestream structure information, which can include high-frequency harmonic level indicators and pitch period. For example, the high-frequency harmonic level indicator is a tapset parameter value, which represents the coefficients of a 5-tap filter and reflects the high-frequency harmonic level. Another example is the pitch period, which is represented by a period parameter value.

[0100] In some embodiments, after pre-processing and MDCT are performed on the audio frame, the signal of the second frequency band in the audio frame is encoded using a BWE coding mode.

[0101] For example, as shown in FIG7 , the bit rate is input into the CELT encoder, and after coarse code control processing, the audio signal (audio frame) is input and pre-processed by pre-emphasis, pre-filtering, signal type detection, and then MDCT (encryption and band energy calculation can also be performed). After that, the high-frequency (second frequency band) signal can be BWE encoded, while the low-frequency (first frequency band) signal continues to be CELT encoded. For example, the low-frequency (first frequency band) signal continues to undergo energy normalization, energy dynamics analysis, energy coarse quantization, time-frequency dynamics encoding, stereo analysis, and other processing processes. BWE encoding can be performed in parallel with any processing process after MDCT, or it can be performed before any processing process after MDCT.

[0102] The method of the above embodiment writes the high-frequency harmonic level identifier into the code stream. The high-frequency harmonic level identifier can be used together with the fundamental frequency period and the signal of the first frequency band to determine the signal of the second frequency band. In this way, the code stream can only carry a small amount of information of the second frequency band, thereby improving transmission efficiency and being suitable for scenarios with low target bit rates. In addition, the signal of the second frequency band is restored based on the high-frequency harmonic level identifier, the fundamental frequency period and the signal of the first frequency band. Considering the characteristics of the audio signal, the restored signal of the second frequency band can be made more accurate, thereby improving the accuracy of audio frame decoding. In addition, combining the two encoding methods of CELT and BWE can solve the problem that the high frequency band in CELT encoding often cannot be allocated enough bits, and the high frequency cannot be encoded with high quality, resulting in spectrum holes in the decoded high frequency, thereby improving the encoding quality.

[0103] The present disclosure also provides a decoding device, which will be described below in conjunction with FIG8 .

[0104] FIG8 is a structural diagram of some embodiments of a decoding device according to the present disclosure. As shown in FIG8 , a decoding device 80 according to this embodiment includes: a first decoding module 810 , a second decoding module 820 , and an audio recovery module 830 .

[0105] The first decoding module 810 is configured to parse out a high-frequency harmonic level identifier and a signal of a first frequency band of an audio frame from a bit stream, wherein the high-frequency harmonic level identifier is used to indicate a harmonic level of a second frequency band of an audio frame, and the frequency of the second frequency band is higher than that of the first frequency band.

[0106] The second decoding module 820 is configured to determine a signal in a second frequency band according to the pitch period of the audio frame, the high frequency harmonic level indicator, and the signal in the first frequency band.

[0107] In some embodiments, the second decoding module 820 is configured to determine a signal in the second frequency band according to the pitch period and the signal in the first frequency band when the harmonic level of the second frequency band indicated by the high-frequency harmonic level indicator is higher than a threshold.

[0108] In some embodiments, the second decoding module 820 is configured to determine the subband copy starting frequency of each subband in one or more subbands in the second frequency band based on the fundamental period; copy the signal corresponding to the subband copy starting frequency of each subband in the first frequency band to each subband; and determine the signal of the second frequency band based on the side information of the signal of each subband and the signal of the second frequency band.

[0109] In some embodiments, the second decoding module 820 is configured to determine, for each sub-band, the frequency of the first harmonic component after the starting frequency of the sub-band based on the fundamental period; and determine the sub-band copy starting frequency of the sub-band based on the fundamental period, the starting frequency of the sub-band and the frequency of the first harmonic component.

[0110] In some embodiments, the second decoding module 820 is configured to determine the reference frequency corresponding to the sub-band according to the following formula: Among them, freq pitch represents the fundamental frequency determined by the fundamental pitch period, Indicates the starting frequency of the i-th subband, i is an integer greater than 0, k i *freq pitch Indicates the frequency of the first harmonic component after the starting frequency of the i-th sub-band, k i is an integer greater than 0, a is a preset value, and offset is an offset value; when the bandwidth between the reference frequency corresponding to the subband and the starting frequency of the first subband in the second frequency band is not less than the width of the subband, the reference frequency corresponding to the subband is determined as the subband copy starting frequency of the subband.

[0111] In some embodiments, the second decoding module 820 is further configured to determine the sub-band copy starting frequency of the sub-band according to the coding scheme of the second frequency band when the frequency band width is smaller than the width of the sub-band.

[0112] In some embodiments, the second decoding module 820 is further configured to determine the sub-band copy starting frequency of each sub-band in one or more sub-bands in the second frequency band according to the coding scheme of the second frequency band when the harmonic level of the second frequency band indicated by the high-frequency harmonic level identifier is not higher than the threshold; copy the signal corresponding to the sub-band copy starting frequency of each sub-band in the first frequency band to each sub-band; and determine the signal of the second frequency band based on the side information of the signal of each sub-band and the signal of the second frequency band.

[0113] The audio restoration module 830 is configured to obtain a decoded signal of the audio frame according to the signal of the first frequency band and the signal of the second frequency band.

[0114] In some embodiments, the bitstream includes a pitch period.

[0115] In some embodiments, the pitch period is obtained from a signal in a first frequency band.

[0116] In some embodiments, the signal of the first frequency band is obtained by decoding the code stream data of the first frequency band in the code stream using a Constrained Energy Overlap Transform (CELT) decoding mode.

[0117] In some embodiments, the side information of the second frequency band is obtained by decoding the code stream data of the second frequency band in the code stream using a bandwidth extension (BWE) decoding mode.

[0118] The present disclosure also provides an encoding device, which is described below in conjunction with FIG9 .

[0119] FIG9 is a structural diagram of some embodiments of the encoding device of the present disclosure. As shown in FIG9 , the encoding device 90 of this embodiment includes: an acquisition module 910 and a generation module 920 .

[0120] The acquisition module 910 is configured to obtain the high-frequency harmonic level identifier and the signal of the first frequency band of the audio frame based on the input audio frame, wherein the high-frequency harmonic level identifier is used to indicate the harmonic level of the second frequency band of the audio frame, and the high-frequency harmonic level identifier is used together with the fundamental period of the audio frame and the signal of the first frequency band to determine the signal of the second frequency band, and the frequency of the second frequency band is higher than that of the first frequency band.

[0121] The generating module 920 is configured to generate a code stream according to the high-frequency harmonic level identifier and the signal of the first frequency band.

[0122] In some embodiments, the generation module 920 is configured to encode the signal of the first frequency band to obtain the code stream data of the first frequency band; determine the side information of the second frequency band as the code stream data of the second frequency band; and write the high-frequency harmonic level identifier, the code stream data of the first frequency band, and the code stream data of the second frequency band into the code stream.

[0123] In some embodiments, the code stream data of the first frequency band is obtained by encoding the signal of the first frequency band in the audio frame using the Constrained Energy Overlap Transform (CELT) coding mode; the side information of the second frequency band is obtained by encoding the signal of the second frequency band using the Bandwidth Extension (BWE) coding mode.

[0124] In some embodiments, the bitstream includes a pitch period.

[0125] It should be noted that the above-mentioned units (modules) are merely logical modules divided according to the specific functions they implement, and are not intended to limit specific implementation methods. For example, they can be implemented in software, hardware, or a combination of software and hardware. In actual implementation, the above-mentioned units (modules) can be implemented as independent physical entities, or can also be implemented by a single entity (for example, a processor (CPU or DSP, etc.), an integrated circuit, etc.). In addition, the above-mentioned units are shown with dotted lines in the drawings to indicate that these units may not actually exist, and the operations / functions they implement can be implemented by the processing circuit itself.

[0126] In addition, although not shown, the device may also include a memory that can store various information generated by the device and the various units contained in the device during operation, programs and data used for operation, data to be sent by the communication unit, etc. The memory can be volatile memory and / or non-volatile memory. For example, the memory can include but is not limited to random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), read-only memory (ROM), and flash memory. Of course, the memory can also be located outside the device. Optionally, although not shown, the device may also include a communication unit that can be used to communicate with other devices. In one example, the communication unit can be implemented in an appropriate manner known in the art, for example, including communication components such as an antenna array and / or a radio frequency link, various types of interfaces, communication units, etc. This will not be described in detail here. In addition, the device may also include other components not shown, such as a radio frequency link, a baseband processing unit, a network interface, a processor, a controller, etc. This will not be described in detail here.

[0127] Some embodiments of the present disclosure also provide an electronic device. Figure 10 shows a block diagram of some embodiments of the electronic device of the present disclosure. For example, in some embodiments, the electronic device 100 can be various types of devices, for example, including but not limited to mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. For example, the electronic device 100 may include a display panel for displaying data and / or execution results utilized in the scheme of the present disclosure. For example, the display panel can be of various shapes, such as a rectangular panel, an elliptical panel, or a polygonal panel. In addition, the display panel can be not only a flat panel, but also a curved panel or even a spherical panel.

[0128] As shown in Figure 10, the electronic device 100 of this embodiment includes a memory 101 and a processor 102 coupled to the memory 101. It should be noted that the components of the electronic device 100 shown in Figure 10 are merely exemplary and non-limiting. The electronic device 100 may also have other components according to actual application requirements. The processor 102 can control the other components in the electronic device 100 to perform the desired functions.

[0129] In some embodiments, the memory 101 is configured to store one or more computer-readable instructions. When the processor 102 is configured to execute the computer-readable instructions, the computer-readable instructions, when executed by the processor 102, implement the decoding method of any embodiment of the present disclosure or the encoding method of any embodiment of the present disclosure. The specific implementation of each step of the method and related explanations can be found in the above-mentioned embodiments, and any repetitive details are not repeated here.

[0130] For example, the processor 102 and the memory 101 may communicate with each other directly or indirectly. For example, the processor 102 and the memory 101 may communicate with each other via a network. The network may include a wireless network, a wired network, and / or any combination of wireless networks and wired networks. The processor 102 and the memory 101 may also communicate with each other via a system bus, which is not limited in this disclosure.

[0131] For example, the processor 102 can be embodied as various appropriate processors, processing devices, etc., such as a central processing unit (CPU), a graphics processing unit (GPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The central processing unit (CPU) can be an X86 or ARM architecture, etc. For example, the memory 101 can include any combination of various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The memory 101 can include, for example, a system memory, which stores, for example, an operating system, an application, a boot loader (Boot Loader), a database, and other programs. Various applications and various data can also be stored in the storage medium.

[0132] In addition, according to some embodiments of the present disclosure, when various operations / processes according to the present disclosure are implemented through software and / or firmware, the programs constituting the software can be installed from a storage medium or a network to a computer system having a dedicated hardware structure, such as the electronic device 110 shown in Figure 11. When the various programs are installed, the electronic device can perform various functions, including functions such as those described above. Figure 11 is a block diagram illustrating an example structure of a computer system that can be used in an electronic device according to an embodiment of the present disclosure.

[0133] In Figure 11, a central processing unit (CPU) 1101 performs various processes according to a program stored in a read-only memory (ROM) 1102 or a program loaded from a storage part 1108 to a random access memory (RAM) 1103. In the RAM 1103, data required when the CPU 1101 performs various processes, etc., is also stored as needed. The central processing unit is merely exemplary and may also be other types of processors, such as the various processors described above. The ROM 1102, RAM 1103, and storage part 1108 may be various forms of computer-readable storage media, as described below. It should be noted that although ROM 1102, RAM 1103, and storage part 1108 are shown separately in Figure 11, one or more of them may be combined or located in the same or different memory or storage modules.

[0134] The CPU 1101, the ROM 1102, and the RAM 1103 are connected to one another via a bus 1104. An input / output interface 1105 is also connected to the bus 1104.

[0135] The following components are connected to the input / output interface 1105: an input portion 1106, such as a touch screen, touchpad, keyboard, mouse, image sensor, microphone, accelerometer, gyroscope, etc.; an output portion 1107, including a display, such as a cathode ray tube (CRT), liquid crystal display (LCD), speaker, vibrator, etc.; a storage portion 1108, including a hard disk, magnetic tape, etc.; and a communication portion 1109, including a network interface card, such as a LAN card, modem, etc. The communication portion 1109 allows communication processing to be performed via a network, such as the Internet. It will be readily understood that although FIG11 shows that the various devices or modules in the electronic device 110 communicate via the bus 1104, they may also communicate via a network or other means, where the network may include a wireless network, a wired network, and / or any combination of wireless and wired networks.

[0136] A drive 1110 is also connected to the input / output interface 1105 as needed. A removable medium 1111 such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc. is mounted on the drive 1110 as needed so that a computer program read therefrom is installed in the storage section 1108 as needed.

[0137] In the case of realizing the above-described series of processing by software, a program constituting the software can be installed from a network such as the Internet or a storage medium such as the removable medium 1111 .

[0138] According to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication part 1109, or installed from the storage part 1108, or installed from the ROM 1102. When the computer program is executed by the CPU 1101, the above-mentioned functions defined in the method of the embodiment of the present disclosure are performed.

[0139] It should be noted that in the context of the present disclosure, a computer-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium or any combination thereof. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to, an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present disclosure, a computer-readable storage medium may be any tangible medium that contains or stores a program that may be used by or in conjunction with an instruction execution system, apparatus, or device. In the present disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries a computer-readable program code. Such propagated data signals may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium, which may send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination thereof.

[0140] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.

[0141] In some embodiments, a computer program is further provided, comprising: instructions, which, when executed by a processor, cause the processor to perform any of the methods of the above embodiments. For example, the instructions may be embodied as computer program codes.

[0142] The present disclosure also provides an audio processing system, which is described below in conjunction with FIG12 .

[0143] FIG12 is a structural diagram of some embodiments of the audio processing system of the present disclosure. As shown in FIG12 , the system 12 of this embodiment includes: a decoding device 80 and an encoding device 90 of any of the aforementioned embodiments.

[0144] In embodiments of the present disclosure, computer program code for performing the operations of the present disclosure can be written in one or more programming languages ​​or combinations thereof, including but not limited to object-oriented programming languages, such as Java, Smalltalk, C++, and conventional procedural programming languages, such as "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In situations involving a remote computer, the remote computer can be connected to the user's computer via any type of network (including a local area network (LAN) or a wide area network (WAN)), or can be connected to an external computer (e.g., using an Internet service provider to connect via the Internet).

[0145] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the module, program segment, or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0146] The modules, components, or units described in the embodiments of the present disclosure may be implemented in software or hardware. The names of the modules, components, or units do not necessarily limit the modules, components, or units themselves.

[0147] The functions described above herein may be performed at least in part by one or more hardware logic components. For example, and without limitation, exemplary hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chip (SOCs), complex programmable logic devices (CPLDs), and the like.

[0148] According to some embodiments of the present disclosure, a decoding method is provided, comprising: parsing a high-frequency harmonic level identifier and a signal of a first frequency band of an audio frame from a bit stream, wherein the high-frequency harmonic level identifier is used to indicate the harmonic level of a second frequency band of the audio frame, and the frequency of the second frequency band is higher than that of the first frequency band; determining the signal of the second frequency band based on the fundamental frequency period of the audio frame, the high-frequency harmonic level identifier and the signal of the first frequency band; and obtaining a decoded signal of the audio frame based on the signal of the first frequency band and the signal of the second frequency band.

[0149] In some embodiments, determining the signal of the second frequency band based on the fundamental period of the audio frame, the high-frequency harmonic level identifier and the signal of the first frequency band includes: when the harmonic level of the second frequency band indicated by the high-frequency harmonic level identifier is higher than a threshold, determining the signal of the second frequency band based on the fundamental period and the signal of the first frequency band.

[0150] In some embodiments, determining the signal of the second frequency band based on the fundamental pitch period and the signal of the first frequency band includes: determining the subband copy starting frequency of each subband in one or more subbands in the second frequency band based on the fundamental pitch period; copying the signal corresponding to the subband copy starting frequency of each subband in the first frequency band to each subband; and determining the signal of the second frequency band based on the side information of the signal of each subband and the signal of the second frequency band.

[0151] In some embodiments, determining the subband copy starting frequency of each subband in one or more subbands in the second frequency band based on the fundamental pitch period includes: determining, for each subband, the frequency of the first harmonic component after the start frequency of the subband based on the fundamental pitch period; and determining the subband copy starting frequency of the subband based on the fundamental pitch period, the start frequency of the subband, and the frequency of the first harmonic component.

[0152] In some embodiments, determining the sub-band copy starting frequency of a sub-band according to the pitch period, the starting frequency of the sub-band, and the frequency of the first harmonic component includes: determining the reference frequency corresponding to the sub-band according to the following formula: Among them, freq pitch represents the fundamental frequency determined by the fundamental pitch period, Indicates the starting frequency of the i-th subband, i is an integer greater than 0, k i *freq pitch Indicates the frequency of the first harmonic component after the starting frequency of the i-th sub-band, k i is an integer greater than 0, a is a preset value, and offset is an offset value; when the bandwidth between the reference frequency corresponding to the subband and the starting frequency of the first subband in the second frequency band is not less than the width of the subband, the reference frequency corresponding to the subband is determined as the subband copy starting frequency of the subband.

[0153] In some embodiments, the decoding method further comprises: determining a sub-band copy starting frequency of the sub-band according to a coding scheme of the second frequency band when the frequency band width is smaller than the width of the sub-band.

[0154] In some embodiments, the decoding method also includes: when the harmonic level of the second frequency band indicated by the high-frequency harmonic level identifier is not higher than a threshold, determining the sub-band copy starting frequency of each sub-band in one or more sub-bands in the second frequency band according to the encoding type of the second frequency band; copying the signal corresponding to the sub-band copy starting frequency of each sub-band in the first frequency band to each sub-band; and determining the signal of the second frequency band based on the side information of the signal of each sub-band and the signal of the second frequency band.

[0155] In some embodiments, the bitstream includes a pitch period.

[0156] In some embodiments, the pitch period is obtained from a signal in a first frequency band.

[0157] In some embodiments, the signal of the first frequency band is obtained by decoding the code stream data of the first frequency band in the code stream using a Constrained Energy Overlap Transform (CELT) decoding mode.

[0158] In some embodiments, the side information of the second frequency band is obtained by decoding the code stream data of the second frequency band in the code stream using a bandwidth extension (BWE) decoding mode.

[0159] According to other embodiments of the present disclosure, a coding method is provided, comprising: obtaining a high-frequency harmonic level identifier and a signal of a first frequency band of an audio frame based on an input audio frame, wherein the high-frequency harmonic level identifier is used to indicate the harmonic level of a second frequency band of the audio frame, and the high-frequency harmonic level identifier is used together with the fundamental frequency period of the audio frame and the signal of the first frequency band to determine the signal of the second frequency band, the frequency of the second frequency band being higher than that of the first frequency band; and generating a code stream based on the high-frequency harmonic level identifier and the signal of the first frequency band.

[0160] In some embodiments, generating a bitstream based on a high-frequency harmonic level identifier and a signal of a first frequency band includes: encoding the signal of the first frequency band to obtain bitstream data of the first frequency band; determining side information of a second frequency band as bitstream data of the second frequency band; and writing the high-frequency harmonic level identifier, the bitstream data of the first frequency band, and the bitstream data of the second band into the bitstream.

[0161] In some embodiments, the code stream data of the first frequency band is obtained by encoding the signal of the first frequency band in the audio frame using the Constrained Energy Overlap Transform (CELT) coding mode; the side information of the second frequency band is obtained by encoding the signal of the second frequency band using the Bandwidth Extension (BWE) coding mode.

[0162] In some embodiments, the bitstream includes a pitch period.

[0163] According to some further embodiments of the present disclosure, a decoding device is provided, including: a first decoding module, configured to parse out a high-frequency harmonic level identifier and a signal of a first frequency band of an audio frame from a bit stream, wherein the high-frequency harmonic level identifier is used to indicate the harmonic level of a second frequency band of the audio frame, and the frequency of the second frequency band is higher than that of the first frequency band; a second decoding module, configured to determine the signal of the second frequency band based on the fundamental period of the audio frame, the high-frequency harmonic level identifier and the signal of the first frequency band; and an audio recovery module, configured to obtain a decoded signal of the audio frame based on the signal of the first frequency band and the signal of the second frequency band.

[0164] According to some further embodiments of the present disclosure, an encoding device is provided, including: an acquisition module, configured to acquire a high-frequency harmonic level identifier and a signal of a first frequency band of an audio frame based on an input audio frame, wherein the high-frequency harmonic level identifier is used to indicate the harmonic level of a second frequency band of the audio frame, and the high-frequency harmonic level identifier is used together with the fundamental frequency period of the audio frame and the signal of the first frequency band to determine the signal of the second frequency band, and the frequency of the second frequency band is higher than that of the first frequency band; and a generation module, configured to generate a code stream based on the high-frequency harmonic level identifier and the signal of the first frequency band.

[0165] According to some further embodiments of the present disclosure, an electronic device is provided, comprising: a memory; and a processor coupled to the memory, the processor being configured to execute a decoding method as in any embodiment of the present disclosure or an encoding method as in any embodiment of the present disclosure based on instructions stored in the memory.

[0166] According to some further embodiments of the present disclosure, a computer-readable storage medium is provided, on which a computer program is stored. When the program is executed by a processor, the decoding method of any embodiment of the present disclosure or the encoding method of any embodiment of the present disclosure is implemented.

[0167] According to some further embodiments of the present disclosure, a computer program is provided, comprising: instructions, which, when executed by a processor, cause the processor to execute the decoding method of any embodiment of the present disclosure or the encoding method of any embodiment of the present disclosure.

[0168] According to some embodiments of the present disclosure, a computer program product is provided, comprising instructions, which, when executed by a processor, implement the decoding method of any embodiment of the present disclosure or the encoding method of any embodiment of the present disclosure.

[0169] The above descriptions are merely some embodiments of the present disclosure and an illustration of the technical principles employed. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by a specific combination of the above-mentioned technical features, but also encompasses other technical solutions formed by any combination of the above-mentioned technical features or their equivalents without departing from the above-mentioned disclosed concepts. For example, a technical solution formed by replacing the above-mentioned features with (but not limited to) technical features with similar functions disclosed in the present disclosure.

[0170] In the description provided herein, numerous specific details are set forth. However, it is understood that embodiments of the present invention may be practiced without these specific details. In other cases, well-known methods, structures, and techniques are not presented in detail in order not to obscure the understanding of the description.

[0171] In addition, although each operation is described in a specific order, this should not be understood as requiring these operations to be performed in the specific order shown or in a sequential order. Under certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although some specific implementation details have been included in the above discussion, these should not be interpreted as limiting the scope of the present disclosure. Some features described in the context of a separate embodiment can also be implemented in a single embodiment in combination. On the contrary, the various features described in the context of a single embodiment can also be implemented in multiple embodiments individually or in any suitable sub-combination mode.

[0172] Although some specific embodiments of the present disclosure have been described in detail by way of examples, those skilled in the art will appreciate that the above examples are for illustrative purposes only and are not intended to limit the scope of the present disclosure. Those skilled in the art will appreciate that modifications may be made to the above embodiments without departing from the scope and spirit of the present disclosure. The scope of the present disclosure is defined by the appended claims.

Claims

1. A decoding method, comprising: Parsing a high-frequency harmonic level identifier and a signal of a first frequency band of an audio frame from a bitstream, wherein the high-frequency harmonic level identifier is used to indicate a harmonic level of a second frequency band of the audio frame, the second frequency band having a higher frequency than the first frequency band; determining a signal in the second frequency band according to the pitch period of the audio frame, the high-frequency harmonic level identifier, and the signal in the first frequency band; A decoded signal of the audio frame is obtained according to the signal of the first frequency band and the signal of the second frequency band.

2. The decoding method according to claim 1, wherein: The determining the signal of the second frequency band according to the pitch period of the audio frame, the high-frequency harmonic level identifier, and the signal of the first frequency band includes: When the harmonic level of the second frequency band indicated by the high-frequency harmonic level identifier is higher than a threshold, a signal of the second frequency band is determined according to the pitch period and the signal of the first frequency band.

3. The decoding method according to claim 2, wherein: The determining the signal of the second frequency band according to the pitch period and the signal of the first frequency band includes: determining, according to the pitch period, a subband copy starting frequency for each of the one or more subbands in the second frequency band; Copying a signal in the first frequency band corresponding to a sub-band copy start frequency of each sub-band to each sub-band; The signal of the second frequency band is determined according to the side information of the signal of each sub-band and the signal of the second frequency band.

4. The decoding method according to claim 3, wherein: The determining, based on the pitch period, a subband copy starting frequency of each of the one or more subbands in the second frequency band includes: for each subband, Determining, according to the pitch period, the frequency of the first harmonic component following the start frequency of the sub-band; A sub-band copy starting frequency of the sub-band is determined according to the pitch period, the starting frequency of the sub-band and the frequency of the first harmonic component.

5. The decoding method according to claim 4, wherein: Determining the sub-band copy starting frequency of the sub-band according to the pitch period, the starting frequency of the sub-band, and the frequency of the first harmonic component includes: The reference frequency corresponding to the sub-band is determined according to the following formula: Among them, freq pitch represents the fundamental frequency determined according to the fundamental pitch period, Indicates the starting frequency of the i-th subband, i is an integer greater than 0, k i *freq pitch Indicates the frequency of the first harmonic component after the starting frequency of the i-th sub-band, k i is an integer greater than 0, a is the preset value, and offset is the offset value; When a frequency bandwidth between a reference frequency corresponding to the subband and a starting frequency of a first subband in the second frequency band is not less than a width of the subband, the reference frequency corresponding to the subband is determined as a subband copy starting frequency of the subband.

6. The decoding method according to claim 5, further comprising: In a case where the frequency bandwidth is smaller than the width of the sub-band, a sub-band copy starting frequency of the sub-band is determined according to a coding scheme of the second frequency band.

7. The decoding method according to any one of claims 2 to 6, further comprising: When the harmonic level of the second frequency band indicated by the high-frequency harmonic level identifier is not higher than a threshold, determining, according to a coding scheme of the second frequency band, a sub-band copy starting frequency for each of one or more sub-bands in the second frequency band; Copying a signal in the first frequency band corresponding to a sub-band copy start frequency of each sub-band to each sub-band; The signal of the second frequency band is determined according to the side information of the signal of each sub-band and the signal of the second frequency band.

8. The decoding method according to any one of claims 1 to 7, wherein: The bit stream includes the pitch period.

9. The decoding method according to any one of claims 1 to 8, wherein: The pitch period is obtained through a signal in the first frequency band.

10. The decoding method according to any one of claims 1 to 9, wherein: The signal of the first frequency band is obtained by decoding the code stream data of the first frequency band in the code stream using a Constrained Energy Overlap Transform (CELT) decoding mode.

11. The decoding method according to any one of claims 3 to 7, wherein: The side information of the second frequency band is obtained by decoding the code stream data of the second frequency band in the code stream using a bandwidth extension (BWE) decoding mode.

12. A coding method comprising: Obtaining, based on an input audio frame, a high-frequency harmonic level identifier and a signal of a first frequency band of the audio frame, wherein the high-frequency harmonic level identifier is used to indicate a harmonic level of a second frequency band of the audio frame, and the high-frequency harmonic level identifier, together with a pitch period of the audio frame and the signal of the first frequency band, is used to determine a signal of the second frequency band, where a frequency of the second frequency band is higher than that of the first frequency band; A code stream is generated according to the high-frequency harmonic level identifier and the signal of the first frequency band.

13. The encoding method according to claim 12, wherein: The generating a code stream according to the high-frequency harmonic level identifier and the signal of the first frequency band includes: Encoding the signal of the first frequency band to obtain code stream data of the first frequency band; determining side information of the second frequency band as code stream data of the second frequency band; The high-frequency harmonic level identifier, the code stream data of the first frequency band, and the code stream data of the second frequency band are written into the code stream.

14. The encoding method according to claim 13, wherein: The bit stream data of the first frequency band is obtained by encoding the signal of the first frequency band in the audio frame using a constrained energy overlapped transform (CELT) coding mode; The side information of the second frequency band is obtained by encoding the signal of the second frequency band in a bandwidth extension (BWE) coding mode.

15. The encoding method according to any one of claims 12 to 14, wherein: The bit stream includes the pitch period.

16. A decoding device comprising: a first decoding module configured to parse a bitstream to obtain a high-frequency harmonic level identifier and a signal of a first frequency band of an audio frame, wherein the high-frequency harmonic level identifier is used to indicate a harmonic level of a second frequency band of the audio frame, the frequency of the second frequency band being higher than that of the first frequency band; a second decoding module configured to determine a signal of the second frequency band according to the pitch period of the audio frame, the high frequency harmonic level identifier, and the signal of the first frequency band; The audio recovery module is configured to obtain a decoded signal of the audio frame according to the signal of the first frequency band and the signal of the second frequency band.

17. An encoding device comprising: an acquisition module configured to acquire, based on an input audio frame, a high-frequency harmonic level identifier of the audio frame and a signal of a first frequency band, wherein the high-frequency harmonic level identifier is used to indicate a harmonic level of a second frequency band of the audio frame, and the high-frequency harmonic level identifier, together with a pitch period of the audio frame and the signal of the first frequency band, is used to determine a signal of the second frequency band, where a frequency of the second frequency band is higher than that of the first frequency band; The generating module is configured to generate a code stream according to the high-frequency harmonic level identifier and the signal of the first frequency band.

18. An electronic device comprising: processor; as well as A memory coupled to the processor, for storing instructions, wherein when the instructions are executed by the processor, the processor executes the decoding method according to any one of claims 1 to 11 and / or the encoding method according to any one of claims 12 to 15.

19. A computer-readable storage medium having a computer program stored thereon, wherein: When the program is executed by a processor, the decoding method according to any one of claims 1 to 11 and / or the encoding method according to any one of claims 12 to 15 are implemented.

20. A computer program product comprising: An instruction, which, when executed by a processor, implements the decoding method described in any one of claims 1 to 11 and / or the encoding method described in any one of claims 12 to 15.

21. An audio processing system comprising: The decoding device according to claim 16 and the encoding device according to claim 17.

22. A computer program comprising: An instruction, which, when executed by a processor, implements the decoding method described in any one of claims 1 to 11 and / or the encoding method described in any one of claims 12 to 15.

Citation Information

Patent Citations

  • Audio coding and decoding method and audio coding and decoding equipment

    CN113192521A

  • Audio coding method and device, electronic equipment and storage medium

    CN113744744A

  • Coding and decoding method of high-frequency audio signal and related device

    CN114550732A

  • Signal processing method and device, computer equipment, storage medium and program product

    CN117334204A

  • Reconstruction of a high-frequency range in low-bitrate audio coding using predictive pattern analysis

    US20140142959A1