Method and device for detecting attack in a sound signal to be encoded and decoded and encoding and decoding the detected attack
By using a non-predictive glottal shape codebook in the CELP codec, the low encoding and codec efficiency and poor quality in the case of voiced start are solved, and more efficient and robust voice signal processing is achieved.
Patent Information
- Application Number
- CN202080033815.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-05-07
- Filing Date
- 2020-05-01
- Publication Date
- 2025-08-26
- Estimated Expiration
- 2040-05-01
AI Technical Summary
When the existing CELP-based voice codecs process initiating frames, the encoding and decoding efficiency is low and the quality is poor, especially in the case of voiced start and transition frames, which leads to the codec status being out of synchronization, affecting the voice quality and encoding and decoding efficiency.
The detected attack frame is coded by using a non-predictive glottal shape codebook. By forcibly using the TC encoding and decoding mode in the attack frame, the glottal shape codebook is used instead of the adaptive codebook to improve the encoding and decoding efficiency and quality.
It improves the encoding and decoding efficiency and quality of the attacking frame, enhances the robustness of the codec in the noise channel, and improves the synthesis quality of the voice signal.
Smart Images

Figure CN113826161B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to techniques for encoding and decoding sound signals (such as speech or audio signals) in order to transmit and synthesize the sound signals.
[0002] More particularly, but not exclusively, the present disclosure relates to methods and apparatus for detecting an attack in a sound signal to be encoded or decoded (eg, a speech or audio signal), and for encoding or decoding the detected attack.
[0003] In this disclosure and the appended claims:
[0004] The term "onset" refers to the change in energy of a signal from low to high, such as the onset of a voiced sound (the transition from an unvoiced speech segment to a voiced speech segment), other sound onsets, transitions, plosives, etc., which is generally characterized by a sudden increase in energy within a sound signal segment.
[0005] The term "onset" refers to the beginning of a significant sound event, such as a speech, note, or other sound.
[0006] The term "plosive" in phonetics refers to consonants in which the vocal tract is blocked, stopping all airflow; and
[0007] The term "detected attack codec" refers to the codec of a sound signal segment, the length of which is typically within a few milliseconds after the start of the attack. Background Art
[0008] A speech encoder converts a speech signal into a digital bit stream, which is transmitted over a communication channel or stored on a storage medium. The speech signal is digitized—typically sampled and quantized at 16 bits per sample. The speech encoder's job is to represent these digital samples using a small number of bits while maintaining good subjective speech quality. A speech decoder or synthesizer operates on the transmitted or stored digital bit stream and converts it back into a speech signal.
[0009] The CELP (Codec Excited Linear Prediction) codec is one of the best technologies for achieving a good compromise between subjective quality and bit rate. This codec forms the basis of several speech codec standards for both wireless and wireline applications. In the CELP codec, the sampled speech signal is processed in consecutive blocks of M samples (often called frames), where M is a predetermined number of speech samples, typically corresponding to 10-30 ms. An LP (Linear Prediction) filter is calculated and transmitted for each frame. The calculation of the LP filter typically requires a lookahead, such as a 5-15 ms speech segment from the next frame. Each M-sample frame is divided into smaller blocks, called subframes. Typically, there are 2-5 subframes, resulting in subframes of 4-10 ms. In each subframe, the excitation is typically derived from two components: a past excitation contribution and an innovative, fixed-codebook excitation contribution. The past excitation contribution is often referred to as the pitch or adaptive codebook excitation contribution. Parameters representing the excitation are encoded and transmitted to the decoder, where the excitation is reconstructed and provided as input to the LP synthesis filter.
[0010] CELP based speech codecs rely heavily on prediction to achieve their high performance. This prediction can be of different types but usually involves the use of an adaptive codebook that stores adaptive codebook excitation contributions selected from previous frames. The CELP codec exploits the quasi-periodicity of voiced speech by searching for the past adaptive codebook excitation contributions for the segment that is most similar to the currently encoded segment. The same past adaptive codebook excitation contributions are also stored in the decoder. The encoder then only needs to send the pitch delay and pitch gain and the decoder can reconstruct the same adaptive codebook excitation contributions as used in the encoder. The evolution (difference) between the previous speech segment and the currently encoded speech segment is further modeled using fixed codebook excitation contributions selected from a fixed codebook.
[0011] The inherent prediction-related problems of CELP-based speech codecs arise when the encoder and decoder states become out of sync in the presence of transmission errors (erased frames or packets). Due to prediction, the effects of an erased frame are not limited to the erased frame but continue to propagate after the frame is erased, often for several subsequent frames. Naturally, the perceptual impact can be quite disturbing. Onsets, such as transitions from unvoiced to voiced speech segments (e.g., between a consonant or silent speech and a vowel) or between two distinct voiced segments (e.g., between two vowels), are among the most problematic cases for frame erasure masking. When the transition from an unvoiced to a voiced speech segment (voiced onset) is lost, the frame preceding the voiced onset frame is unvoiced or silent, and therefore no meaningful excitation contribution is found within the adaptive codebook buffer. In the encoder, the past excitation contribution is built up in the adaptive codebook during the voiced onset frame, and the following voiced frame is encoded and decoded using this past adaptive codebook excitation contribution. Most frame error concealment techniques use information from the last correctly received frame to conceal the lost frame. When a voiced onset frame is lost, the decoder's adaptive codebook buffer will therefore be updated with the noise-like adaptive codebook excitation contribution of the previous frame (unvoiced or inactive frame). Therefore, after the loss of the voiced onset, the periodic excitation part (adaptive codebook excitation contribution) is completely missing from the decoder's adaptive codebook, and the decoder may need several frames to recover from this loss. A similar situation occurs in the case of a transition from lost voiced to voiced speech. In this case, the excitation contribution stored in the adaptive codebook before the transition frame usually has very different characteristics from the excitation contribution stored in the adaptive codebook after the transition. Also, since the decoder usually uses past frame information to conceal the lost frame, the state of the encoder and the state of the decoder will be very different, and the synthesized signal will suffer from significant distortion. A solution to this problem is introduced in [2], where the inter-frame dependent adaptive codebook is replaced by a non-predictive glottal-shape codebook in the frames after the transition frame.
[0012] Another issue when encoding and decoding transition frames in CELP-based codecs is the codec efficiency. Codec efficiency decreases when the codec handles transitions where the excitations of the previous and current segments are very different. These situations usually occur in frames where the coding starts, such as voiced onset (transition from an unvoiced speech segment to a voiced speech segment), other sound onsets, transitions between two different voiced segments (e.g. transitions between two vowels), plosives, etc. The following two issues are the main reasons for this efficiency drop (mainly refer to [1]). The first issue is that the long-term prediction is very inefficient, so the contribution of the adaptive codebook to the total excitation is very weak. The second issue is related to the gain quantizer, which is usually designed as a vector quantizer with a limited bit budget, which usually cannot respond adequately to sudden energy increases within the frame. The closer this sudden energy increase occurs to the end of the frame, the more critical the second issue is.
[0013] To overcome the above problems, there is a need for a method and apparatus for improving the coding efficiency of frames including attack frames such as start frames and transition frames, and more generally, for improving the coding quality of CELP-based codecs. Summary of the Invention
[0014] According to a first aspect, the present disclosure relates to a method for detecting an attack in a sound signal to be encoded or decoded, wherein the sound signal is processed in consecutive frames, each frame comprising a plurality of subframes. The method comprises a first-stage attack detection for detecting an attack in a last subframe of a current frame, and a second-stage attack detection for detecting an attack in one of the subframes of the current frame, including a subframe preceding the last subframe.
[0015] The present disclosure also relates to a method for encoding and decoding an attack tone in a sound signal, comprising the above-defined attack detection method, wherein the encoding and decoding method comprises encoding and decoding a subframe containing a detected attack tone using an encoding and decoding mode with a non-predictive codebook.
[0016] According to another aspect, the present disclosure relates to a device for detecting an attack in a sound signal to be encoded or decoded, wherein the sound signal is processed in consecutive frames, each frame including a plurality of subframes. The device includes a first-stage attack detector for detecting an attack in a last subframe of a current frame, and a second-stage attack detector for detecting an attack in one of the subframes of the current frame, including a subframe preceding the last subframe.
[0017] The present disclosure also relates to a device for encoding and decoding an attack tone in a sound signal, comprising the above-defined attack detection device and a codec, and encoding and decoding a subframe containing a detected attack tone using a coding mode with a non-predictive codebook.
[0018] The above and other objects, advantages and features of the method and apparatus for detecting an attack in a sound signal to be encoded and decoded and for encoding and decoding a detected attack will become more apparent upon reading the following non-restrictive description of illustrative embodiments thereof, which is given by way of example only with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] In the attached figure:
[0020] Figure 1 is a schematic block diagram of a sound processing and communication system, depicting a possible context for implementation of a method and apparatus for detecting an attack in a sound signal to be encoded and decoded and encoding and decoding the detected attack;
[0021] Figure 2 is a schematic block diagram illustrating the structure of a CELP-based encoder and decoder, forming Figure 1 Part of the sound processing and communication system;
[0022] Figure 3 is a block diagram illustrating the operation of the modules of the EVS (Enhanced Voice Service) codec mode classification method and the EVS codec mode classifier;
[0023] Figure 4 is a block diagram illustrating both a method for detecting an attack in a sound signal to be encoded and decoded and the operation of a module of an attack detector for implementing the method;
[0024] FIG5 shows Figure 4 Graph of a first non-limiting illustrative example of the influence of an onset detector and a TC (transitional codec) codec mode on the quality of a decoded speech signal, wherein curve a) represents an input speech signal, curve b) represents a reference speech signal synthesis, and curve c) represents the effect of a decoded speech signal on the quality of a decoded speech signal. Figure 4 The onset detector and TC codec mode are used to improve speech signal synthesis when processing the start frame;
[0025] Figure 6 It shows Figure 4 A diagram of a second non-limiting illustrative example of the effect of an onset detector and a TC codec mode on the quality of a decoded speech signal, wherein curve a) represents an input speech signal, curve b) represents a reference speech signal synthesis, and curve c) represents the effect of a decoded speech signal on the quality of a decoded speech signal. Figure 4 The onset detector and TC codec mode are used to improve speech signal synthesis when processing the start frame; and
[0026] Figure 7 is a simplified block diagram of an example configuration of hardware components for implementing methods and apparatus for detecting an attack in a sound signal to be encoded and decoded and encoding and decoding the detected attack. DETAILED DESCRIPTION
[0027] Although non-limiting illustrative embodiments of methods and apparatus for detecting an attack in a sound signal to be encoded and decoded and for encoding and decoding the detected attack will be described in the following description in conjunction with speech signals and a CELP-based codec, it should be remembered that these methods and apparatus are not limited to application to speech signals and a CELP-based codec, and the principles and concepts thereof can be applied to any other type of sound signals and codecs.
[0028] The following description relates to detecting an onset in a sound signal, such as a speech or audio signal, and forcing a transition codec (TC) mode in a subframe in which the onset is detected. The onset detection can also be used to select a subframe in which a glottal shaped codebook is used instead of an adaptive codebook as part of the TC codec mode.
[0029] In the EVS codec described in reference [4], when the detection algorithm detects an onset in the last subframe of the current frame, the glottal shape codebook of the TC codec mode is used in the last subframe. In the present disclosure, the detection algorithm is supplemented with a second-stage logic that not only detects more frames containing an onset, but also forces the use of the TC codec mode and the corresponding glottal shape codebook in all subframes where an onset is detected when encoding and decoding these frames.
[0030] The above technique not only improves the efficiency of encoding and decoding detected attack sounds in the sound signal to be encoded and decoded, but also improves the efficiency of encoding and decoding certain musical segments (such as castanets). More generally, the encoding and decoding quality is improved.
[0031] Figure 1 is a schematic block diagram of a sound processing and communication system 100 , depicting a possible context for implementation of a method and apparatus for detecting an attack in a sound signal to be encoded and decoding and encoding and decoding the detected attack, as disclosed in the following description.
[0032] Figure 1 The sound processing and communication system 100 supports the transmission of sound signals across a communication channel 101. The communication channel 101 may comprise, for example, an electrical wire or fiber optic link. Alternatively, the communication channel 101 may comprise, at least in part, a radio frequency link. Radio frequency links typically support multiple simultaneous communications requiring shared bandwidth resources, such as those found in cellular telephones. Although not shown, in a single-device implementation of the system 100, the communication channel 101 may be replaced by a storage device that records and stores the encoded sound signals for later playback.
[0033] Still refer to Figure 1, for example, a microphone 102 generates an original analog sound signal 103. As indicated in the foregoing description, the sound signal 103 may include, in particular but not limited to, speech and / or audio.
[0034] The analog sound signal 103 is provided to an analog-to-digital (A / D) converter 104 for conversion into a raw digital sound signal 105. The raw digital sound signal 105 may also be recorded and provided from a storage device (not shown).
[0035] The sound encoder 106 encodes the digital sound signal 105, thereby producing a set of coding parameters that are multiplexed in the form of a bit stream 107, which is delivered to an optional error correction channel encoder 108. The optional error correction channel encoder 108, when present, adds redundancy to the binary representation of the coding parameters in the bit stream 107 before transmitting the resulting bit stream 111 over the communication channel 101.
[0036] On the receiver side, an optional error correction channel decoder 109 utilizes the redundant information in the received digital bit stream 111 to detect and correct errors that may have occurred during transmission over the communication channel 101, generating an error-corrected bit stream 112 having received coding parameters. An audio decoder 110 converts the received coding parameters in the bit stream 112 to create a synthesized digital audio signal 113. The digital audio signal 113 reconstructed in the audio decoder 110 is converted into a synthesized analog audio signal 114 in a digital-to-analog (D / A) converter 115.
[0037] The synthesized analog sound signal 114 is played back in the speaker unit 116 (the speaker unit 116 may obviously be replaced by headphones). Alternatively, the digital sound signal 113 from the sound decoder 110 may also be supplied to a storage device (not shown) and recorded in the storage device.
[0038] As a non-limiting example, the method and apparatus for detecting an attack in a sound signal to be encoded and decoded and for encoding and decoding the detected attack according to the present disclosure may be used in Figure 1 The voice encoder 106 and decoder 110 are implemented. It should be noted that Figure 1 The sound processing and communication system 100, as well as the method and apparatus for detecting an attack in a sound signal to be encoded and decoded and for encoding and decoding the detected attack, can be extended to cover the case of stereo, where the input of the encoder 106 and the output of the decoder 110 include the left and right channels of the stereo signal. Figure 1The sound processing and communication system 100, as well as the method and apparatus for detecting an attack in a sound signal to be encoded and decoded and for encoding and decoding the detected attack, can be further extended to cover multi-channel and / or scene-based audio and / or independent stream encoding and decoding scenarios (e.g., surround and Ambisonics).
[0039] Figure 2 is a schematic block diagram illustrating the structure of a CELP-based encoder and decoder, according to an illustrative embodiment, the encoder and decoder are Figure 1 Part of the sound processing and communication system 100. Figure 2 As shown, the audio codec includes two basic parts: audio encoder 106 and audio decoder 110. Figure 1 The encoder 106 is supplied with the original digital sound signal 105 and determines the encoding parameters 107 described below that represent the original analog sound signal 103. These parameters 107 are encoded into a digital bit stream 111. As already explained, the bit stream 111 is transmitted using a communication channel, such as Figure 1 The synthesized digital sound signal 113 is transmitted to the decoder 110 via the communication channel 101. The sound decoder 110 reconstructs the synthesized digital sound signal 113 so that it is as similar to the original digital sound signal 105 as possible.
[0040] Currently, the most common speech codec technology is based on linear prediction (LP), especially CELP. In LP-based codecs, the synthesized digital sound signal 230 ( Figure 2 ) is generated by filtering the excitation 214 through an LP synthesis filter 216 with a transfer function 1 / A(z). An example of a procedure for finding the filter parameters A(z) of the LP filter can be found in reference [4].
[0041] In CELP, the excitation 214 typically consists of two parts: the first stage, the adaptive codebook contribution 222, which selects the past excitation signal v(n) from the adaptive codebook 218 in response to the index t (pitch lag) and adds it to the adaptive codebook gain g p 226 amplifies the past excitation signal v(n) and generates; in the second stage, the fixed codebook contribution 224 is generated by selecting the innovative code vector c from the fixed codebook 220 k (n) with response index k and fixed codebook gain g c 228 zoom innovation code vector c k In general, the adaptive codebook contribution 222 models the periodic part of the excitation, and the fixed codebook excitation contribution 224 is added to model the evolution of the sound signal.
[0042] The sound signal is processed in frames of typically 20 ms, and the filter parameters A(z) of the LP filter are transmitted from the encoder 106 to the decoder 110 once per frame. In CELP, the frame is further divided into several subframes to encode the excitation. The length of the subframe is typically 5 ms.
[0043] CELP uses a principle known as "Analysis-by-Synthesis," in which possible decoder outputs are tried (synthesized) during the encoding and decoding process of the encoder 106 and then compared to the original digital sound signal 105. Thus, the encoder 106 includes elements similar to those of the decoder 110. These elements include an adaptive codebook excitation contribution 250 (corresponding to the adaptive codebook contribution 222 of the decoder 110), which is selected from an adaptive codebook 242 (corresponding to the adaptive codebook 218 of the decoder 110) in response to an index t (pitch lag). The adaptive codebook 242 provides the past excitation signal v(n) convolved with the impulse response of a weighted synthesis filter H(z) 238 (a cascade of the LP synthesis filter 1 / A(z) and the perceptual weighting filter W(z)). The output of the weighted synthesis filter H(z) 238, y1(n), is multiplied by the adaptive codebook gain g. p 240 (corresponding to the adaptive codebook gain 226 of the decoder 110). These elements also include a fixed codebook excitation contribution 252 (corresponding to the fixed codebook contribution 224 of the decoder 110), which is selected from the fixed codebook 244 (corresponding to the fixed codebook 220 of the decoder 110) in response to the index k, and the fixed codebook 244 provides the innovation code vector c that is convolved with the impulse response of the weighted synthesis filter H(z) 246. k (n), the weighted synthesis filter H(z) 246 output y2(n) is fixed by the codebook gain g c The codebook gain is amplified by 248 (corresponding to the fixed codebook gain 228 of the decoder 110).
[0044] The encoder 106 includes a calculator 234 for calculating the zero-input response of the perceptual weighting filter W(z) 233 and the cascade (H(z)) of the LP synthesis filter 1 / A(z) and the perceptual weighting filter W(z). Subtractors 236, 254, and 256 respectively subtract the zero-input response of the calculator 234, the adaptive codebook contribution 250, and the fixed codebook contribution 252 from the original digital sound signal 105 filtered by the perceptual weighting filter 233 to provide the original digital sound signal 105 and the synthesized digital sound signal 113 ( Figure 1 ) is an error signal of the mean square error 232 between .
[0045] The adaptive codebook 242 and the fixed codebook 244 are searched to minimize the mean square error 232 between the original digital sound signal 105 and the synthesized digital sound signal 113 in the perceptual weighted domain, where the discrete time index n=0, 1, ..., N-1, and N is the length of the subframe. Minimization of the mean square error 232 provides the best candidate past excitation signal v(n) (identified by index t) and the innovation code vector c for encoding and decoding the digital sound signal 105. k (n) (identified by index k). The perceptual weighting filter W(z) exploits the frequency masking effect and is usually derived from the LP filter A(z). Examples of perceptual weighting filters W(z) for WB (wideband, typically 50-7000 Hz) signals can be found in reference [4].
[0046] Since the memory of LP synthesis filter 1 / A(z) and weighted filter W(z) is the same as the searched innovative code vector c k (n), this memory (the zero-input response of the cascade of the LP synthesis filter 1 / A(z) and the perceptual weighting filter W(z) (H(z))) can be subtracted (subtractor 236) from the original digital sound signal 105 before the fixed codebook search. Figure 2 The candidate innovation code vector c is completed by convolving the impulse response of the cascade of filters 1 / A(z) and W(z) represented by H(z) k (n) filtering.
[0047] The digital bit stream 111 transmitted from the encoder 106 to the decoder 110 typically contains the following parameters 107: the quantization parameter of the LP filter A(z), the index t of the adaptive codebook 242 and the index k of the fixed codebook 244, and the gains g of the adaptive codebook 242 and the fixed codebook 244. p 240 and g c 248. In decoder 110:
[0048] - the received quantization parameters of the LP filter A(z) are used to establish the LP synthesis filter 216;
[0049] - the received index t is applied to the adaptive codebook 218;
[0050] - the received index k is applied to the fixed codebook 220;
[0051] - Gain received g p is used as the gain 226 of the adaptive codebook; and
[0052] - Gain received g c is used as the fixed codebook gain 228.
[0053] Further explanation of the structure and operation of CELP based encoders and decoders can be found, for example, in reference [4].
[0054] Additionally, although the following description refers to the EVS standard (reference [4]), it should be remembered that the concepts, principles, structures, and operations described therein are applicable to other sound / speech processing and communication standards.
[0055] Voiced Onset Codec
[0056] To achieve better codec performance, the LP-based core of the EVS codec described in reference [4] uses a signal classification algorithm and six (6) different codec modes tailored for each class of signal, namely, Inactive Codec (IC) mode, Unvoiced Codec (UC) mode, Transition Codec (TC) mode, Voiced Codec (VC) mode, General Codec (GC) mode, and Audio Codec (AC) mode (not shown).
[0057] Figure 3 is a simplified high-level block diagram illustrating both the operation of the EVS codec mode classification method 300 and the modules of the EVS codec mode classifier 320 .
[0058] Reference Figure 3 The coding mode classification method 300 includes an active frame detection operation 301 , an unvoiced frame detection operation 302 , a post-initial frame detection operation 303 , and a stable voiced frame detection operation 304 .
[0059] To perform the active frame detection operation 301, the active frame detector 311 determines whether the current frame is active or inactive. To this end, sound activity detection (SAD) or voice activity detection (VAD) can be used. If an inactive frame is detected, the IC codec mode 321 is selected and the program is terminated.
[0060] If detector 311 detects an active frame during active frame detection operation 301, unvoiced frame detection operation 302 is performed using unvoiced frame detector 312. Specifically, if an unvoiced frame is detected, unvoiced frame detector 312 selects UC codec mode 322 to encode and decode the detected unvoiced frame. The UC codec mode is designed to encode and decode unvoiced frames. In the UC codec mode, an adaptive codebook is not used, and the excitation includes two vectors selected from a linear Gaussian codebook. Alternatively, the codec mode in UC can include a fixed algebraic codebook and a Gaussian codebook.
[0061] If the current frame is not classified as unvoiced by the detector 312, then the post-start frame detection operation 303 and the corresponding post-start frame detector 313, as well as the stable voiced frame detection operation 304 and the corresponding stable voiced frame detector 314 are used.
[0062] In post-onset frame detection operation 303, detector 313 detects voiced frames after the onset of voiced speech and selects TC codec mode 323 to encode and decode these frames. TC codec mode 323 is designed to improve codec performance in the presence of frame erasures by limiting the use of past information (adaptive codebook). To minimize the impact of TC codec mode 323 on clean channel performance (without frame erasures), mode 323 is only used for the most critical frames from a frame erasure perspective. These most critical frames are voiced frames after the onset of voiced speech.
[0063] If the current frame is not a voiced frame following a voiced onset, a stable voiced frame detection operation 304 is performed. During this operation, a stable voiced frame detector 314 is designed to detect quasi-periodic stable voiced frames. If the current frame is detected as a quasi-periodic stable voiced frame, the detector 314 selects the VC codec mode 324 to encode the stable voiced frame. The selection of the VC codec mode by the detector 314 is conditioned on smooth pitch evolution. This uses the Algebraic Code Excited Linear Prediction (ACELP) technique, but since the pitch evolution is smooth throughout the frame, more bits are allocated to the fixed (algebraic) codebook compared to the GC codec mode.
[0064] If during operations 301 - 304 the current frame is not classified into one of the above frame categories, the frame may contain a non-stationary speech segment and the detector 314 selects a GC codec mode 325 , such as the general ACELP codec mode, for encoding such a frame.
[0065] Finally, the EVS standard speech / music classification algorithm (not shown) is run to decide whether the current frame should be encoded or decoded using the AC mode. The AC mode is designed to efficiently encode and decode general audio signals, especially but not limited to music.
[0066] To improve the performance of the codec in noisy channels, the reference described in the previous paragraphs was applied. Figure 3A refinement of the codec mode classification method is called frame classification for Frame Error Concealment (FEC) (Ref. [4]). The basic idea behind using different frame classification methods for FEC is that the ideal FEC strategy should be different for quasi-static speech segments and speech segments with rapidly changing characteristics. In the EVS standard (Ref. [4]), the FEC frame classification used by the codec defines the following five (5) different classes. The unvoiced class includes all unvoiced speech frames and all frames without active speech. A voiced offset frame can also be classified as unvoiced if the end of the frame tends to be unvoiced. The unvoiced transition class includes unvoiced frames with possible voiced onset at the end of the frame. The voiced transition class includes voiced frames with relatively weak voiced characteristics. The voiced class includes voiced frames with stable characteristics. The onset class includes all voiced frames with stable characteristics following a frame classified as unvoiced or unvoiced transition.
[0067] about Figure 3 A further explanation of the EVS codec mode classification method 300 and the EVS codec mode classifier 320 can be found, for example, in reference [4].
[0068] Initially, the TC codec mode was introduced to transition frames to help stop error propagation in case of loss of transition frames (Ref. [4]). Furthermore, the TC codec mode can be used in transition frames to improve codec efficiency. In particular, before the onset of voiced speech, the adaptive codebook usually contains a noise-like signal that is not very useful or efficient for encoding and decoding the start of voiced segments. The goal is to supplement the adaptive codebook with a better, non-predictive codebook filled with a quantized version of the simplified glottal pulse shape to encode the onset of voiced speech. The glottal shape codebook is only used for the subframe containing the first glottal pulse in the frame, more precisely, for the subframe in which the LP residual signal ( Figure 2 s in w (n)) The subframe with its maximum energy in the first pitch period of the frame. Figure 3 A further explanation of the TC codec mode can be found, for example, in reference [4].
[0069] The present disclosure proposes to further extend the concept of EVS for encoding and decoding voiced onsets by using a glottal shaped codebook in a TC codec mode. When the onset occurs at the end of a frame, it is recommended to use as much of the bit budget (the number of available bits) as possible to encode and decode the excitation at the end of the frame, because it is sufficient to encode and decode the previous part of the frame (including the subframe before the subframe of the onset) with a low number of bits. Unlike the TC codec mode of EVS described in reference [4], the glottal shaped codebook is usually used in the last subframe in the frame, regardless of the actual maximum energy of the LP residual signal in the first pitch period of the frame.
[0070] By forcing most of the bit budget to be used at the end of the coding frame, the waveform of the sound signal at the beginning of the frame may not be well modeled, especially at low bit rates, where the fixed codebook includes, for example, only one or two pulses per subframe. However, the sensitivity of the human ear is exploited here. The human ear is not sensitive to inaccurate coding and decoding of the sound signal before the attack, but is more sensitive to any imperfect coding and decoding of the sound signal segment after the attack (such as the voiced segment). By forcing a larger number of bits to construct the attack, the adaptive codebook in the subsequent sound signal frame is more efficient because it benefits from past excitations corresponding to the well-modeled attack segment. The subjective quality is thus improved.
[0071] This disclosure proposes a method for detecting onset and a corresponding onset detector. These detectors operate on frames intended for encoding and decoding in the GC codec mode to determine whether these frames should be encoded and decoded in the TC codec mode. Specifically, when an onset is detected, these frames are encoded and decoded in the TC codec mode. As a result, the relative number of frames encoded and decoded in the TC codec mode increases. Furthermore, since the TC codec mode does not use past excitation, this approach can improve the inherent robustness of the codec to frame erasures.
[0072] Attack detection method and onset detector
[0073] Figure 4 is a block diagram illustrating the operation of the modules of both the onset detection method 400 and the onset detector 450 .
[0074] The onset detection method 400 and the onset detector 450 appropriately select frames to be coded or decoded using the TC codec mode. Figure 4 An example of an onset detection method 400 and an onset detector 450 is described that can be used with a codec, in this illustrative example, a CELP codec with an internal sampling rate of 12.8 kbps and with a frame length of 20 ms comprising four (4) subframes. An example of such a codec is the lower bit rate (≤ 13.2 kbps) EVS codec (reference [4]). Application to other types of codecs having different internal bit rates, frame lengths, and numbers of subframes is also contemplated.
[0075] Onset detection begins with preprocessing, where the energy of several segments of the input sound signal in the current frame is calculated. This is followed by detection and a final decision in two stages. The first stage of detection is based on comparing the calculated energy of the current frame, while the second stage also takes into account the energy values of past frames.
[0076] Energy of segment
[0077] exist Figure 4 In the energy calculation operation 401, the energy calculator 451 calculates the perceptually weighted input sound signal s w (n) wherein n=0, ..., N-1, and wherein N is the length of the frame in samples. To calculate this energy, calculator 451 may use, for example, the following equation (1):
[0078]
[0079] Wherein K is the length of the samples of the analysis sound signal segment, i is the index of the segment, and N / K is the total number of segments. In the EVS standard operating at an internal sampling rate of 12.8k bps, the length of the frame is N=256 samples, and the length of the segment can be set to, for example, K=8, which results in a total number of N / K=32 analysis segments. Thus, segment i=0,...,7 corresponds to the first subframe, segment i=8,...,15 corresponds to the second subframe, segment i=16,...,23 corresponds to the third subframe, and finally, segment i=24,...,31 corresponds to the last (fourth) subframe of the current frame. In the non-limiting illustrative example of equation (1), these segments are continuous. In another possible embodiment, partially overlapping segments can be used.
[0080] Next, in the maximum energy segment search operation 402, the maximum energy segment finder 452 searches for the segment i with the maximum energy. To do this, the finder 452 may use, for example, the following equation (2):
[0081]
[0082] The segment with the maximum energy represents the location of a candidate attack, which is verified in the following two stages (referred to herein as the first stage and the second stage).
[0083] In the illustrative embodiment given as an example in this description, only active frames (VAD=1, where local VAD is taken into account in the current frame) that were previously classified as being processed using the GC codec mode are subjected to the following first and second stage onset detection. Further explanation of VAC (Voice Activity Detection) can be found, for example, in reference [4]. In decision operation 403, decision module 453 determines whether VAD=1 and the current frame has been classified as being processed using the GC codec mode. If so, the first stage onset detection is performed on the current frame. Otherwise, no onset is detected and the current frame is processed according to its previous classification, e.g. Figure 3 shown.
[0084] Both speech and music frames can be classified in GC codec mode, so onset detection is applicable not only to encoding and decoding speech signals, but also to encoding and decoding general sound signals.
[0085] First stage attack detection
[0086] Now refer to Figure 4 A first stage attack detection operation 404 and a corresponding first stage attack detector 454 are described.
[0087] The first stage onset detection operation 404 includes an average energy calculation operation 405. To perform operation 405, the first stage onset detector 454 includes a calculator 455 that calculates the average energy of the entire analysis segment before the last subframe in the current frame using, for example, the following equation (3):
[0088]
[0089] Where P is the number of segments before the last subframe. In a non-limiting example implementation, where N / K=32, the parameter P is equal to 24.
[0090] Similarly, in the average energy calculation operation 405, the calculator 455 calculates the average energy from the segment 1 using the following equation (4) as an example. att The average energy of the entire analysis segment starting from the end of the current frame.
[0091]
[0092] The first stage attack detection operation 404 also includes a comparison operation 406. To perform the comparison operation 406, the first stage attack detector 454 includes a comparator 456 for comparing the ratio of the average energy E1 from equation (3) and the average energy E2 from equation (4) with a threshold value that depends on the signal classification of the previous frame denoted as "last_class" performed by the frame classification for frame error concealment (FEC) discussed above (reference [4]). The comparator 456 determines the attack position I from the first stage attack detection. att1 , as a non-limiting example, using the following logic of equation (5):
[0093]
[0094] Where β1 and β2 are thresholds, which, according to a non-limiting example, can be set to β1=8 and β2=20, respectively. att1 = 0, no attack is detected. Using the logic of equation (5), all attacks that are not strong enough are eliminated.
[0095] In order to further reduce the number of falsely detected attacks, the first stage attack detection operation 404 further includes a segment energy comparison operation 407. To perform the segment energy comparison operation 407, the first stage attack detector 454 includes a segment energy comparator 457 for comparing the segment with the maximum energy E seg (I att ) and the energy E of other analysis segments of the current frame seg (i) Comparison. Therefore, if I determined by operation 406 and comparator 456 att1 >0, the comparator 457 performs the comparison of equation (6) for i=2, ..., P-3 as a non-limiting example:
[0096]
[0097] The threshold β3 is determined experimentally to minimize false detections of onsets without hindering the efficiency of detecting true onsets. In a non-limiting experimental implementation, the threshold β3 is set to 2. Similarly, when I att1 = 0, no attack is detected.
[0098] Second stage attack detection
[0099] Now refer to Figure 4 A second stage attack detection operation 410 and a corresponding second stage attack detector 460 are described.
[0100] The second-stage onset detection operation 410 includes a voiced class comparison operation 411. To perform the voiced class comparison operation 411, the second-stage onset detector 460 includes a voiced class decision module 461 to obtain information from the EVS FEC classification method discussed above to determine whether the current frame class is voiced. If the current frame class is voiced, the decision module 461 outputs a decision that no onset is detected.
[0101] If no attack is detected in the first stage attack detection operation 404 and the first stage attack detector 454 (specifically, the comparison operation 406 and the comparator 456 or the comparison operation 407 and the comparator 457), that is, I att1 =0, and the class of the current frame is other than voiced, then the second stage onset detection operation 410 and the second stage onset detector 460 are applied.
[0102] The second stage attack detection operation 410 includes an average energy calculation operation 412. To perform operation 412, the second stage attack detector 460 includes an average energy calculator 462 for calculating the average energy of the candidate attack I using, for example, equation (7). att Average energy of the previous N / K analysis segments (including segments from previous frames):
[0103]
[0104] Among them E seg,past (i) is the energy of each segment from the previous frame.
[0105] The second stage attack detection operation 410 includes a logic decision operation 413. To perform operation 413, the second stage attack detector 460 includes a logic decision module 463 to find the attack position I from the second stage attack detector by applying the following logic, such as equation (8), to the average energy from equation (7): att2 :
[0106]
[0107] Among them I att is found in equation (2), and β4 and β5 are thresholds, which in this non-limiting example implementation are set to β4=16 and β5=12, respectively. att2 = 0, no attack is detected.
[0108] The second stage attack detection operation 410 finally includes an energy comparison operation 414. To perform operation 414, the second stage attack detector 460 includes an energy comparator 464 to compare the I determined in operation 413 with the I determined in comparator 463. att2 When is greater than 0, the following ratio is compared to the following threshold, for example, as shown in equation (9), to further reduce the number of falsely detected onsets:
[0109]
[0110] Where β6 is a threshold value set to β6=20 in this non-limiting example implementation, E LT is the long-term energy calculated using equation (10) as a non-limiting example.
[0111]
[0112] In this non-limiting example implementation, the parameter α is set to 0.95. att2 = 0, no attack is detected.
[0113] Finally, in the energy comparison operation 414, if an attack is detected in the previous frame, the energy comparator 464 compares the attack position I att2 Set to 0. In this case, no attack is detected.
[0114] Final attack detection decision
[0115] Based on the attack I obtained during the first stage 404 and the second stage 410 detection operations, respectively att1 and I att2 The final decision is whether to determine the current frame as the onset frame to be encoded and decoded using the TC codec mode.
[0116] If the current frame is active (VAD=1) and was previously classified as codec in GC codec mode as determined in decision operation 403 and decision module 453, the following logic, such as equation (11), applies:
[0117]
[0118] Specifically, the attack detection method 400 includes a first stage attack decision operation 430. To perform operation 430, if the current frame is active (VAD=1) and was previously classified as a codec in the GC codec mode determined in decision operation 403 and decision module 453, the attack detector 450 also includes a first stage attack decision module 470 to determine I att1 ≥P. If I att1 ≥P, then I att1 is the position of the onset detected in the last subframe of the current frame I att,final , and the glottal shape codebook for determining the TC codec mode is used in the last subframe. Otherwise, no attack is detected.
[0119] Regarding the second stage attack detection, if the comparison of equation (9) is true, or if an attack was detected in the previous frame as determined in energy comparison operation 414 and energy comparator 464, then I att2 = 0 and no attack is detected. Otherwise, in the attack decision operation 440 of the attack detection method 400, the attack decision module 480 of the attack detector 450 determines the attack at position I in the current frame. att,final =I att2 The attack is detected at the position I att,final The glottal shape codebook is used to determine in which subframe to use the TC codec mode.
[0120] The final position of the detected attack I att,final The information of is used to determine in which subframe of the current frame the glottal shape codebook in the TC codec mode is used, and which TC mode configuration is used (see reference [3]). For example, in the case of a frame of N = 256 samples, the frame is divided into four (4) subframes and N / K = 32 analysis segments, if the final attack position I is detected in segments 1-7 att,final , then the glottal shape codebook is used in the first subframe; if the final attack position I is detected in segments 8-15att,final , then the glottal shape codebook is used in the second subframe, and if the final attack position I is detected in segment 16-23 att,final , then the glottal shape codebook is used in the third subframe, and if the final attack position I is detected in segment 24-31 att,final , then the glottal shape codebook is used in the last (fourth) subframe of the current frame. att,final =0 signals that no tone onset is found and the current frame is coded according to the original classification (usually using GC codec mode).
[0121] Illustrative implementation in immersive speech / audio codecs
[0122] The onset detection method 400 includes a glottal shape codebook assignment operation 445. To perform operation 445, the onset detector 450 includes a glottal shape codebook assignment module 485 to assign a glottal shape codebook within the TC codec mode to a specific subframe of a current frame including four subframes using the following logic of equation (12):
[0123]
[0124] Here, sbfr is a subframe index, sbfr=0, . . . 3, where index 0 represents the first subframe, index 1 represents the second subframe, index 2 represents the third subframe, and index 3 represents the fourth subframe.
[0125] The above description of the non-limiting embodiment assumes that the pre-processing module operates at an internal sampling rate of 12.8 kHz, with four (4) subframes, so a frame has a number of samples N = 256. If the core codec uses ACELP at an internal sampling rate of 12.8 kHz, the final attack position I att,final is assigned to the subframes defined in equation (12). However, the situation is different when the core codec operates at a different internal sampling rate, for example at higher bit rates (16.4 kbps or higher in the case of EVS) where the internal sampling rate is 16 kHz. Assuming a frame length of 20 ms, in this case the frame consists of 5 subframes and the length of such a frame is N 16 =320 samples. In this example of implementation, since the pre-processing classification and analysis may still be performed in the internal sampling nominal domain of 12.8 kHz, the glottal shape codebook allocation module 485 uses the logic of the following equation (13) in the glottal shape codebook allocation operation 445 to select the subframe to be encoded using the glottal shape codebook within the TC codec mode:
[0126]
[0127] where the operator represents the largest integer less than or equal to x. In the case of equation (13), sbfr = 0, ... 4 is different from equation (12), while the number of analysis segments is the same as equation (12), that is, N / K = 32. Therefore, if the final attack position I att,final is detected in segments 1-6, the glottal shape codebook is used in the first subframe; if the final onset position I att,final is detected in segments 7-12, then the glottal shape codebook is used in the second subframe; if the final attack position I att,final If the glottal shape codebook is used in the third subframe, the final attack position I att,final is detected in the 20-25 segment, the glottal shape codebook is used in the fourth subframe; finally, if the final onset position I att,final If segment 26-31 is detected, the glottal shape codebook is used in the last (fifth) subframe of the current frame.
[0128] FIG5 shows Figure 4 A first non-limiting illustrative example of the effect of an attack detector and TC codec mode on the quality of a decoded music signal is shown in FIG5 . Specifically, FIG5 shows a musical segment of castanets, wherein curve a) represents the input (unencoded) music signal, curve b) represents the decoded reference signal synthesis using only the first-stage attack detection, and curve c) represents the decoded improved synthesis using both the first-stage and second-stage attack detection and the TC codec mode for encoding. Comparing curves b) and c), it can be seen that the attack (a low-to-high amplitude onset, such as 500 in FIG5 ) in the synthesis of curve c) is significantly more accurately reconstructed, both in terms of preserving the energy and sharpness of the castanets signal at the beginning of the onset.
[0129] Figure 6 It shows Figure 4 FIGURE 4 is a second non-limiting illustrative example of the effect of an onset detector and a TC codec mode on the quality of a decoded speech signal, wherein curve a) represents an input (uncoded) speech signal, curve b) represents a decoded reference speech signal synthesis when the onset frame is coded using the GC codec mode, and curve c) represents a decoded improved speech signal synthesis when the entire first and second stage onset detection is employed in the onset frame and the TC codec mode is employed. Comparing curves b) and c), it can be seen that when the onset detection operation 400 and the onset detector 450 and the TC codec mode are employed in the onset frame, the coding of the onset (the onset of low to high amplitude, such as Figure 6Furthermore, in curves b) and c), the frames after the start are all encoded and decoded using the GC codec mode. It can be seen that in curve c), the encoding and decoding quality of the frames after the start is also improved. This is because the adaptive codebook in the GC codec mode in the frames after the start utilizes the good excitation established when the start frame was encoded and decoded using the TC codec mode.
[0130] Figure 7 is a simplified block diagram of an example configuration of hardware components forming an apparatus for detecting an attack in a sound signal to be encoded and for encoding and decoding the detected attack, and implementing a method for detecting an attack in a sound signal to be encoded and encoding the detected attack.
[0131] The device for detecting an attack in a sound signal to be coded and for coding the detected attack can be implemented as part of a mobile terminal, as part of a portable media player or in any similar device. Figure 7 ) includes an input 702, an output 704, a processor 706, and a memory 708.
[0132] The input 702 is configured to receive, for example, a digital input sound signal 105 ( Figure 1 ). The output 704 is configured to provide the coded bit stream 111. The input 702 and the output 704 may be implemented in a common module such as a serial input / output device.
[0133] The processor 706 is operatively connected to the input 702, to the output 704, and to the memory 708. The processor 706 is implemented as one or more processors for executing code instructions to support various modules of the voice encoder 106 (including Figure 2 、 Figure 3 and Figure 4 module) functions.
[0134] The memory 708 may include non-transitory memory for storing code instructions executable by the processor 706. Specifically, the memory may include processor-readable memory including non-transitory memory instructions that, when executed, cause the processor to implement the operations and modules of the voice encoder 106, including Figure 2 、 Figure 3 and Figure 4 The memory 708 may also include random access memory or one or more buffers to store intermediate processing data of various functions performed by the processor 706.
[0135] Those skilled in the art will recognize that the description of the method and apparatus for detecting an attack in a sound signal to be coded and encoding the detected attack is illustrative only and is not intended to be limiting in any way. Other embodiments will readily suggest themselves to those skilled in the art having the benefit of this disclosure. Furthermore, the disclosed method and apparatus for detecting an attack in a sound signal to be coded and encoding the detected attack can be customized to provide valuable solutions to existing needs and problems associated with the allocation or distribution of bit budgets.
[0136] For the sake of clarity, not all conventional features of embodiments of the method and apparatus for detecting an attack in a sound signal to be encoded and for encoding and decoding a detected attack are shown and described. Of course, it will be appreciated that in developing any such actual implementation of the method and apparatus for detecting an attack in a sound signal to be encoded and for encoding and decoding a detected attack, many specific implementation decisions may need to be made to achieve the developer's specific goals, such as compliance with application, system, network, and business-related constraints, and that these specific goals will vary from one embodiment to another and from one developer to another. Furthermore, it will be appreciated that the development effort may be complex and time-consuming, but will nevertheless be a routine undertaking for a person of ordinary skill in the art of sound processing who benefits from the present disclosure.
[0137] According to the present disclosure, the modules, processing operations and / or data structures described herein can be implemented using various types of operating systems, computing platforms, network devices, computer programs and / or general-purpose machines. In addition, those skilled in the art will recognize that less versatile devices such as hard-wired devices, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), or similar devices can also be used. When a method comprising a series of operations and sub-operations is implemented by a processor, computer, or machine, and these operations and sub-operations can be stored as a series of non-transitory code instructions that can be read by a processor, computer, or machine, they can be stored on a tangible and / or non-transitory medium.
[0138] The modules of the method and apparatus for detecting an attack in a sound signal to be encoded and decoded and for encoding and decoding the detected attack as described herein may include software, firmware, hardware, or any combination of software, firmware, or hardware suitable for the purposes described herein.
[0139] In the methods and apparatus for detecting an attack in a sound signal to be encoded and decoded and for encoding and decoding the detected attack as described herein, various operations and sub-operations may be performed in various orders, and some operations and sub-operations may be optional.
[0140] Although the present invention, as disclosed above, has been presented by way of non-limiting illustrative embodiments, these embodiments may be modified at will within the scope of the appended claims without departing from the spirit and nature of the invention.
[0141] References
[0142] The following references are mentioned in this specification, the entire contents of which are incorporated herein by reference.
[0143] [1]V.Eksler, R.Salami, and M.Jelínek, "Efficient handling of moderating and speech transitions in the EVS codec," in Proc.IEEE Int.Conf.onAcoustics, Speech and Signal Processing (ICASSP), Brisbane, Australia, 2015.
[0144] [2] V.Eksler, M.Jelínek, and R.Salami, "Method and Device for theEncoding of Transition Frames in Speech and Audio," WIPO Patent Application No.WO / 2008 / 049221, 24Oct.2006.
[0145] [3]V.Eksler and M.Jelínek, "Glottal-Shape Codebook to Improve Robustness of CELP Codecs," IEEE Trans.on Audio, Speech and Language Processing, vol.18, no.6, pp.1208–1217, Aug.2010.
[0146] [4]3GPP TS 26.445: "Codec for Enhanced Voice Services (EVS); DetailedAlgorithmic Description".
[0147] As additional disclosure, the following is pseudo-code for a non-limiting example of implementing the disclosed onset detector in an Immersive Voice and Audio Services (IVAS) codec.
[0148] This pseudocode is based on EVS. The new IVAS logic is highlighted in the shaded background.
[0149]
[0150]
[0151]
[0152]
[0153]
[0154]
[0155]
Claims
1. A device for encoding and decoding an attack tone in a sound signal, comprising: The apparatus for detecting the attack of the sound signal includes: a first-stage onset detector for detecting an onset in a last subframe of a current frame; and a second-stage attack detector used if the first-stage attack detector does not detect an attack, for detecting an attack in one of the subframes of the current frame, including a subframe before the last subframe; and An encoder encodes a subframe including a detected onset using a codec mode with a non-predictive codebook, wherein the codec mode is a transition codec mode and the non-predictive codebook is a glottal shape codebook populated with a glottal pulse shape.
2. The apparatus of claim 1 , further comprising a decision module configured to determine that a current frame is an active frame that was previously classified as being encoded using the universal codec mode, and to indicate that no onset is detected when the current frame is not determined to be an active frame that was previously classified as being encoded using the universal codec mode.
3. The device according to claim 1 or 2, comprising: a calculator for calculating energy of a sound signal in a plurality of analysis segments in a current frame, wherein a subset of the plurality of analysis segments corresponds to a subframe; as well as A finder that finds one of the analysis segments having a maximum energy that represents a candidate attack position to be verified by the first stage attack detector and the second stage attack detector.
4. The device according to claim 3, wherein The first stage attack detector includes: a calculator for calculating a first average energy across the analysis segment prior to the last subframe in the current frame; and A calculator calculates a second average energy across the analysis segments in the current frame starting from the analysis segment with the maximum energy to the last analysis segment in the current frame.
5. The device according to claim 4, wherein The first stage attack detector includes: A first comparator that compares the ratio between the first average energy and the second average energy with: a first threshold; or A second threshold when the classification of the previous frame is voiced.
6. The device according to claim 5, wherein When the comparison indication of the first comparator detects the first stage attack, the first stage attack detector includes: A second comparator compares the ratio between the energy of the analysis segment of maximum energy and the energies of the other analysis segments of the current frame with a third threshold.
7. The apparatus according to claim 6, when the comparison between the first comparator and the second comparator indicates that the first-stage attack position is an analysis segment with maximum energy representing a candidate attack position, comprising: A decision module is used to determine whether the first-stage onset position is equal to or greater than the number of analysis segments before the last subframe of the current frame. If the first-stage onset position is equal to or greater than the number of analysis segments before the last subframe, the detected onset position is determined to be the first-stage onset position in the last subframe of the current frame.
8. The apparatus of claim 1 , comprising a decision module for determining whether a current frame is classified as voiced speech, and wherein: When the current frame is not classified as voiced, the second stage onset detector is used.
9. The apparatus according to claim 3, wherein The second stage attack detector comprises a calculator that calculates the average energy of the sound signal across the analysis segments preceding the analysis segment with the maximum energy representing the candidate attack position.
10. The apparatus according to claim 9, wherein The analysis segments preceding the analysis segment with the maximum energy representing the candidate attack position include analysis segments from the previous frame.
11. The apparatus according to claim 9, wherein The second stage attack detector includes: A first comparator that compares the ratio between the energy of the analysis segment representing the candidate attack position and the calculated average energy with: - a first threshold; or - A second threshold when the classification of the previous frame was unvoiced.
12. The apparatus according to claim 11, wherein When the comparison indication of the first comparator of the second stage attack detector is detected, the second stage attack detector includes: A second comparator compares the ratio between the energy of the analysis segment representing the candidate attack position and the long-term energy of the analysis segment with a third threshold.
13. The apparatus according to claim 12, wherein When the attack is detected in the previous frame, the second comparator of the second stage attack detector does not detect the attack.
14. The apparatus of claim 12, when the comparison between the first comparator and the second comparator of the second stage attack detector indicates that the second stage attack position is an analysis segment with maximum energy representing a candidate attack position, comprising: The decision module is used to determine the detected onset position as the second-stage onset position.
15. The apparatus according to claim 1, wherein The onset detection device determines a subframe to be coded using the transitional coding mode based on the detected position of the onset.
16. A device for encoding and decoding an attack tone in a sound signal, comprising: at least one processor; as well as a memory coupled to the processor and comprising non-transitory instructions that, when executed, cause the processor to: The apparatus for detecting the attack of the sound signal includes: a first-stage onset detector for detecting an onset in a last subframe of a current frame; and a second-stage attack detector used if the first-stage attack detector does not detect an attack, for detecting an attack in a subframe before the last subframe of the current frame; and An encoder encodes a subframe including a detected onset using a codec mode with a non-predictive codebook, wherein the codec mode is a transition codec mode and the non-predictive codebook is a glottal shape codebook populated with a glottal pulse shape.
17. A device for encoding and decoding an attack tone in a sound signal, comprising: at least one processor; as well as a memory coupled to the processor and comprising non-transitory instructions that, when executed, cause the processor to: Detecting the onset of the sound signal, wherein the sound signal is processed in consecutive frames, each frame including a plurality of subframes, and detecting the onset of the sound signal includes: In the first stage, the onset is detected in the last subframe of the current frame; and If no attack is detected in the first stage, in the second stage, detecting an attack in a subframe before the last subframe of the current frame; and A subframe including the detected onset is encoded using a codec mode having a non-predictive codebook, wherein the codec mode is a transition codec mode and the non-predictive codebook is a glottal shape codebook populated with a glottal pulse shape.
18. A method for encoding and decoding an attack tone in a sound signal, comprising: Detecting the onset of the sound signal, wherein the sound signal is processed in consecutive frames, each frame including a plurality of subframes, and detecting the onset of the sound signal includes: The first stage attack detection is used to detect the attack in the last subframe of the current frame; and a second stage attack detection used if the first stage attack detection does not detect an attack, for detecting an attack in one of the subframes of the current frame including a subframe before the last subframe; and A subframe including the detected onset is encoded using a codec mode having a non-predictive codebook, wherein the codec mode is a transition codec mode and the non-predictive codebook is a glottal shape codebook populated with a glottal pulse shape.
19. The method of claim 18, comprising determining that the current frame is an active frame previously classified as being encoded using the universal codec mode, and indicating that no attack was detected when the current frame is not determined to be an active frame previously classified as being encoded using the universal codec mode.
20. The method according to claim 18 or 19, comprising: calculating energy of a sound signal in a plurality of analysis segments in a current frame, wherein a subset of the plurality of analysis segments corresponds to a subframe; as well as A search is performed for one of the analysis segments with the maximum energy that represents a candidate attack position to be verified by the first stage attack detection and the second stage attack detection.
21. The method according to claim 20, wherein The first stage of attack detection includes: Calculating a first average energy across the analysis segment before the last subframe in the current frame; and A second average energy is calculated across the analysis segments in the current frame starting from the analysis segment with the maximum energy to the last analysis segment of the current frame.
22. The method according to claim 21, wherein The first stage of attack detection includes: Using a first comparator, the ratio between the first average energy and the second average energy is compared to: a first threshold; or A second threshold when the classification of the previous frame is voiced.
23. The method according to claim 22, wherein When the comparison indication of the first comparator detects the first stage attack, the first stage attack detection includes: Using a second comparator, the ratio between the energy of the analysis segment of maximum energy and the energies of the other analysis segments of the current frame is compared with a third threshold.
24. The method according to claim 23, when the comparison between the first comparator and the second comparator indicates that the first-stage attack position is an analysis segment with the maximum energy representing the candidate attack position, comprising: Determine whether the first-stage attack position is equal to or greater than the number of analysis segments before the last subframe of the current frame. If the first-stage attack position is equal to or greater than the number of analysis segments before the last subframe, determine that the detected attack position is the first-stage attack position in the last subframe of the current frame.
25. The method of claim 18, comprising determining whether a current frame is classified as voiced speech, wherein: When the current frame is not classified as voiced, the second stage onset detection is used.
26. The method according to claim 20, wherein The second stage attack detection comprises calculating the average energy of the acoustic signal across the analysis segments preceding the analysis segment with the maximum energy representing the candidate attack position.
27. The method according to claim 26, wherein The analysis segments preceding the analysis segment with the maximum energy representing the candidate attack position include analysis segments from the previous frame.
28. The method according to claim 26, wherein The second stage of attack detection includes: Using a first comparator, the ratio between the energy of the analysis segment representing the candidate attack position and the calculated average energy is compared with: - a first threshold; or - A second threshold when the classification of the previous frame was unvoiced.
29. The method according to claim 28, wherein When the comparison indication of the first comparator of the second-stage attack detection detects the second-stage attack, the second-stage attack detection includes: Using a second comparator, the ratio between the energy of the analysis segment representing the candidate attack position and the long-term energy of the analysis segment is compared with a third threshold value.
30. The method according to claim 29, wherein When the attack is detected in the previous frame, the comparison of the second comparator of the second stage attack detection does not detect the attack.
31. The method of claim 29, when the comparison between the first comparator and the second comparator of the second stage attack detection indicates that the second stage attack position is an analysis segment with maximum energy representing a candidate attack position, comprising: The detected attack position is determined as the second-stage attack position.
32. The method of claim 18, comprising determining a subframe to be coded using a transitional codec mode based on the detected location of the onset.
Citation Information
Patent Citations
Method and apparatus for robust speech classification
US20020111798A1
Method and Device for Coding Transition Frames in Speech Signals
US20100241425A1