Stereo signal coding method and stereo signal coding apparatus
By optimizing residual signal encoding based on downmix and residual signal energies, the method addresses low spatial sense and sound image stability at low coding rates and high-frequency distortion at high coding rates, enhancing encoding quality.
Patent Information
- Application Number
- JP2025178710
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2018-05-31
- Filing Date
- 2025-10-23
- Publication Date
- 2026-01-21
AI Technical Summary
Existing stereo signal encoding methods suffer from low spatial sense and sound image stability at low coding rates, and high-frequency distortion at high coding rates due to insufficient bit allocation for high-frequency information.
A method that determines residual signal coding parameters based on downmix and residual signal energies of subbands to decide whether to encode residual signals, optimizing bit allocation to improve spatial sense and sound image stability while minimizing high-frequency distortion.
The method enhances the spatial sense and sound image stability of decoded stereo signals while reducing high-frequency distortion, thereby improving overall encoding quality.
Smart Images

Figure 2026010182000001_ABST
Abstract
Description
[Technical Field]
[0001] This application claims priority to Chinese Patent Application No. 201810549237.3, entitled "STEREO SIGNAL ENCODING METHOD AND APPARATUS," filed with the China Patent Office on May 31, 2018, which is incorporated herein by reference in its entirety.
[0002] The present application relates to the audio field, and more particularly to a method and an apparatus for encoding a stereo signal. [Background technology]
[0003] The general process for encoding a stereo signal using time-domain or time-frequency-domain stereo coding techniques is as follows. performing time-domain preprocessing on the left channel time-domain signal and the right channel time-domain signal; performing a time domain analysis on the left channel time domain signal and the right channel time domain signal obtained by the time domain preprocessing; performing a time-frequency domain transform on the left channel time domain signal and the right channel time domain signal obtained by the time domain preprocessing to obtain a left channel frequency domain signal and a right channel frequency domain signal; determining an Inter-channel Time Difference (ITD) parameter in the time domain; performing a time shift adjustment on the left frequency domain signal and the right channel frequency domain signal based on the ITD parameters; The stereo parameters, the downmix signal, and the residual signal are calculated based on the left channel frequency domain signal and the right channel frequency domain signal obtained by the time shift adjustment, and the stereo parameters, the downmix signal, and the residual signal are coded.
[0004] It is known in the prior art that when the coding rate is relatively low, only the stereo parameters and the downmix signal are generally coded, and only when the coding rate is relatively high, a part or all of the residual signal is coded, which results in a relatively low spatial sense of the decoded stereo signal and a relatively low sound image stability of the decoded stereo signal.
[0005] In other prior art, when the coding rate is relatively low, in addition to the downmix signal, a residual signal of a subband satisfying a preset bandwidth range is also coded. This coding method can improve the spatial sense and sound image stability of the decoded stereo signal, but because the total number of coding bits used for coding the residual signal and the downmix signal is fixed and low-frequency information is preferentially coded during the downmix signal coding, when the downmix signal is to be coded, there may not be enough bits to code some signals with the richer high-frequency information in the downmix signal. Therefore, the high-frequency distortion of the decoded stereo signal is relatively large, which affects the coding quality. Summary of the Invention [Means for solving the problem]
[0006] The present application provides a stereo signal encoding method that improves the spatial sense and sound image stability of the decoded stereo signal and can reduce high frequency distortion of the decoded stereo signal as much as possible, thereby improving the encoding quality.
[0007] According to a first aspect, there is provided a stereo signal coding method, comprising: determining a residual signal coding parameter of a current frame of the stereo signal based on a downmix signal energy and a residual signal energy of each of M subbands of the current frame, where the residual signal coding parameter of the current frame is used to indicate whether to code the residual signal of the M subbands, the M subbands being at least a portion of the N subbands, N being a positive integer greater than 1, and M≦N, where M is a positive integer; and determining whether to code the residual signal of the M subbands of the current frame based on the residual signal coding parameter of the current frame.
[0008] The residual signal coding parameter is determined based on the downmix signal energy and the residual signal energy of M subbands within the N subbands that satisfy a preset bandwidth range, and whether to encode the residual signal of each of the M subbands is determined based on the residual signal coding parameter. This avoids encoding only the downmix signal when the coding rate is relatively low. Alternatively, whether to encode all the residual signals of the subbands that satisfy the preset bandwidth range is determined based on the residual signal coding parameter. This improves the spatial sense and sound image stability of the decoded stereo signal while minimizing high-frequency distortion of the decoded stereo signal, thereby improving coding quality.
[0009] Referring to the first aspect, in one possible implementation example of the first aspect, the M subbands are M subbands within the N subbands whose subband index numbers are less than or equal to a preset maximum subband index number.
[0010] Optionally, in one embodiment, the M subbands are M subbands within the N subbands whose subband index numbers are greater than or equal to a preset minimum subband index number and less than or equal to a preset maximum subband index number.
[0011] The minimum subband index number and / or the maximum subband index number are set based on different coding rates. Residual signal coding parameters are determined based on the different coding rates and the downmix signal energies and residual signal energies of specific subbands among the N subbands. Whether to code the residual signal of each of the M subbands is determined based on the residual signal coding parameters. This avoids coding only the downmix signal when the coding rate is relatively low. Alternatively, whether to code all the residual signals of subbands that satisfy a preset bandwidth range is determined based on the residual signal coding parameters. This improves the spatial sense and sound image stability of the decoded stereo signal while minimizing high-frequency distortion of the decoded stereo signal, thereby improving coding quality.
[0012] Referring to the first aspect, in one possible implementation example of the first aspect, the step of determining whether to encode the residual signal of each of the M subbands based on the residual signal coding parameter of the current frame includes the steps of comparing the residual signal coding parameter of the current frame with a preset first threshold, where the first threshold is greater than 0 and less than 1.0, and determining not to encode the residual signal of each of the M subbands if the residual signal coding parameter of the current frame is less than or equal to the first threshold, or determining to encode the residual signal of each of the M subbands if the residual signal coding parameter of the current frame is greater than the first threshold.
[0013] A first threshold is set, and the determined residual signal coding parameter is compared with the first threshold. Whether to code the residual signal of each of the M subbands is determined based on the comparison result between the residual signal coding parameter and the first threshold. This avoids encoding only the downmix signal when the coding rate is relatively low. Alternatively, whether to code all the residual signals of the subbands that satisfy the preset bandwidth range is determined based on the comparison result between the residual signal coding parameter and the first threshold. This improves the spatial sense and sound image stability of the decoded stereo signal, while simultaneously minimizing high-frequency distortion of the decoded stereo signal, thereby improving coding quality.
[0014] Referring to the first aspect, in one possible implementation example of the first aspect, determining a residual signal coding parameter of the current frame based on the downmix signal energy and the residual signal energy of each of the M subbands includes determining a residual signal coding parameter based on the downmix signal energy, the residual signal energy, and a side gain of each of the M subbands.
[0015] The residual signal coding parameters for each of the M subbands are determined based on the downmix signal energy, the residual signal energy, and the side gain, and whether to code the residual signal for each of the M subbands is determined based on the residual signal coding parameters. This avoids coding only the downmix signal when the coding rate is relatively low. Alternatively, whether to code all the residual signals of the subbands that satisfy a preset bandwidth range is determined based on the residual signal coding parameters. This improves the spatial sense and sound image stability of the decoded stereo signal while minimizing high-frequency distortion of the decoded stereo signal, thereby improving coding quality.
[0016] Referring to the first aspect, in one possible implementation example of the first aspect, the step of determining a residual signal coding parameter based on the downmix signal energy, the residual signal energy, and the side gain of each of the M subbands comprises the steps of: determining a first parameter based on the downmix signal energy, the residual signal energy, and the side gain of each of the M subbands, wherein the first parameter indicates a value relationship between the downmix signal energy and the residual signal energy of each of the M subbands; and determining a second parameter based on the downmix signal energy and the residual signal energy of each of the M subbands, the second parameter indicating a value relationship between the first energy sum and the second energy sum, the first energy sum being a sum of residual signal energies and downmix signal energies of the M subbands, the second energy sum being a sum of residual signal energies and downmix signal energies of the M subbands in a frequency domain signal of a frame preceding the current frame, the M subbands of the current frame having the same subband index numbers as the M subbands of the previous frame; and determining a residual signal coding parameter of the current frame based on the first parameter, the second parameter, and a long-term smoothing parameter of the frame preceding the current frame.
[0017] Referring to the first aspect, in one possible implementation example of the first aspect, the step of determining the first parameter based on the downmix signal energy, the residual signal energy, and the side gain of each of the M subbands includes the steps of determining M energy parameters based on the downmix signal energy, the residual signal energy, and the side gain of each of the M subbands, where the M energy parameters respectively indicate a value relationship between the downmix signal energy and the residual signal energy of each of the M subbands, and the M energy parameters correspond one-to-one to the M subbands; and determining an energy parameter having a maximum value among the M energy parameters as the first parameter.
[0018] Referring to the first aspect, in one possible implementation example of the first aspect, an energy parameter of a subband having a subband index number b among the M energy parameters satisfies the following formula: res_dmx_ratio[b]=res_cod_NRG_S[b] / (res_cod_NRG_S[b]+(1-g(b))·(1-g(b))·res_cod_NRG_M[b]+1) In the formula, res_dmx_ratio[b] represents the energy parameter of the subband having subband index number b, where b is greater than or equal to 0 and less than or equal to the preset maximum subband index number, res_cod_NRG_S[b] represents the residual signal energy of the subband having subband index number b, res_cod_NRG_M[b] represents the downmix signal energy of the subband having subband index number b, and g(b) represents a function of the side gain side_gain[b] of the subband having subband index number b.
[0019] Referring to the first aspect, in one possible implementation example of the first aspect, the step of determining a residual signal coding parameter of the current frame based on the downmix signal energy and the residual signal energy of each of the M subbands includes the steps of: determining a first parameter based on the downmix signal energy and the residual signal energy of each of the M subbands, where the first parameter indicates a value relationship between the downmix signal energy and the residual signal energy of each of the M subbands; and determining a second parameter based on the downmix signal energy and the residual signal energy of each of the M subbands, where the second parameter indicates a value relationship between the downmix signal energy and the residual signal energy of each of the M subbands. the parameter indicates a value relationship between the first energy sum and the second energy sum, the first energy sum being the sum of the residual signal energy and the downmix signal energy of the M subbands, the second energy sum being the sum of the residual signal energy and the downmix signal energy of the M subbands in the frequency domain signal of the frame preceding the current frame, and the M subbands of the current frame having the same subband index numbers as the M subbands of the previous frame; and determining a residual signal coding parameter of the current frame based on the first parameter, the second parameter, and the long-term smoothing parameter of the frame preceding the current frame.
[0020] Referring to the first aspect, in one possible implementation example of the first aspect, the step of determining the first parameter based on the downmix signal energy and the residual signal energy of each of the M subbands includes the steps of determining M energy parameters based on the downmix signal energy and the residual signal energy of each of the M subbands, where the M energy parameters respectively indicate a value relationship between the downmix signal energy and the residual signal energy of each of the M subbands, and the M energy parameters correspond one-to-one to the M subbands; and determining an energy parameter having a maximum value among the M energy parameters as the first parameter.
[0021] Optionally, in one embodiment, the sum of M energy parameters is determined as the first parameter (to be corrected) res_dmx_ratio1, and res_dmx_ratio1 is corrected based on the maximum value res_dmx_ratio_max among the M energy parameters and the downmix signal energy res_cod_NRG_M[b] of each of the M subbands, and res_dmx_ratio2 obtained by the correction is determined.
[0022] For example, the encoder side corrects res_dmx_ratio1 according to the following formula, where M=5: The res_dmx_ratio2 obtained by the correction satisfies the following formula.
number
[0023] Optionally, in one embodiment, the res_dmx_ratio2 obtained by the correction may be further corrected.
[0024] For example, the final res_dmx_ratio3 obtained by the correction satisfies the following formula: res_dmx_ratio3=pow(res_dmx_ratio2,1.2) In the formula, the pow() function represents an exponential function, and pow(res_dmx_ratio2,1.2) represents res_dmx_ratio2 raised to the power of 1.2.
[0025] Optionally, in an embodiment, the encoder side determines the first parameter based on a sum of residual signal energies of the M subbands and a sum of downmix signal energies of the M subbands.
[0026] Specifically, the encoder side separately determines the sum of downmix signal energies of the M subbands dmx_nrg_all_curr and the sum of residual signal energies of the M subbands res_nrg_all_curr, and determines the first parameter based on dmx_nrg_all_curr and res_nrg_all_curr.
[0027] Optionally, in one embodiment, the sum of the downmix signal energies of the M subbands dmx_nrg_all_curr satisfies the following equation:
number
[0028] Optionally, in one embodiment, the sum of the residual signal energies of the M subbands, res_nrg_all_curr, satisfies the following equation:
number
[0029] The encoder side determines the first parameter res_dmx_ratio based on dmx_nrg_all_curr and res_nrg_all_curr.
[0030] For example, the first parameter res_dmx_ratio finally determined by the encoder side satisfies the following formula: res_dmx_ratio=res_nrg_all_curr / dmx_nrg_all_curr
[0031] Referring to the first aspect, in one possible implementation example of the first aspect, an energy parameter of a subband having a subband index number b among the M energy parameters satisfies the following formula: res_dmx_ratio[b]=res_cod_NRG_S[b] / res_cod_NRG_M[b] In the formula, res_dmx_ratio[b] represents the energy parameter of the subband having subband index number b, where b is greater than or equal to 0 and less than or equal to the preset maximum subband index number, res_cod_NRG_S[b] represents the residual signal energy of the subband having subband index number b, and res_cod_NRG_M[b] represents the downmix signal energy of the subband having subband index number b.
[0032] Referring to the first aspect, in one possible implementation of the first aspect, the residual signal coding parameter of the current frame is a long-term smoothing parameter of the current frame, and the long-term smoothing parameter of the current frame satisfies the following formula: res_dmx_ratio_lt=res_dmx_ratio·α+res_dmx_ratio_lt_prev·(1-α) where res_dmx_ratio_lt represents the long-term smoothing parameter of the current frame, res_dmx_ratio represents the first parameter, and res_dmx_ratio_lt_prev represents the long-term smoothing parameter of the frame previous to the current frame, and 0<α<1; When the second parameter is greater than a third preset threshold, the value of α when the first parameter is less than the second preset threshold is greater than the value of α when the first parameter is equal to or greater than the second preset threshold, the second threshold is greater than or equal to 0 and less than or equal to 0.6, and the third threshold is greater than or equal to 2.7 and less than or equal to 3.7; or When the second parameter is less than a fifth preset threshold, the value of α when the first parameter is greater than a fourth preset threshold is greater than the value of α when the first parameter is equal to or less than the fourth preset threshold, the fourth threshold is greater than or equal to 0.9, and the fifth threshold is greater than or equal to 0.71; or When the second parameter is greater than or equal to a preset fifth threshold and less than or equal to a preset third threshold, the value of α is less than the value of α when the first parameter is less than the preset second threshold and the second parameter is greater than the preset third threshold, the second threshold being greater than or equal to 0 and less than or equal to 0.6, the third threshold being greater than or equal to 2.7 and less than or equal to 3.7, and the fifth threshold being greater than or equal to 0 and less than or equal to 0.71.
[0033] Referring to the first aspect, in one possible implementation example of the first aspect, the method further includes a step of encoding the downmix signal and the residual signal of the M subbands when it is determined to encode the residual signals of the M subbands, or a step of encoding the downmix signal of the M subbands when it is determined not to encode the residual signals of the M subbands.
[0034] According to a second aspect, there is provided an encoding apparatus, including: a first decision module configured to determine residual signal coding parameters for a current frame of a stereo signal based on downmix signal energies and residual signal energies of each of M subbands of the current frame, where the residual signal coding parameters of the current frame are used to indicate whether to encode the residual signal of the M subbands, the M subbands being at least a portion of the N subbands, N being a positive integer greater than 1, and M≦N, where M is a positive integer; and a second decision module configured to determine whether to encode the residual signal of the M subbands based on the residual signal coding parameters of the current frame.
[0035] According to a third aspect, there is provided an encoding apparatus, the apparatus comprising: a memory and a processor, the memory configured to store a program, and the processor configured to execute the program, the program, when executed, causing the processor to perform a method according to the first aspect or any one of the possible implementations of the first aspect.
[0036] According to a fourth aspect, there is provided a computer-readable storage medium storing program code to be executed by a device, the program code including instructions used to perform a method according to the first aspect or various implementations of the first aspect.
[0037] According to a fifth aspect, there is provided a chip, the chip including a processor and a communication interface, the communication interface configured to communicate with an external device, the processor configured to perform a method according to the first aspect or any one of the possible implementations of the first aspect.
[0038] Optionally, in one embodiment, the chip may further include a memory, the memory storing instructions, and the processor configured to execute the instructions stored in the memory, the instructions, when executed, configuring the processor to perform a method according to the first aspect or any one of the possible implementations of the first aspect.
[0039] Optionally, in one embodiment, the chip is incorporated into a terminal device or a network device. [Brief explanation of the drawings]
[0040] [Figure 1] FIG. 1 is a schematic structural diagram of stereo encoding and decoding in the time domain according to an embodiment of the present application; [Figure 2] 1 is a schematic diagram of a mobile terminal according to an embodiment of the present application; [Figure 3] 1 is a schematic diagram of a network element according to an embodiment of the present application; [Figure 4]1 is a schematic flow diagram of a method for encoding a stereo signal in the frequency domain; [Figure 5] 1 is a schematic flow diagram of a method for encoding a stereo signal in the time-frequency domain; [Figure 6] 1 is a schematic flow chart of a stereo signal encoding method according to an embodiment of the present application; [Figure 7] 4 is another schematic flow chart of a stereo signal encoding method according to an embodiment of the present application; [Figure 8] 1 is a schematic block diagram of a stereo signal encoding device according to an embodiment of the present application; [Figure 9] FIG. 2 is another schematic block diagram of a stereo signal encoding device according to an embodiment of the present application; DETAILED DESCRIPTION OF THE INVENTION
[0041] The technical solutions of the present application are described below with reference to the accompanying drawings.
[0042] 1 is a schematic structural diagram of a stereo encoding and decoding system in the time domain according to an exemplary embodiment of the present application. The stereo encoding and decoding system includes an encoding component 110 and a decoding component 120.
[0043] The encoding component 110 is configured to encode the stereo signal in the time domain. Optionally, the encoding component 110 may be implemented using software, or hardware, or a combination of software and hardware, which is not limited in this embodiment.
[0044] The encoding component 110 encodes the stereo signal in the time domain and includes the following steps:
[0045] (1) Time-domain preprocessing is performed on the obtained stereo signal to obtain a time-domain preprocessed left channel signal and a time-domain preprocessed right channel signal.
[0046] The stereo signal is collected by a collection component and sent to an encoding component 110. Optionally, the collection component and the encoding component 110 may be located on the same device. Alternatively, the collection component and the encoding component 110 may be located on different devices.
[0047] The left channel signal obtained by preprocessing and the right channel signal obtained by preprocessing are two channel signals of a stereo signal obtained by preprocessing.
[0048] Optionally, the pre-processing includes at least one of high-pass filtering, pre-emphasis, sampling rate conversion, and channel conversion, which is not limited in this embodiment.
[0049] (2) Performing delay estimation based on the pre-processed left channel signal and the pre-processed right channel signal to obtain an inter-channel time difference between the pre-processed left channel signal and the pre-processed right channel signal.
[0050] (3) Based on the inter-channel time difference, a delay adjustment process is performed on the left channel signal obtained by preprocessing and the right channel signal obtained by preprocessing to obtain a left channel signal obtained by delay matching process and a right channel signal obtained by delay matching process.
[0051] (4) Encoding the inter-channel time difference to obtain an inter-channel time difference coding index.
[0052] (5) Calculating stereo parameters used in the time-domain down-mixing process, encoding the stereo parameters used in the time-domain down-mixing process, and obtaining encoding indexes of the stereo parameters used in the time-domain down-mixing process.
[0053] The stereo parameters used for the time-domain downmixing process are used to perform the time-domain downmixing process on the left channel signal obtained by the delay matching process and the right channel signal obtained by the delay matching process.
[0054] (6) Based on the stereo parameters used in the time-domain downmixing process, a time-domain downmixing process is performed on the left channel signal obtained by the delay matching process and the right channel signal obtained by the delay matching process to obtain a primary channel signal and a secondary channel signal.
[0055] The primary channel signal is used to represent information about the correlation between channels. The secondary channel signal is used to represent information about the difference between channels. When the left channel signal obtained by delay matching processing and the right channel signal obtained by delay matching processing are matched in the time domain, the secondary channel signal is minimized. In this case, the stereo signal has the best effect.
[0056] (7) Separately encode the primary channel signal and the secondary channel signal to obtain a first mono encoded bitstream corresponding to the primary channel signal and a second mono encoded bitstream corresponding to the secondary channel signal.
[0057] (8) Writing the inter-channel time difference coding index, the stereo parameter coding index, the first mono coded bitstream, and the second mono coded bitstream into a stereo coded bitstream.
[0058] The decoding component 120 is configured to decode the stereo encoded bitstream produced by the encoding component 110 to obtain a stereo signal.
[0059] Optionally, the encoding component 110 is connected to the decoding component 120 via a wired or wireless connection, and the decoding component 120 obtains the stereo encoded bitstream generated by the encoding component 110 over this connection. Alternatively, the encoding component 110 stores the generated stereo encoded bitstream in a memory, and the decoding component 120 reads the stereo encoded bitstream from the memory.
[0060] Optionally, the decoding component 120 may be implemented using software, or may be implemented using hardware, or may be implemented in the form of a combination of software and hardware, which is not limited in this embodiment.
[0061] The decoding component 120 decodes the stereo encoded bitstream to obtain a stereo signal, which includes the following steps:
[0062] (1) A first mono coded bitstream and a second mono coded bitstream in the stereo coded bitstream are decoded to obtain a primary channel signal and a secondary channel signal.
[0063] (2) Based on the stereo coded bitstream, a coding index of stereo parameters to be used for a time-domain upmix process is obtained, and the time-domain upmix process is performed on the primary channel signal and the secondary channel signal to obtain a left channel signal obtained by the time-domain upmix process and a right channel signal obtained by the time-domain upmix process.
[0064] (3) A coding index of the inter-channel time difference is obtained based on the stereo coded bitstream, and a stereo signal is obtained by performing delay adjustment on the left channel signal obtained by the time domain upmix processing and the right channel signal obtained by the time domain upmix processing.
[0065] Optionally, the encoding component 110 and the decoding component 120 may be located in the same device or in different devices. The device may be a mobile terminal having an audio signal processing function, such as a mobile phone, a tablet computer, a laptop portable computer, a desktop computer, a Bluetooth speaker, a pen recorder, or a wearable device, or may be a network element having an audio signal processing capability in a core network or a wireless network. This is not limited in this embodiment.
[0066] For example, as shown in FIG. 2, in this embodiment, the encoding component 110 is located in the mobile terminal 130, the decoding component 120 is located in the mobile terminal 140, and the mobile terminal 130 and the mobile terminal 140 are mutually independent devices having audio signal processing capabilities, and may be, for example, a mobile phone, a wearable device, a virtual reality (VR) device, an augmented reality (AR) device, etc., and the description will be given using an example in which the mobile terminal 130 is connected to the mobile terminal 140 using a wireless or wired network.
[0067] Optionally, the mobile terminal 130 includes a collection component 131, an encoding component 110, and a channel encoding component 132. The collection component 131 is connected to the encoding component 110, and the encoding component 110 is channel The encoder component 132 is connected to the encoder component 132 .
[0068] Optionally, the mobile terminal 140 includes an audio playback component 141, a decoding component 120, and a channel decoding component 142. The audio playback component 141 includes a decoding component 120. 120 connected to the decryption component 120 is the channel decoding component 142 is connected to.
[0069] After collecting the stereo signal using the collection component 131, the mobile terminal 130 encodes the stereo signal using the encoding component 110 to obtain a stereo encoded bitstream, and then encodes the stereo encoded bitstream using the channel encoding component 132 to obtain a transmission signal.
[0070] Mobile terminal 130 sends transmissions to mobile terminal 140 using a wireless or wired network.
[0071] After receiving the transmitted signal, the mobile terminal 140 decodes the transmitted signal using a channel decoding component 142 to obtain a stereo coded bitstream, and 120 to obtain a stereo signal, and an audio playback component 141 to play a stereo signal.
[0072] For example, as shown in FIG. 3, in this embodiment, an example is given in which the encoding component 110 and the decoding component 120 are located in a network element 150 having audio signal processing capabilities within the same core network or wireless network.
[0073] Optionally, network element 150 includes a channel decoding component 151, a decoding component 120, an encoding component 110, and a channel encoding component 152. Channel decoding component 151 is connected to decoding component 120, decoding component 120 is connected to encoding component 110, and encoding component 110 is connected to channel encoding component 152.
[0074] After receiving the transmission signal sent by other equipment, the channel decoding component 151 decodes the transmission signal to obtain a first stereo encoded bitstream, and the decoding component 120 1stThe stereo encoded bitstream is decoded to obtain a stereo signal, the encoding component 110 encodes the stereo signal to obtain a second stereo encoded bitstream, and the channel encoding component 152 encodes the second stereo encoded bitstream to obtain a transmission signal.
[0075] The other device may be a mobile terminal having an audio signal processing capability, or may be another network element having an audio signal processing capability, which is not limited in this embodiment.
[0076] Optionally, encoding component 110 and decoding component 120 in the network element may transcode the stereo encoded bitstream transmitted by the mobile terminal.
[0077] Optionally, in this embodiment, the device in which the encoding component 110 is installed is called an audio encoding device. In actual implementation, the audio encoding device may also have an audio decoding function, which is not limited in this embodiment.
[0078] Optionally, this embodiment is described using only a stereo signal as an example. In this application, the audio encoding apparatus may further process a multi-channel signal, and the multi-channel signal includes signals of at least two channels.
[0079] To facilitate understanding of the stereo signal encoding method in the embodiment of the present application, the entire encoding process of the frequency domain stereo encoding method and the time-frequency domain stereo encoding method will first be generally described below with reference to FIG. 4 and FIG. 5, respectively.
[0080] 4 is a schematic flow chart of a frequency domain stereo signal encoding method, which specifically includes steps 101 to 107.
[0081] 101: Convert the time domain stereo signal into a frequency domain stereo signal.
[0082] 102: Extract frequency domain stereo parameters in the frequency domain.
[0083] 103: Perform a downmix process on the frequency domain stereo signal to obtain a downmix signal and a residual signal.
[0084] The downmix signal is also called the central channel signal or primary channel signal. residual The signals may be referred to as side channel signals or secondary channel signals.
[0085] 104: Encode the downmix signal to obtain coding parameters corresponding to the downmix signal, and write the coding parameters into the coded bitstream.
[0086] 106: Encode the frequency-domain stereo parameters to obtain coding parameters corresponding to the frequency-domain stereo parameters, and write the coding parameters into a coded bitstream.
[0087] In an optional implementation, the method may further include: 105: encoding the residual signal to obtain coding parameters corresponding to the residual signal, and writing the coding parameters into a coded bitstream.
[0088] 107: Multiplex the bitstream.
[0089] 5 is a schematic flow chart of a stereo signal encoding method in the time-frequency domain, which specifically includes steps 201 to 208.
[0090] 201: A time domain analysis is performed on the stereo signal to extract time domain stereo parameters.
[0091] 202: Transform the time domain stereo signal into a frequency domain stereo signal.
[0092] 203: Extract frequency domain stereo parameters in the frequency domain.
[0093] 204: Perform a downmix process on the frequency domain stereo signal to obtain a downmix signal and a residual signal.
[0094] 205: Encode the downmix signal to obtain coding parameters corresponding to the downmix signal, and write the coding parameters into the coded bitstream.
[0095] 207: Encode the time domain stereo parameters and the frequency domain stereo parameters to obtain coding parameters corresponding to the time domain stereo parameters and coding parameters corresponding to the frequency domain stereo parameters, and write the coding parameters into a coded bitstream.
[0096] Optionally, the method further comprises: 206: encoding the residual signal to obtain coding parameters corresponding to the residual signal, and writing the coding parameters into a coded bitstream.
[0097] 208: Multiplex the bitstream.
[0098] When the coding rate is relatively low, for example, when the coding bandwidth is wideband, such as 26 kilobytes per second (kbps), 16.4 kbps, 24.4 kbps, or 32 kbps, all residual signals of subbands satisfying a predetermined bandwidth range are coded when the downmix signal of each frame of the stereo signal is coded to improve spatial sensation and stability during playback and reduce high-frequency distortion of the stereo signal. Alternatively, when the coding rate is relatively low, only the stereo parameters and the downmix signal are coded. Part or all of the residual signal is coded only when the coding rate is relatively high, such as 48 kbps, 64 kbps, or 96 kbps. This application provides a stereo signal coding method. This method improves the spatial sensation and sound image stability of the decoded stereo signal while simultaneously minimizing high-frequency distortion of the decoded stereo signal, thereby improving overall coding quality.
[0099] 6 is a schematic flow chart of a stereo signal encoding method 300 according to an embodiment of the present application. The method 300 may be performed by an encoder side, which may be an encoder or a device having a stereo signal encoding function. The method 300 includes the following steps:
[0100] The stereo signal encoding method of the present application may be a stereo encoding method that can be applied independently or may be a stereo encoding method that is applied to multi-channel signal encoding. The encoder side processes the stereo signal frame by frame. In the following, the stereo signal encoding method of the method 300 will be described in detail using a wideband stereo signal in which the signal length of each frame is 20 ms as an example, and a frame (e.g., a current frame) being processed by the encoder side as an example.
[0101] 301: Determine a residual signal coding parameter of a current frame of a stereo signal based on the downmix signal energy and the residual signal energy of each of M subbands of the current frame, where the residual signal coding parameter of the current frame is used to indicate whether to code the residual signal of the M subbands, where the M subbands are at least a part of the N subbands, N is a positive integer greater than 1, and M≦N, M is a positive integer.
[0102] Specifically, the encoder side divides the spectral coefficients of the current frame of the stereo signal to obtain N subbands, and determines residual signal coding parameters of the current frame based on the downmix signal energy and residual signal energy of each of at least a portion of the N subbands (e.g., M subbands within the N subbands, M≦N), and the encoder side can use the residual signal coding parameters of the current frame to determine whether to code the residual signal of each of the M subbands.
[0103] 302: Determine whether to encode the residual signals of the M subbands of the current frame based on the residual signal encoding parameters of the current frame.
[0104] Specifically, the encoder side determines whether to encode the residual signal of each of the M subbands of the current frame based on the residual signal encoding parameters of the current frame determined in step 301.
[0105] If it is decided to encode the residual signal of each of the M subbands, the downmix signal and the residual signal of each of the M subbands are encoded.
[0106] If it is decided not to encode the residual signal of each of the M subbands, the downmix signal of each of the M subbands is encoded.
[0107] In one embodiment, by way of example and not limitation, the M subbands are M subbands within the N subbands whose subband index numbers are less than a preset maximum subband index number. In other words, the M subbands are subbands with relatively low frequencies within the N subbands, and specifically, the frequencies of the M subbands are lower than the frequencies of NM subbands other than the M subbands within the N subbands.
[0108] Specifically, different maximum subband index numbers are preset based on different coding rates, so that M subbands whose subband index numbers are equal to or less than the preset maximum subband index number are selected from the N subbands based on the preset maximum subband index number, and residual signal coding parameters of the current frame are determined based on the M subbands.
[0109] For example, if the coding rate is 26 kbps, N=10, M=5, and the preset maximum subband index number is set to 4, this indicates that the residual signal coding parameters of the current frame are determined based on five subbands with subband index numbers 0 to 4 among the 10 subbands.
[0110] In another example, if the coding rate is 44 kbps, N=12, M=6, and the preset maximum subband index number is set to 5, this indicates that the residual signal coding parameters of the current frame are determined based on 6 subbands with subband index numbers 0 to 5 among the 12 subbands.
[0111] In another example, if the coding rate is 56 kbps, N=12, M=7, and the preset maximum subband index number is set to 6, this indicates that the residual signal coding parameters of the current frame are determined based on seven subbands among the 12 subbands, with subband index numbers from 0 to 6.
[0112] In another embodiment, for different coding rates, the maximum subband index numbers and minimum subband index numbers of the M subbands at the different coding rates may be preset, so that M subbands whose subband index numbers are greater than or equal to the preset minimum subband index number and less than or equal to the preset maximum subband index number are selected from the N subbands based on the preset maximum subband index number and the preset minimum subband index number, and the residual signal coding parameters of the current frame are determined based on the M subbands.
[0113] For example, if the coding rate is 26 kbps, N=10, M=4, the preset minimum subband index number is set to 4, and the preset maximum subband index number is set to 7, this indicates that the residual signal coding parameters of the current frame are determined based on four subbands with subband index numbers 4 to 7 among the 10 subbands.
[0114] By way of example and not limitation, determining whether to encode the residual signal of each of the M subbands based on the residual signal coding parameters of the current frame includes determining whether to encode the residual signal of each of the M subbands based on a comparison result between the residual signal coding parameters of the current frame and a preset first threshold, where the first threshold is greater than 0 and less than 1.0; and determining not to encode the residual signal of each of the M subbands if the residual signal coding parameters of the current frame are less than or equal to the first threshold, or determining to encode the residual signal of each of the M subbands if the residual signal coding parameters are greater than the first threshold.
[0115] Specifically, the encoder side compares the residual signal coding parameter of the current frame with a preset first threshold, and decides to code the residual signal of each of the M subbands if the residual signal coding parameter of the current frame is greater than the first threshold, or decides not to code the residual signal of each of the M subbands if the residual signal coding parameter of the current frame is equal to or less than the first threshold.
[0116] For example, in one embodiment, the first threshold is 0.075. If the value of the residual signal coding parameter for the current frame is 0.06, the encoder side does not code the residual signal of each of the M subbands.
[0117] It should be understood that the value of the first threshold is merely an example, and that the first threshold may alternatively be other values greater than 0 and less than 1.0. For example, the first threshold may be 0.55, 0.46, 0.86, or 0.9.
[0118] In another optional embodiment, the encoder side may further indicate the comparison result between the residual signal coding parameter of the current frame and the first threshold using 0 or 1. For example, 0 is used to indicate that the residual signal of each of the M subbands should not be coded, and 1 is used to indicate that the residual signal of each of the M subbands should be coded. Of course, 1 may alternatively be used to indicate that the residual signal of each of the M subbands should not be coded, and 0 may alternatively be used to indicate that the residual signal of each of the M subbands should be coded.
[0119] In the following, to explain in detail how the encoder side determines the residual signal coding parameters of the current frame, an example is used in which the M subbands are subbands whose subband index numbers are less than or equal to a preset maximum subband index number (e.g., the maximum subband index number is M-1).
[0120] Method 1
[0121] The encoder side determines residual signal coding parameters of the current frame based on the downmix signal energy, the residual signal energy, and the side gain of each of the M subbands.
[0122] In one possible implementation, the encoder side determines a first parameter based on the downmix signal energy, the residual signal energy, and the side gain of each of the M subbands, where the first parameter indicates a value relationship between the downmix signal energy and the residual signal energy of each of the M subbands; determine a second parameter according to the down-mix signal energy and the residual signal energy of each of the M subbands, the second parameter indicating a value relationship between the first energy sum and the second energy sum, the first energy sum being the sum of the residual signal energy and the down-mix signal energy of the M subbands, the second energy sum being the sum of the residual signal energy and the down-mix signal energy of the M subbands in the frequency domain signal of a frame previous to the current frame, and the M subbands of the current frame have the same subband index numbers as the M subbands of the previous frame; A residual signal coding parameter of the current frame is finally determined based on the first parameter, the second parameter, and the long-term smoothing parameter of the frame preceding the current frame.
[0123] Specifically, when determining the first parameter, the encoder side determines M energy parameters based on the downmix signal energy, residual signal energy, and side gain of each of the M subbands, where the M energy parameters respectively indicate a value relationship between the downmix signal energy and the residual signal energy of one of the M subbands, and the M energy parameters correspond one-to-one to the M subbands, and the encoder side finally determines the energy parameter with the maximum value among the M energy parameters as the first parameter.
[0124] Optionally, the energy parameter of the subband with subband index number b among the M energy parameters may be determined using the following function: res_dmx_ratio[b]=f(g(b),res_cod_NRG_M[b],res_cod_NRG_S[b])(1) In the formula, res_dmx_ratio[b] represents the energy parameter of the subband having subband index number b among the M energy parameters, where b is greater than or equal to 0 and less than or equal to the preset maximum subband index number, res_cod_NRG_S[b] represents the residual signal energy of the subband having subband index number b, res_cod_NRG_M[b] represents the downmix signal energy of the subband having subband index number b, and g(b) represents a function of the side gain side_gain[b] of the subband having subband index number b.
[0125] Specifically, in one embodiment, among the M energy parameters, the energy parameter of the subband with subband index number b satisfies the following equation: res_dmx_ratio[b]=res_cod_NRG_S[b] / (res_cod_NRG_S[b]+(1-g(b))·(1-g(b))·res_cod_NRG_M[b]+1)(2)
[0126] The first parameter is denoted as res_dmx_ratio, and res_dmx_ratio satisfies the following formula: res_dmx_ratio=max(res_dmx_ratio[0],res_dmx_ratio[1],…,res_dmx_ratio[M-1])(3)
[0127] When determining the second parameter, the encoder first calculates the residual signal of the M subbands as Energy and the downmix signal of M subbands EnergyThe sum of the downmix signals of the M subbands is determined as dmx_nrg_all_curr and the residual signals of the M subbands are determined as dmx_nrg_all_curr and dmx_nrg_all_curr. Energy The sum is denoted as res_nrg_all_curr.
[0128] Optionally, the sum of the downmix signal energies of the M subbands, dmx_nrg_all_curr, satisfies the following equation:
number
[0129] It should be understood that the value of γ1 is merely an example, and that the value of γ1 may alternatively be other values greater than or equal to 0 and less than or equal to 1. For example, γ1 is 0.3, 0.5, 0.6, or 0.8.
[0130] Optionally, the sum of the residual signal energies of the M subbands, res_nrg_all_curr, satisfies the following equation:
number
[0131] It should be understood that the value of γ2 is merely an example, and that the value of γ2 may alternatively be other values greater than or equal to 0 and less than or equal to 1. For example, γ2 is 0.2, 0.5, 0.7, or 0.9.
[0132] The encoder side determines the sum of the downmix signal energies and the residual signal energies of the M subbands of the current frame (i.e., the first energy sum) based on dmx_nrg_all_curr and res_nrg_all_curr. The first energy sum is denoted as dmx_res_all.
[0133] Optionally, dmx_res_all satisfies the following formula: dmx_res_all=res_nrg_all_curr+dmx_nrg_all_curr(6)
[0134] The encoder side may further determine a sum of residual signal energies and downmix signal energies of M subbands in the frequency domain signal of a frame previous to the current frame (i.e., a second energy sum), where the M subbands of the frame previous to the current frame are: Current frame It has the same subband index numbers as the M subbands. The second energy sum is denoted as dmx_res_all_prev.
[0135] For the determination of the second energy sum dmx_res_all_prev, please refer to the method for determining the first energy sum dmx_res_all described above, and for the sake of brevity, the details will not be repeated here.
[0136] After determining the first energy sum and the second energy sum, the encoder side may determine a second parameter based on the first energy sum and the second energy sum.
[0137] Optionally, the second parameter is an inter-frame energy variation ratio, which is denoted as frame_nrg_ratio.
[0138] Optionally, in one implementation, the frame-to-frame energy variation ratio frame_nrg_ratio satisfies the following equation: frame_nrg_ratio=dmx_res_all / dmx_res_all_prev(7)
[0139] Optionally, in other implementations, the frame-to-frame energy variation ratio frame_nrg_ratio satisfies the following equation: frame_nrg_ratio=min(5.0,max(0.2,dmx_res_all / dmx_res_all_prev))(8)
[0140] The max function is used to return the larger value for the given parameters (0.2, frame_nrg_ratio_prev), and the min function is used to return the minimum value for the given parameters (5.0, max(0.2, frame_nrg_ratio_prev)). Compared with Equation (7), Equation (8) further has a correction operation, so that the frame_nrg_ratio determined using Equation (8) can better reflect the inter-frame energy variation between the current frame and the previous frame.
[0141] After determining the first parameter and the second parameter, the encoder side may determine a residual signal coding parameter of the current frame based on the first parameter, the second parameter, and the long-term smoothing parameter of the frame previous to the current frame.
[0142] For example, but not by way of limitation, the residual signal coding parameter of the current frame may be a long-term smoothing parameter of the current frame. In other words, the encoder side may determine the long-term smoothing parameter of the current frame based on the first parameter, the second parameter, and the long-term smoothing parameter of the frame preceding the current frame, and then compare the long-term smoothing parameter of the current frame with a preset first threshold to determine whether to encode the residual signal of each of the M subbands.
[0143] For example, the long-term smoothing parameters of the current frame satisfy the following equation: res_dmx_ratio_lt=res_dmx_ratio α+res_dmx_ratio_lt_prev·(1-α)(9) where res_dmx_ratio_lt represents the long-term smoothing parameter of the current frame, res_dmx_ratio represents the first parameter, and res_dmx_ratio_lt_prev represents the long-term smoothing parameter of the frame previous to the current frame, with 0<α<1.
[0144] When res_dmx_ratio_lt is calculated according to equation (9), if the value of the first parameter and / or the value of the second parameter changes, the value of the parameter α in equation (9) may also change accordingly. In other words, when the value of the first parameter and / or the value of the second parameter changes, the weight of the long-term smoothing parameter of the frame previous to the current frame in equation (9) may also change accordingly.
[0145] For example, when the second parameter is greater than a third preset threshold, the value of α when the first parameter is less than the second preset threshold is greater than the value of α when the first parameter is equal to or greater than the second preset threshold, the second threshold is greater than or equal to 0.6, and the third threshold is greater than or equal to 2.7 and less than or equal to 3.7; or When the second parameter is less than a fifth preset threshold, the value of α when the first parameter is greater than a fourth preset threshold is greater than the value of α when the first parameter is equal to or less than the fourth preset threshold, the fourth threshold is greater than or equal to 0.9, and the fifth threshold is greater than or equal to 0.71; or The value of α when the first parameter is smaller than a preset second threshold and the second parameter is greater than a preset third threshold is greater than the value of α when the second parameter is greater than or equal to a preset fifth threshold and less than or equal to the preset third threshold, the second threshold being greater than or equal to 0 and less than or equal to 0.6, the third threshold being greater than or equal to 2.7 and less than or equal to 3.7, and the fifth threshold being greater than or equal to 0 and less than or equal to 0.71.
[0146] For example, the value of the second threshold may be 0.1 and the value of the third threshold may be 3.2. Specifically, when the second parameter frame_nrg_ratio is greater than 3.2, the value of α when the first parameter res_dmx_ratio is less than 0.1 is greater than the value of α when res_dmx_ratio is 0.1 or greater; or The fourth threshold value may be 0.4, and the fifth threshold value may be 0.21. Specifically, when frame_nrg_ratio is less than 0.21, the value of α when res_dmx_ratio is greater than 0.4 is greater than the value of α when res_dmx_ratio is less than or equal to 0.4; or The second threshold value may be 0.1, the third threshold value may be 3.2, and the fifth threshold value may be 0.21. Specifically, the value of α when res_dmx_ratio is less than 0.1 and frame_nrg_ratio is greater than 3.2 is greater than the value of α when frame_nrg_ratio is greater than or equal to 0.21 and less than or equal to 3.2; or The fourth threshold value may be 0.4 and the fifth threshold value may be 0.21, and specifically, the value of α when res_dmx_ratio is greater than 0.4 and frame_nrg_ratio is less than 0.21 is greater than the value of α when frame_nrg_ratio is greater than or equal to 0.21 and less than or equal to 3.2.
[0147] Furthermore, for example, if res_dmx_ratio is less than 0.1 and frame_nrg_ratio is greater than 3.2, the value of α is 0.5, or if frame_nrg_ratio is greater than or equal to 0.21 and less than or equal to 3.2, the value of α is 0.1.
[0148] It should be noted that the values of the second to fifth thresholds and the value of α described are merely illustrative examples and do not constitute any limitation to the present application. The values of the second to fifth thresholds and the value of α may alternatively be other values in a given interval.
[0149] It should be further noted that if the current frame is the first frame processed by the encoder side, the current frame does not have a previous frame. In this case, when the long-term smoothing parameter of the current frame is determined, the long-term smoothing parameter of the frame previous to the current frame in the above equation is the preset long-term smoothing parameter. By way of example and not limitation, the value of the preset long-term smoothing parameter may be 1.0, or of course, other values such as 0.9 or 1.1.
[0150] Method 2
[0151] The method for determining residual signal coding parameters in Method 2 is similar to that of Method 1, and the difference lies in the different method for determining the first parameter. Therefore, reference may be made to the related description of determining residual signal coding parameters in Method 1. For brevity, only the method for determining the first parameter in Method 2 will be described in this specification.
[0152] By way of example and not limitation, the encoder side determines a first parameter based on the downmix signal energy and the residual signal energy of each of the M subbands, and the first parameter indicates a value relationship between the downmix signal energy and the residual signal energy of each of the M subbands.
[0153] Specifically, when determining the first parameter, the encoder side determines M energy parameters based on the downmix signal energy and residual signal energy of each of the M subbands, where the M energy parameters respectively indicate a value relationship between the downmix signal energy and the residual signal energy of one of the M subbands, the M energy parameters correspond one-to-one to the M subbands, and the encoder side finally determines the energy parameter with the maximum value among the M energy parameters as the first parameter.
[0154] Optionally, the energy parameter of the subband with subband index number b among the M energy parameters determined by the encoder side may be determined using the following function: res_dmx_ratio[b]=f(res_cod_NRG_M[b],res_cod_NRG_S[b])(10) In the formula, res_dmx_ratio[b] represents the energy parameter of the subband having subband index number b among the M energy parameters, where b is greater than or equal to 0 and less than or equal to the preset maximum subband index number, res_cod_NRG_S[b] represents the residual signal energy of the subband having subband index number b, and res_cod_NRG_M[b] represents the downmix signal energy of the subband having subband index number b.
[0155] For example, the energy parameter of the subband with subband index number b among the M energy parameters satisfies the following equation. res_dmx_ratio[b]=res_cod_NRG_S[b] / res_cod_NRG_M[b](11)
[0156] The first parameter is denoted as res_dmx_ratio, and res_dmx_ratio satisfies the following formula: res_dmx_ratio=max(res_dmx_ratio[0],res_dmx_ratio[1],…,res_dmx_ratio[M-1])(12)
[0157] After determining the first parameter, the encoder side may determine the second parameter according to the method described in Method 1, finally determine the residual signal coding parameter according to the method described in Method 1, and determine whether to code the residual signal of each of the M subbands based on the residual signal coding parameter.
[0158] Method 3
[0159] The method for determining residual signal coding parameters in Method 3 is similar to that of Method 1, and the difference lies in the different method for determining the first parameter. Therefore, reference may be made to the related description of determining residual signal coding parameters in Method 1. For brevity, this specification will only describe the method for determining the first parameter in Method 3.
[0160] By way of example and not limitation, the encoder side determines a first parameter based on the downmix signal energy and the residual signal energy of each of the M subbands, corrects the first parameter, and determines the first parameter obtained by the correction as a final first parameter, where the first parameter indicates a value relationship between the downmix signal energy and the residual signal energy of each of the M subbands.
[0161] Specifically, when determining the first parameter, the encoder side determines M energy parameters based on the downmix signal energy and residual signal energy of each of the M subbands, where the M energy parameters respectively indicate a value relationship between the downmix signal energy and the residual signal energy of one of the M subbands, the M energy parameters correspond one-to-one to the M subbands, and the encoder side determines a sum of the M energy parameters as the first parameter.
[0162] Optionally, the energy parameter of the subband with subband index number b among the M energy parameters determined by the encoder side may be determined using function (1).
[0163] For example, the energy parameter of the subband with the subband index number b among the M energy parameters satisfies equation (2).
[0164] Optionally, the energy parameter of the subband with subband index number b among the M energy parameters determined by the encoder side may be determined using function (11).
[0165] For example, the energy parameter of the subband with the subband index number b among the M energy parameters satisfies equation (11).
[0166] For example, the first parameter res_dmx_ratio1 determined by the encoder side based on the M energy parameters satisfies the following equation:
number
[0167] In addition, the encoder side may further determine a maximum value res_dmx_ratio_max among the M energy parameters, where res_dmx_ratio_max satisfies equation (12).
[0168] The encoder side corrects res_dmx_ratio1 based on res_dmx_ratio_max of each of the M subbands and the downmix signal energy res_cod_NRG_M[b], and determines res_dmx_ratio2 obtained by the correction.
[0169] For example, the encoder side corrects res_dmx_ratio1 according to the following formula, where M=5: The res_dmx_ratio2 obtained by the correction satisfies the following formula.
number
[0170] Optionally, the res_dmx_ratio2 obtained by the correction may be further corrected.
[0171] For example, the final res_dmx_ratio3 obtained by the correction satisfies the following formula: res_dmx_ratio3=pow(res_dmx_ratio2,1.2)(15) In the formula, the pow() function represents an exponential function, and pow(res_dmx_ratio2,1.2) represents res_dmx_ratio2 raised to the power of 1.2.
[0172] After determining the first parameter obtained by correction (res_dmx_ratio3 obtained by correction), the encoder side may determine the second parameter according to the method described in Method 1, finally determine the residual signal coding parameter according to the method described in Method 1, and determine whether to code the residual signal of each of the M subbands based on the residual signal coding parameter.
[0173] Method 4
[0174] The method for determining residual signal coding parameters in Method 4 is similar to that of Method 1, and the difference lies in the different method for determining the first parameter, so reference may be made to the relevant description of determining residual signal coding parameters in Method 1. For brevity, this specification will only describe the method for determining the first parameter in Method 4.
[0175] By way of example and not limitation, the encoder side determines the first parameter based on a sum of the residual signal energies of the M subbands and a sum of the downmix signal energies of the M subbands.
[0176] Specifically, the encoder side separately determines the sum of downmix signal energies of the M subbands dmx_nrg_all_curr and the sum of residual signal energies of the M subbands res_nrg_all_curr, and determines the first parameter based on dmx_nrg_all_curr and res_nrg_all_curr.
[0177] Optionally, the sum of the downmix signal energies of the M subbands, dmx_nrg_all_curr, satisfies equation (4): 。
[0178] Optionally, the sum of the residual signal energies of the M subbands, res_nrg_all_curr, satisfies equation (5): 。
[0179] The encoder side determines the first parameter res_dmx_ratio based on dmx_nrg_all_curr and res_nrg_all_curr.
[0180] For example, the first parameter res_dmx_ratio finally determined by the encoder side satisfies the following formula: res_dmx_ratio=res_nrg_all_curr / dmx_nrg_all_curr(16)
[0181] After determining the first parameter, the encoder side may determine the second parameter according to the method described in Method 1, finally determine the residual signal coding parameter according to the method described in Method 1, and determine whether to code the residual signal of each of the M subbands based on the residual signal coding parameter.
[0182] In order to better understand the whole stereo signal encoding, hereinafter, a wideband stereo signal in which the signal length of each frame is 20 ms is used as an example, and a frame being processed by the encoder side (for example, the current frame) is used as an example, and the stereo signal encoding method 300 of this embodiment of the present application is described with reference to Fig. 7. The stereo signal encoding method shown in Fig. 7 includes at least the following steps:
[0183] 401: Perform time-domain pre-processing on the left channel time-domain signal and the right channel time-domain signal to obtain left channel time-domain signal and right channel time-domain signal obtained by time-domain pre-processing.
[0184] Specifically, the signal length of the current frame is 20 ms. The sampling frequency is 16 kHz ( kHz ), then after sampling, the frame length of the current frame H=320, in other words, the current frame contains 320 sampling points.
[0185] The stereo signal of the current frame includes a left channel time-domain signal of the current frame and a right channel time-domain signal of the current frame. The left channel time-domain signal of the current frame is x L (n), and the right channel time domain signal of the current frame is denoted by x R (n), where n is the sequence number of the sampling point, and n=0, 1, ..., and H-1. The left channel time-domain signal and the right channel time-domain signal may be referred to as the left and right channel time-domain signals.
[0186] The step of performing time-domain pre-processing on the left channel time-domain signal and the right channel time-domain signal of the current frame may include performing high-pass filtering on the left channel time-domain signal and the right channel time-domain signal of the current frame, respectively, to obtain the left channel time-domain signal and the right channel time-domain signal of the current frame obtained by time-domain pre-processing. The left channel time-domain signal of the current frame obtained by pre-processing is represented by x L_HP (n), and the right channel time domain signal of the current frame obtained by preprocessing is denoted as x R_HP The high-pass filtering process may be performed using an infinite impulse response (IIR) digital filter with a cutoff frequency of 20 Hz, or other types of filters.
[0187] For example, when the sampling rate of a stereo signal is 16 kHz, the corresponding transfer function of a high-pass filter with a cutoff frequency of 20 Hz may be:
number
[0188] where b0=0.994461788958195, b1=-1.988923577916390, b2=0.994461788958195, a1=1.98892905899653, a2=-0.988954249933127, and z represents the transform coefficients of the Z transform. The corresponding time domain filter is: x L_HP (n)=b0·x L (n)+b1·x L (n-1)+b2·x L (n-2)-a1·x L_HP (n-1)-a2·x L_HP (n-2)(18)
[0189] 402: Perform time-domain analysis on the left channel time-domain signal and the right channel time-domain signal obtained by time-domain pre-processing.
[0190] Specifically, the time-domain analysis may include transient detection, etc. The transient detection may be performing energy detection separately on the left channel time-domain signal and the right channel time-domain signal of the current frame obtained by pre-processing, to detect whether an energy burst occurs in the current frame.
[0191] For example, the energy E of the left channel time domain signal of the current frame obtained by preprocessing cur_L In order to obtain a transient detection result for the left channel time domain signal of the current frame obtained by preprocessing, the energy E pre_Land the energy E of the left channel time domain signal of the current frame obtained by preprocessing. cur_L The transient detection can be performed on the right channel time domain signal of the current frame obtained by pre-processing using the same method.
[0192] The time domain analysis may include other time domain analyses in addition to transient detection, such as time domain Inter-channel Time Difference (ITD) parameter determination, time domain delay matching processing, and bandwidth extension pre-processing.
[0193] 403: Perform time-frequency transformation on the left channel time-domain signal and the right channel time-domain signal obtained by time-domain pre-processing to obtain a left channel frequency-domain signal and a right channel frequency-domain signal.
[0194] Specifically, a discrete Fourier transform may be performed on the left channel time domain signal obtained by time domain pre-processing to obtain a left channel frequency domain signal, and a discrete Fourier transform is performed on the right channel time domain signal obtained by time domain pre-processing to obtain a right channel frequency domain signal.
[0195] To overcome the problem of spectral aliasing, the overlap-add method may be used to process between two consecutive times of the Discrete Fourier Transform, and in some cases, zeros may be added to the input signal of the Discrete Fourier Transform.
[0196] The discrete Fourier transform may be performed once per frame, or each frame of the signal may be divided into P subframes (P is a positive integer greater than or equal to 2) and the discrete Fourier transform may be performed once per subframe.
[0197] For example, a discrete Fourier transform is performed once on the current frame, and the left channel frequency domain signal of the current frame on which the discrete Fourier transform is performed is denoted as L(k), and the right channel frequency domain signal of the current frame on which the discrete Fourier transform is performed is denoted as R(k), where k represents a frequency bin index number, k=0, 1, ..., L-1, and L represents the frame length of the current frame on which the discrete Fourier transform is performed. In other words, the current frame on which the discrete Fourier transform is performed includes L frequency bins.
[0198] In other examples , current The current frame is divided into P subframes, where P is a positive integer equal to or greater than 2. The left channel frequency domain signal of the subframe with index number i on which the discrete Fourier transform is performed is denoted by L. i The right channel frequency domain signal of the subframe where the discrete Fourier transform is performed is denoted as (k) and the index number is i. i (k), where i represents a subframe index number, i = 0, 1, ..., P-1, k represents a frequency bin index number, k = 0, 1, ..., L-1, and L represents the frame length of each subframe on which the discrete Fourier transform is performed. In other words, each subframe on which the discrete Fourier transform is performed includes L frequency bins.
[0199] 404: Determine ITD parameters and encode the determined ITD parameters.
[0200] Specifically, there are multiple methods for determining the ITD parameters. The ITD parameters may be determined only in the frequency domain, or only in the time domain, or in the time-frequency domain. This is not a limitation in this application.
[0201] The ITD parameters can be extracted in the time domain using cross-correlation coefficients. For example, 0≦i≦T max In the range of
number
number
[0202]
number
number
[0203] Alternatively, the ITD parameters may be determined in the frequency domain based on the left and right channel frequency domain signals. For example, the time domain signals may be converted to frequency domain signals using a time-frequency transform technique such as a Discrete Fourier Transform (DFT), a Fast Fourier Transform (FFT), and a Modified Discrete Cosine Transform (MDCT).
[0204] In this embodiment of the present application, the left channel frequency domain signal of the subframe whose index number is i and for which the Discrete Fourier Transform is performed is L i (k), where k=0, 1, ..., L / 2-1, and the right channel frequency domain signal of the subframe where the index number is i and the transformation is performed is R i(k), where k = 0, 1, ..., L / 2-1 and i = 0, 1, ..., P-1. The frequency domain of the subframe with index number i. mutual The correlation coefficient is XCORR i (k)=L i (k)·R * i (k) is calculated according to R * i (k) represents the conjugate of the right channel frequency domain signal of the i-th subframe on which the transformation is performed.
[0205] The frequency domain cross-correlation coefficient is the time domain xcorr i (n), where n=0, 1, …, L-1, and the ITD parameter value of the subframe with index number i is
number
[0206] In addition, a search range -T is calculated based on the left channel frequency domain signal and the right channel frequency domain signal of the subframe whose index number is i and for which the DFT transformation is performed. max ≦j≦T max In
number
number
[0207] After the ITD parameters are determined, the ITD parameters may be coded to obtain coding parameters, which are written into the stereo coded bitstream.
[0208] 405: Perform a time shift adjustment on the left frequency domain signal and the right channel frequency domain signal based on the ITD parameters.
[0209] Specifically, the time-shift adjustment may be performed on the left channel frequency domain signal and the right channel frequency domain signal according to any technique, which is not limited in this embodiment of the present application.
[0210] For example, if the current frame of a signal is divided into P subframes, where P is a positive integer equal to or greater than 2, the left channel frequency domain signal obtained by time shift adjustment of the subframe with index number i is L' i (k), where k=0, 1, ..., L / 2-1, and the right channel frequency domain signal obtained by time shift adjustment of the subframe with index number i is R' i It may be written as (k), where k represents the frequency bin index number, k=0, 1, ..., L / 2-1, and i represents the subframe index number, i=0, 1, ..., P-1.
number
[0211] T i represents the ITD parameter value of the subframe with index number i, L represents the length of the subframe over which the Discrete Fourier Transform is performed, and L i (k) represents the left channel frequency domain signal of the i-th subframe where the index number is i and the transformation is performed, and R i (k) represents the right channel frequency domain signal of the subframe with index number i on which the transformation is performed, where i represents the subframe index number, i=0, 1, . . . , P-1.
[0212] 406: Calculate other frequency domain stereo parameters based on the left channel frequency domain signal and the right channel frequency domain signal obtained by the time shift adjustment, and encode the other frequency domain stereo parameters.
[0213] Specifically, other frequency domain stereo parameters may include, but are not limited to, an inter-channel phase difference (IPD) parameter, an inter-channel level difference (ILD) parameter, and / or subband side gains, etc. ILD may also be referred to as inter-channel amplitude difference.
[0214] After the stereo parameters of the other frequency domains are obtained by calculation, the stereo parameters of the other frequency domains may be coded to obtain coding parameters, which are written into a stereo coded bitstream.
[0215] 407: Determine M subbands that satisfy a preset condition from the N subbands contained in the frequency domain signal of the current frame.
[0216] Specifically, the frequency domain signal of the current frame obtained by time shift adjustment is divided into subbands. For example, the frequency domain signal of the current frame is divided into N subbands (N is a positive integer greater than or equal to 2), and the frequency bins included in the subband with subband index number b are k∈[band_limits(b),band_limits(b+1)-1], where band_limits(b) represents the minimum index number of the frequency bins included in the subband with subband index number b, and band_limits(b+1) represents the minimum index number of the frequency bins included in the subband with subband index number b+1. According to the preset condition, M subbands that satisfy the preset condition are determined from the N subbands.
[0217] For example, the preset condition may be that the subband index number is less than or equal to a preset maximum subband index number, i.e., b≦res_cod_band_max, where res_cod_band_max represents the preset maximum subband index number.
[0218] Alternatively, the preset condition may be that the subband index number is less than or equal to a preset maximum subband index number and greater than or equal to a preset minimum subband index number, i.e., res_cod_band_min≦b≦res_cod_band_max, where res_cod_band_max represents the preset maximum subband index number and res_cod_band_min represents the preset minimum subband index number.
[0219] Furthermore, for wideband stereo signals, different preset conditions can be set based on different coding rates. For example, when the coding rate is 26 kbps, the preset condition is that the subband index number b is 5 or less, that is, the maximum preset subband index number is 5. When the coding rate is 44 kbps, the preset condition is that the subband index number b is 6 or less, that is, the maximum preset subband index number is 6. When the coding rate is 56 kbps, the preset condition is that the subband index number b is 7 or less, that is, the maximum preset subband index number is 7.
[0220] For example, if the preset condition is that the subband index number b≦4, then five subbands with index numbers 0 to 4 may be determined as subbands that satisfy the preset condition from among the N subbands of the current frame.
[0221] In addition, if the current frame of the signal is divided into P subframes (P is a positive integer equal to or greater than 2), each subframe obtained by time shift adjustment is divided into subbands. For example, if a subframe with index number i (i=0, 1, …, P-1) is divided into N subbands, and the subband with index number b in the subframe with index number i contains k frequency bins. i ∈[band_limits(b),band_limits(b+1)-1], where band_limits(b) represents the minimum index number of the frequency bin included in the subband with index number b in the subframe with index number i, and band_limits(b+1) represents the minimum index number of the frequency bin included in the subband with index number b+1 in the subframe with index number i.
[0222] According to a preset condition, M subbands that satisfy the preset condition are determined from among N subbands included in each frame.
[0223] The preset condition may be that the subband index number is equal to or greater than a preset minimum subband index number and equal to or less than a preset maximum subband index number, ie, res_cod_band_min≦b≦res_cod_band_max.
[0224] For example, if the preset condition is 4≦b≦8, five subbands with index numbers 4 to 8 are determined as subbands that satisfy the preset condition from among the N subbands in each subframe.
[0225] 408: Calculate a downmix signal and a residual signal of the subbands that satisfy a preset condition according to the left channel frequency domain signal and the right channel frequency domain signal obtained by the time shift adjustment.
[0226] Specifically, the method for calculating the downmix signal and the residual signal of the subbands that satisfy the preset conditions will be described using an example in which the current frame is divided into P subframes (P is a positive integer greater than or equal to 2) (e.g., the current frame can be divided into two subframes or four subframes).
[0227] For example, if the preset condition is that the subband index number b is less than or equal to 5, the downmix signals and residual signals of the subbands with index numbers from 0 to 5 in each subframe are calculated.
[0228] The downmix signal of the subband with index number b (b≦5) in the subframe with index number i is DMX i The residual signal of the subband with index number b in the subframe with index number i is denoted as RES i It is written as '(k)' and is DMX i (k) and RES i '(k) satisfies the following equation:
number
number
number
number
[0229] IPD i(b) represents the IPD parameters of the subband with index number b in the subframe with index number i, and g_ILD i represents the side gain of the subband with index number b in the subframe with index number i, and L' i (k) represents the left channel frequency domain signal obtained by time shift adjustment of the subband with index number b in the subframe with index number i, and R' i (k) is ,stomach In the subframe with index number i The index number of is b represents the right channel frequency domain signal obtained by time-shift adjustment of the subband, L'' i (k) represents the left channel frequency domain signal obtained by adjusting the stereo parameters of the subband with index number b in the subframe with index number i, and R'' i (k) represents a right channel frequency domain signal obtained by adjusting multiple stereo parameters of a subband with index number b in a subframe with index number i, where i represents a subframe index number, i = 0, 1, ..., P-1, k represents a frequency bin index number, k ∈ [band_limits(b),band_limits(b+1)-1], band_limits(b) represents the minimum index number of a frequency bin included in a subband with index number b in a subframe with index number i, and band_limits(b+1) represents the minimum index number of a frequency bin included in a subband with index number b+1 in a subframe with index number i.
[0230] In another example, the downmix signal DMX of the subband with index number b in the subframe with index number i is i (k) may alternatively be calculated according to the following method. DMX i (k) = [L''(k) + R''(k)] · c(26), and
number
[0231] L'' i (k) represents the left channel frequency domain signal obtained by adjusting the stereo parameters of the subband with index number b in the subframe with index number i, and R'' i (k) represents a right channel frequency domain signal obtained by adjusting a plurality of stereo parameters of a subband with index number b in a subframe with index number i, where i represents a subframe index number, i = 0, 1, ..., P-1, k represents a frequency bin index number, k ∈ [band_limits(b), band_limits(b+1)-1], band_limits(b) represents the minimum index number of a frequency bin included in the subband with subband index number b, and band_limits(b+1) represents the minimum index number of a frequency bin included in the subband with index number b+1 in the subframe with index number i. The methods for calculating the downmix signal energy and the residual signal energy are not limited in this embodiment of the present application.
[0232] 409: Determine residual signal coding parameters based on the downmix signal energy and the residual signal energy of the subbands that satisfy the preset condition.
[0233] 410: Based on the residual signal coding parameters, determine whether the residual signal of each of the M subbands of the current frame needs to be coded. If it is determined that the residual signal needs to be coded, perform 412. If it is determined that the residual signal does not need to be coded, perform 411.
[0234] 411: Encode the downmix signal of each of the M subbands of the current frame based on the residual signal encoding parameters, where the residual signal does not need to be encoded.
[0235] 412: Encode the downmix signal and the residual signal of each of the M subbands of the current frame based on the residual signal coding parameters.
[0236] For specific implementations of steps 409 to 411, please refer to the related description of method 300. For the sake of brevity, the details will not be repeated here.
[0237] method 300 , where the encoder side divides the current frame into P subframes, where P is a positive integer greater than or equal to 2, and divides the spectral coefficients of each of the P subframes into N subbands, and the residual signal coding parameters are determined based on the downmix signal energies and residual signal energies of M subbands (the M subbands are at least a part of the N subbands) in each subframe that satisfy the preset condition. Therefore, it should be noted that in method 300, the residual signal energy res_cod_NRG_S[b] of the subband with index number b in the current frame is the sum of the residual signal energies of the subband with index number b in all P subframes, and the downmix signal energy res_cod_NRG_M[b] of the subband with index number b in the current frame is the sum of the downmix signal energies of the subband with index number b in all P subframes.
[0238] For example, the current frame is divided into two subframes, and the spectral coefficients of each of the two subframes are divided into N subbands. Thus, in method 300, the downmix signal energy res_cod_NRG_M[b] of the subband with index number b in the current frame is the sum of the downmix signal energy of the subband with index number b in subframe 1 and the downmix signal energy of the subband with index number b in subframe 2, and the residual signal energy res_cod_NRG_S[b] of the subband with index number b in the current frame is the sum of the residual signal energy of the subband with index number b in subframe 1 and the residual signal energy of the subband with index number b in subframe 2.
[0239] The stereo signal encoding method according to the embodiment of the present application has been described in detail above with reference to Figures 1 to 7. Hereinafter, the stereo signal encoding device according to the embodiment of the present application will be described with reference to Figures 8 and 9. It should be understood that both the devices in Figures 8 and 9 correspond to the stereo signal encoding method according to the embodiment of the present application. In addition, both the devices in Figures 8 and 9 can perform the stereo signal encoding method according to the embodiment of the present application. For the sake of brevity, repeated description will be omitted below as appropriate.
[0240] 8 is a schematic block diagram of a stereo signal encoding device according to an embodiment of the present application. The device 500 of FIG. a first determining module 501 configured to determine a residual signal coding parameter of a current frame of the stereo signal based on a downmix signal energy and a residual signal energy of each of M subbands of the current frame, where the residual signal coding parameter of the current frame is used to indicate whether to code a residual signal of the M subbands, the M subbands being at least a part of the N subbands, N being a positive integer greater than 1, and M≦N, where M is a positive integer; a second decision module 502 configured to decide whether to code the residual signal of the M subbands of the current frame based on the residual signal coding parameters of the current frame; Includes:
[0241] In this application, the residual signal coding parameters are determined based on the downmix signal energy and the residual signal energy of M subbands within the N subbands that satisfy a preset bandwidth range, and whether to encode the residual signal of each of the M subbands is determined based on the residual signal coding parameters. This avoids encoding only the downmix signal when the coding rate is relatively low. Alternatively, whether to encode all the residual signals of the subbands that satisfy the preset bandwidth range is determined based on the residual signal coding parameters. This improves the spatial sense and sound image stability of the decoded stereo signal while minimizing high-frequency distortion of the decoded stereo signal, thereby improving coding quality.
[0242] Optionally, in one embodiment, the M subbands are M subbands whose subband index numbers are less than or equal to a preset maximum subband index number in the N subbands.
[0243] Optionally, in one embodiment, the M subbands are M subbands within the N subbands whose subband index numbers are greater than or equal to a preset minimum subband index number and less than or equal to a preset maximum subband index number.
[0244] Optionally, in one embodiment, the second decision module 502 is further configured to compare the residual signal coding parameter with a preset first threshold, and determine not to code the residual signal of each of the M subbands if the first threshold is greater than 0 and less than 1.0 and the residual signal coding parameter is less than or equal to the first threshold, or determine to code the residual signal of each of the M subbands if the residual signal coding parameter is greater than the first threshold.
[0245] Optionally, in one implementation, the first determining module 501 is further configured to determine a residual signal coding parameter based on the downmix signal energy, the residual signal energy and the side gain of each of the M subbands.
[0246] Optionally, in one embodiment, the first determination module 501 is further configured to: determine a first parameter based on the downmix signal energy, the residual signal energy, and the side gain of each of the M subbands, where the first parameter indicates a value relationship between the downmix signal energy and the residual signal energy of each of the M subbands; determine a second parameter based on the downmix signal energy and the residual signal energy of each of the M subbands, where the second parameter indicates a value relationship between a first energy sum and a second energy sum, where the first energy sum is the sum of the residual signal energy and the downmix signal energy of the M subbands and the second energy sum is the sum of the residual signal energy and the downmix signal energy of the M subbands in the frequency domain signal of a frame previous to the current frame, where the M subbands of the current frame have the same subband index numbers as the M subbands of the previous frame; and determine a residual signal coding parameter based on the first parameter, the second parameter, and a long-term smoothing parameter of the frame previous to the current frame.
[0247] Optionally, in one embodiment, the first determination module 501 is further configured to determine M energy parameters based on the downmix signal energy, residual signal energy, and side gain of each of the M subbands, wherein the M energy parameters each indicate a value relationship between the downmix signal energy and the residual signal energy of one of the M subbands, the M energy parameters having a one-to-one correspondence with the M subbands, and determine an energy parameter having a maximum value among the M energy parameters as the first parameter.
[0248] Optionally, in one embodiment, an energy parameter of a subband with a subband index number b among the M energy parameters determined by the first determining module 501 satisfies the following formula: res_dmx_ratio[b]=res_cod_NRG_S[b] / (res_cod_NRG_S[b]+(1-g(b))·(1-g(b))·res_cod_NRG_M[b]+1) In the formula, res_dmx_ratio[b] represents the energy parameter of the subband having subband index number b, where b is greater than or equal to 0 and less than or equal to the preset maximum subband index number, res_cod_NRG_S[b] represents the residual signal energy of the subband having subband index number b, res_cod_NRG_M[b] represents the downmix signal energy of the subband having subband index number b, and g(b) represents a function of the side gain side_gain[b] of the subband having subband index number b.
[0249] Optionally, in one embodiment, the first determination module 501 is further configured to: determine a first parameter based on the downmix signal energy and the residual signal energy of each of the M subbands, the first parameter indicating a value relationship between the downmix signal energy and the residual signal energy of each of the M subbands; determine a second parameter based on the downmix signal energy and the residual signal energy of each of the M subbands, the second parameter indicating a value relationship between a first energy sum and a second energy sum, the first energy sum being the sum of the residual signal energy and the downmix signal energy of the M subbands, and the second energy sum being the sum of the residual signal energy and the downmix signal energy of the M subbands in the frequency domain signal of a frame previous to the current frame, the M subbands of the current frame having the same subband index numbers as the M subbands of the previous frame; and determine a residual signal coding parameter based on the first parameter, the second parameter, and a long-term smoothing parameter of the frame previous to the current frame.
[0250] Optionally, in one embodiment, the first determination module 501 is further configured to determine M energy parameters based on the downmix signal energy and the residual signal energy of each of the M subbands, the M energy parameters respectively indicating a value relationship between the downmix signal energy and the residual signal energy of each of the M subbands, the M energy parameters having a one-to-one correspondence with the M subbands, and determine an energy parameter having a maximum value among the M energy parameters as the first parameter.
[0251] Optionally, in one embodiment, an energy parameter of a subband with a subband index number b among the M energy parameters determined by the first determining module 501 satisfies the following formula: res_dmx_ratio[b]=res_cod_NRG_S[b] / res_cod_NRG_M[b] In the formula, res_dmx_ratio[b] represents the energy parameter of the subband having subband index number b, where b is greater than or equal to 0 and less than or equal to the preset maximum subband index number, res_cod_NRG_S[b] represents the residual signal energy of the subband having subband index number b, and res_cod_NRG_M[b] represents the downmix signal energy of the subband having subband index number b.
[0252] Optionally, in one embodiment, the first determination module 501 is further configured to determine the sum of M energy parameters as a first parameter (to be corrected) res_dmx_ratio1, correct res_dmx_ratio1 based on the maximum value res_dmx_ratio_max among the M energy parameters and the downmix signal energy res_cod_NRG_M[b] of each of the M subbands, and determine res_dmx_ratio2 obtained by the correction.
[0253] For example, the encoder side corrects res_dmx_ratio1 according to the following formula, where M=5: The res_dmx_ratio2 obtained by the correction satisfies the following formula.
number
[0254] Optionally, in one embodiment, the res_dmx_ratio2 obtained by the correction may be further corrected.
[0255] For example, the final res_dmx_ratio3 obtained by the correction satisfies the following formula: res_dmx_ratio3=pow(res_dmx_ratio2,1.2) In the formula, the pow() function represents an exponential function, and pow(res_dmx_ratio2,1.2) represents res_dmx_ratio2 raised to the power of 1.2.
[0256] Optionally, in one implementation, the first determining module 501 is further configured to determine the first parameter based on a sum of residual signal energies of the M subbands and a sum of downmix signal energies of the M subbands.
[0257] Specifically, the encoder side separately determines the sum of downmix signal energies of the M subbands dmx_nrg_all_curr and the sum of residual signal energies of the M subbands res_nrg_all_curr, and determines the first parameter based on dmx_nrg_all_curr and res_nrg_all_curr.
[0258] Optionally, in one embodiment, the sum of the downmix signal energies of the M subbands dmx_nrg_all_curr satisfies the following equation:
number
[0259] Optionally, in one embodiment, the sum of the residual signal energies of the M subbands, res_nrg_all_curr, satisfies the following equation:
number
[0260] The encoder side determines the first parameter res_dmx_ratio based on dmx_nrg_all_curr and res_nrg_all_curr.
[0261] For example, the first parameter res_dmx_ratio finally determined by the encoder side satisfies the following formula: res_dmx_ratio=res_nrg_all_curr / dmx_nrg_all_curr
[0262] Optionally, in one embodiment, an energy parameter of a subband with a subband index number b among the M energy parameters determined by the first determining module 501 satisfies the following formula: res_dmx_ratio[b]=res_cod_NRG_S[b] / res_cod_NRG_M[b] In the formula, res_dmx_ratio[b] represents the energy parameter of the subband having subband index number b among the M energy parameters, where b is greater than or equal to 0 and less than or equal to the preset maximum subband index number, res_cod_NRG_S[b] represents the residual signal energy of the subband having subband index number b, and res_cod_NRG_M[b] represents the downmix signal energy of the subband having subband index number b.
[0263] Optionally, in one embodiment, the residual signal coding parameter of the current frame determined by the first determination module 501 is a long-term smoothing parameter of the current frame, and the long-term smoothing parameter of the current frame satisfies the following equation: res_dmx_ratio_lt=res_dmx_ratio·α+res_dmx_ratio_lt_prev·(1-α) where res_dmx_ratio_lt represents the long-term smoothing parameter of the current frame, res_dmx_ratio represents the first parameter, and res_dmx_ratio_lt_prev represents the long-term smoothing parameter of the frame previous to the current frame, and 0<α<1; When the second parameter is greater than a third preset threshold, the value of α when the first parameter is less than the second preset threshold is greater than the value of α when the first parameter is equal to or greater than the second preset threshold, the second threshold is greater than or equal to 0 and less than or equal to 0.6, and the third threshold is greater than or equal to 2.7 and less than or equal to 3.7; or When the second parameter is greater than a fifth preset threshold, the value of α when the first parameter is greater than a fourth preset threshold is greater than the value of α when the first parameter is equal to or less than the fourth preset threshold, the fourth threshold is greater than or equal to 0.9, and the fifth threshold is greater than or equal to 0.71; or The value of α when the first parameter is smaller than a preset second threshold and the second parameter is greater than a preset third threshold is greater than the value of α when the second parameter is greater than or equal to a preset fifth threshold and less than or equal to the preset third threshold, the second threshold being greater than or equal to 0 and less than or equal to 0.6, the third threshold being greater than or equal to 2.7 and less than or equal to 3.7, and the fifth threshold being greater than or equal to 0 and less than or equal to 0.71.
[0264] Optionally, in one embodiment, the second decision module 502 is further configured to: encode the downmix signal and the residual signal of the M subbands when it is determined to encode the residual signals of the M subbands, or encode the downmix signal of the M subbands when it is determined not to encode the residual signals of the M subbands.
[0265] 9 is a schematic block diagram of a stereo signal encoding device according to an embodiment of the present application. The device 600 of FIG. a memory 601 configured to store a program; a processor 602 configured to execute a program stored in a memory 601, the processor 602 being particularly configured, when the program in the memory is executed, to determine a residual signal coding parameter for a current frame of the stereo signal based on a downmix signal energy and a residual signal energy of each of M subbands of the current frame, the residual signal coding parameter for the current frame being used to indicate whether to code the residual signal of the M subbands, the M subbands being at least a portion of the N subbands, N being a positive integer greater than 1, and M≦N, N being a positive integer; and Includes:
[0266] Optionally, in one embodiment, the M subbands are M subbands whose subband index numbers are less than or equal to a preset maximum subband index number in the N subbands.
[0267] Optionally, in one embodiment, the M subbands are M subbands within the N subbands whose subband index numbers are greater than or equal to a preset minimum subband index number and less than or equal to a preset maximum subband index number.
[0268] In one optional implementation, the processor 602 compares the residual signal coding parameter to a preset first threshold, and determines if the first threshold is greater than 0 and less than 1.0 and the residual signal coding parameter is greater than the first threshold. is If the residual signal coding parameter is greater than the first threshold, determine not to code the residual signal of each of the M subbands, or determine to code the residual signal of each of the M subbands if the residual signal coding parameter is greater than the first threshold.
[0269] In an optional implementation, the processor 602 is further configured to determine a residual signal coding parameter based on the downmix signal energy, the residual signal energy, and the side gain for each of the M subbands.
[0270] In an optional implementation, the processor 602 is further configured to: determine a first parameter based on the downmix signal energy, the residual signal energy, and the side gain of each of the M subbands, wherein the first parameter indicates a value relationship between the downmix signal energy and the residual signal energy of each of the M subbands; determine a second parameter based on the downmix signal energy and the residual signal energy of each of the M subbands, wherein the second parameter indicates a value relationship between a first energy sum and a second energy sum, wherein the first energy sum is the sum of the residual signal energy and the downmix signal energy of the M subbands and the second energy sum is the sum of the residual signal energy and the downmix signal energy of the M subbands in the frequency-domain signal of a frame previous to the current frame, wherein the M subbands of the current frame have the same subband index numbers as the M subbands of the previous frame; and determine a residual signal coding parameter based on the first parameter, the second parameter, and a long-term smoothing parameter of the frame previous to the current frame.
[0271] In one optional implementation, the processor 602 is further configured to determine M energy parameters based on the downmix signal energy, the residual signal energy, and the side gain of each of the M subbands, wherein the M energy parameters each indicate a value relationship between the downmix signal energy and the residual signal energy of one of the M subbands, the M energy parameters having a one-to-one correspondence with the M subbands, and determine an energy parameter having a maximum value among the M energy parameters as a first parameter.
[0272] Optionally, in one embodiment, an energy parameter of a subband having a subband index number b among the M energy parameters determined by the processor 602 satisfies the following equation: res_dmx_ratio[b]=res_cod_NRG_S[b] / (res_cod_NRG_S[b]+(1-g(b))·(1-g(b))·res_cod_NRG_M[b]+1) In the formula, res_dmx_ratio[b] represents the energy parameter of the subband having subband index number b, where b is greater than or equal to 0 and less than or equal to the preset maximum subband index number, res_cod_NRG_S[b] represents the residual signal energy of the subband having subband index number b, res_cod_NRG_M[b] represents the downmix signal energy of the subband having subband index number b, and g(b) represents a function of the side gain side_gain[b] of the subband having subband index number b.
[0273] In an optional implementation, the processor 602 is further configured to: determine a first parameter based on the downmix signal energy and the residual signal energy of each of the M subbands, the first parameter indicating a value relationship between the downmix signal energy and the residual signal energy of each of the M subbands; determine a second parameter based on the downmix signal energy and the residual signal energy of each of the M subbands, the second parameter indicating a value relationship between a first energy sum and a second energy sum, the first energy sum being the sum of the residual signal energy and the downmix signal energy of the M subbands and the second energy sum being the sum of the residual signal energy and the downmix signal energy of the M subbands in the frequency-domain signal of a frame previous to the current frame, the M subbands of the current frame having the same subband index numbers as the M subbands of the previous frame; and determine a residual signal coding parameter based on the first parameter, the second parameter, and a long-term smoothing parameter of the frame previous to the current frame.
[0274] In one optional implementation, the processor 602 is further configured to determine M energy parameters based on the downmix signal energy and the residual signal energy of each of the M subbands, the M energy parameters respectively indicating a value relationship between the downmix signal energy and the residual signal energy of each of the M subbands, the M energy parameters having a one-to-one correspondence with the M subbands, and to determine an energy parameter having a maximum value among the M energy parameters as a first parameter.
[0275] Optionally, in one embodiment, an energy parameter of a subband having a subband index number b among the M energy parameters determined by the processor 602 satisfies the following equation: res_dmx_ratio[b]=res_cod_NRG_S[b] / res_cod_NRG_M[b] In the formula, res_dmx_ratio[b] represents the energy parameter of the subband having subband index number b, where b is greater than or equal to 0 and less than or equal to the preset maximum subband index number, res_cod_NRG_S[b] represents the residual signal energy of the subband having subband index number b, and res_cod_NRG_M[b] represents the downmix signal energy of the subband having subband index number b.
[0276] In one optional implementation, the processor 602 is further configured to determine the sum of M energy parameters as a first parameter (to be corrected) res_dmx_ratio1, correct res_dmx_ratio1 based on the maximum value res_dmx_ratio_max among the M energy parameters and the downmix signal energy res_cod_NRG_M[b] of each of the M subbands, and determine res_dmx_ratio2 obtained by the correction.
[0277] For example, the encoder side corrects res_dmx_ratio1 according to the following formula, where M=5: The res_dmx_ratio2 obtained by the correction satisfies the following formula.
number
[0278] Optionally, in one embodiment, the res_dmx_ratio2 obtained by the correction may be further corrected.
[0279] For example, the final res_dmx_ratio3 obtained by the correction satisfies the following formula: res_dmx_ratio3=pow(res_dmx_ratio2,1.2) In the formula, the pow() function represents an exponential function, and pow(res_dmx_ratio2,1.2) represents res_dmx_ratio2 raised to the power of 1.2.
[0280] Optionally, in an implementation, the processor 602 is further configured to determine the first parameter based on a sum of the residual signal energies of the M subbands and a sum of the downmix signal energies of the M subbands.
[0281] Specifically, the encoder side separately determines the sum of downmix signal energies of the M subbands dmx_nrg_all_curr and the sum of residual signal energies of the M subbands res_nrg_all_curr, and determines the first parameter based on dmx_nrg_all_curr and res_nrg_all_curr.
[0282] Optionally, in one embodiment, the sum of the downmix signal energies of the M subbands dmx_nrg_all_curr satisfies the following equation:
number
[0283] Optionally, in one embodiment, the sum of the residual signal energies of the M subbands, res_nrg_all_curr, satisfies the following equation:
number
[0284] The encoder side determines the first parameter res_dmx_ratio based on dmx_nrg_all_curr and res_nrg_all_curr.
[0285] For example, the first parameter res_dmx_ratio finally determined by the encoder side satisfies the following formula: res_dmx_ratio=res_nrg_all_curr / dmx_nrg_all_curr
[0286] Optionally, in one embodiment, an energy parameter of a subband having a subband index number b among the M energy parameters determined by the processor 602 satisfies the following equation: res_dmx_ratio[b]=res_cod_NRG_S[b] / res_cod_NRG_M[b] In the formula, res_dmx_ratio[b] represents the energy parameter of the subband having subband index number b among the M energy parameters, where b is greater than or equal to 0 and less than or equal to the preset maximum subband index number, res_cod_NRG_S[b] represents the residual signal energy of the subband having subband index number b, and res_cod_NRG_M[b] represents the downmix signal energy of the subband having subband index number b.
[0287] Optionally, in one embodiment is the If the first parameter is smaller than the preset second threshold and the second parameter is greater than the preset third threshold, the residual signal coding parameter of the current frame determined by the processor 602 is a long-term smoothing parameter of the current frame, and the long-term smoothing parameter of the current frame satisfies the following formula: res_dmx_ratio_lt=res_dmx_ratio·α+res_dmx_ratio_lt_prev·(1-α) where res_dmx_ratio_lt represents the long-term smoothing parameter of the current frame, res_dmx_ratio represents the first parameter, and res_dmx_ratio_lt_prev represents the long-term smoothing parameter of the frame previous to the current frame, and 0<α<1; When the second parameter is greater than a third preset threshold, the value of α when the first parameter is less than the second preset threshold is greater than the value of α when the first parameter is equal to or greater than the second preset threshold, the second threshold is greater than or equal to 0 and less than or equal to 0.6, and the third threshold is greater than or equal to 2.7 and less than or equal to 3.7; or When the second parameter is greater than a fifth preset threshold, the value of α when the first parameter is greater than a fourth preset threshold is greater than the value of α when the first parameter is equal to or less than the fourth preset threshold, the fourth threshold is greater than or equal to 0.9, and the fifth threshold is greater than or equal to 0.71; or The value of α when the first parameter is smaller than a preset second threshold and the second parameter is greater than a preset third threshold is greater than the value of α when the second parameter is greater than or equal to a preset fifth threshold and less than or equal to the preset third threshold, the second threshold being greater than or equal to 0 and less than or equal to 0.6, the third threshold being greater than or equal to 2.7 and less than or equal to 3.7, and the fifth threshold being greater than or equal to 0 and less than or equal to 0.71.
[0288] Optionally, in one embodiment, the processor 602 is further configured to: encode the downmix signal and the residual signal of the M subbands when it is determined to encode the residual signals of the M subbands, or encode the downmix signal of the M subbands when it is determined not to encode the residual signals of the M subbands.
[0289] The present application further provides a chip, which includes a processor and a communication interface, the communication interface being configured to communicate with an external device, and the processor being configured to perform the stereo signal encoding method in an embodiment of the present application.
[0290] Optionally, in one embodiment, the chip may further include a memory for storing instructions, and the processor is configured to execute the instructions stored in the memory, and when the instructions are executed, the processor is configured to perform the stereo signal encoding method in the embodiment of the present application.
[0291] Optionally, in one embodiment, the chip is incorporated into a terminal device or a network device.
[0292] The present application provides a computer-readable storage medium, which stores program code to be executed by a device, the program code including instructions for performing a stereo signal encoding method in an embodiment of the present application.
[0293] It should be understood that the processor referred to in the embodiments of the present invention may be a Central Processing Unit (CPU), or may be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field Programmable Gate Arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor, or the processor may be any conventional processor, etc.
[0294] It will be understood that the memory referred to in the embodiments of the present invention may be volatile or nonvolatile, or may include both volatile and nonvolatile memory. Nonvolatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory may be random access memory (RAM) used as an external cache. By way of example and not limitation, many forms of RAM may be used, such as static random access memory (Static RAM, SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (Synchronous DRAM, SDRAM), double data rate synchronous dynamic random access memory (Double Data Rate SDRAM, DDR SDRAM), enhanced synchronous dynamic random access memory (Enhanced SDRAM, ESDRAM), Synchlink dynamic random access memory (Synchlink DRAM, SLDRAM), and direct Rambus random access memory (Direct Rambus RAM, DR RAM).
[0295] It should be noted that if the processor is a general-purpose processor, DSP, ASIC, FPGA, or other programmable logic device, discrete gate, transistor logic circuit, or discrete hardware component, the memory (storage module) is integrated into the processor.
[0296] Note that memory as described herein may include, but is not limited to, these and any other suitable types of memory.
[0297] In combination with the examples described in the embodiments disclosed herein, those skilled in the art will understand that each unit and algorithm step can be realized by electronic hardware or a combination of computer software and electronic hardware. Whether a function is performed by hardware or software depends on the individual application and design constraints of the technical solution. Those skilled in the art may use various methods to implement the described functions for each specific application, but the implementation should not be considered as going beyond the scope of this application.
[0298] For ease of description, the detailed operation processes of the aforementioned systems, devices and units shall refer to the corresponding processes in the aforementioned method embodiments, and will not be described in detail herein, as will be clearly understood by those skilled in the art.
[0299] In some embodiments provided in the present application, it should be understood that the disclosed system, device, and method may be realized in other ways. For example, the described device embodiment is merely an example. For example, the division into units is merely a logical division of function, and other divisions are possible in actual implementation. For example, multiple units or components may be combined or integrated into other systems, and some features may be ignored or not implemented. In addition, the illustrated or described mutual couplings or direct couplings or communication connections may be realized using some interfaces. Indirect couplings or communication connections between devices or units may be realized in electronic, mechanical, or other forms.
[0300] Units described as separate components may or may not be physically separated, and components illustrated as units may or may not be physical units, located in one place, or distributed over multiple network units. Some or all of the units may be selected based on actual requirements to achieve the objectives of the solutions of each embodiment.
[0301] In addition, the functional units in the embodiments of the present application may be integrated into one processing unit, or each of the units may exist physically independently, or two or more units may be integrated into one unit.
[0302] When each function is realized in the form of a software functional unit and sold or used as an independent product, the function may be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application may essentially be realized, or a part that contributes to the prior art may be realized, or a part of the technical solution may be realized in the form of a software product. A computer software product is stored in a storage medium and includes several instructions for instructing a computer device (which may be a personal computer, a server, a network device, etc.) to perform all or part of the steps of the method described in the embodiments of the present application. The aforementioned storage medium includes any medium that can store program code, such as a USB flash drive, a removable hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0303] The above description is merely a specific embodiment of the present application and is not intended to limit the scope of protection of the present application. Any modifications or replacements that can be easily conceived by those skilled in the art within the technical scope disclosed in the present application shall fall within the scope of protection of the present application. Therefore, the scope of protection of the present application shall be subject to the scope of protection of the claims. [Explanation of symbols]
[0304] 110 Coding Components 120 Decoding Components 130 Mobile Terminals 131 Collection Components 132 Channel Coding Components 140 Mobile Terminals 141 Audio Playback Components 142 Channel Decoding Components 150 network elements 151 Channel Decoding Components 152 Channel Coding Components 300 Stereo signal encoding method 500 devices 501 First Decision Module 502 Second Decision Module 600 equipment 601 memory 602 processor
Claims
1. 1. A method for encoding a stereo signal, comprising: determining a residual signal coding parameter of a current frame of a stereo signal based on a downmix signal energy and a residual signal energy of each of M subbands of the current frame, wherein the residual signal coding parameter of the current frame is used to indicate whether to code a residual signal of the M subbands, the M subbands being at least a portion of N subbands, N being a positive integer greater than 1, and M≦N, where M is a positive integer; determining whether to encode the residual signal of the M subbands of the current frame based on the residual signal coding parameters of the current frame; A stereo signal encoding method comprising:
2. determining whether to encode the residual signal of the M subbands based on the residual signal coding parameters of the current frame, comparing the residual signal coding parameters of the current frame with a preset first threshold, the first threshold being greater than 0 and less than 1.0; determining not to code the residual signal of the M subbands if the residual signal coding parameter of the current frame is less than or equal to the first threshold; or determining to code the residual signal of the M subbands if the residual signal coding parameter is greater than the first threshold; 2. The method of claim 1, comprising:
3. determining a residual signal coding parameter of the current frame based on the downmix signal energy and the residual signal energy of each of the M subbands, determining the residual signal coding parameters of the current frame based on the downmix signal energy, the residual signal energy, and a side gain for each of the M subbands; 3. The method of claim 1 or 2, comprising:
4. determining the residual signal coding parameters of the current frame based on the downmix signal energy, the residual signal energy, and a side gain for each of the M subbands, determining a first parameter based on the downmix signal energy, the residual signal energy, and the side gain for each of the M subbands, wherein the first parameter indicates a value relationship between the downmix signal energy and the residual signal energy for each of the M subbands; determining a second parameter based on the downmix signal energy and the residual signal energy of each of the M subbands, wherein the second parameter indicates a value relationship between a first energy sum and a second energy sum, the first energy sum being a sum of the residual signal energy and the downmix signal energy of the M subbands, the second energy sum being a sum of the residual signal energy and the downmix signal energy of the M subbands in a frequency domain signal of a frame previous to the current frame, and the M subbands of the current frame have the same subband index numbers as the M subbands of the previous frame; determining the residual signal coding parameters for the current frame based on the first parameters, the second parameters, and a long-term smoothing parameter of the previous frame of the current frame; 4. The method of claim 3, comprising:
5. determining a first parameter based on the downmix signal energy, the residual signal energy, and the side gain for each of the M subbands, determining M energy parameters based on the downmix signal energy, the residual signal energy, and the side gains for each of the M subbands, wherein the M energy parameters respectively indicate the value relationship between the downmix signal energy and the residual signal energy for each of the M subbands, and the M energy parameters correspond one-to-one to the M subbands; determining an energy parameter having a maximum value among the M energy parameters as the first parameter; 5. The method of claim 4, comprising:
6. Among the M energy parameters, the energy parameter of the subband having the subband index number b satisfies the following formula: res_dmx_ratio[b]=res_cod_NRG_S[b] / (res_cod_NRG_S[b]+(1-g(b))・(1-g(b))res_cod_NRG_M[b]+1) 6. The method of claim 5, wherein res_dmx_ratio[b] represents the energy parameter of the subband having subband index number b, where b is greater than or equal to 0 and less than or equal to a preset maximum subband index number, res_cod_NRG_S[b] represents the residual signal energy of the subband having subband index number b, res_cod_NRG_M[b] represents the downmix signal energy of the subband having subband index number b, and g(b) represents a function of side gain side_gain[b] of the subband having subband index number b.
7. determining a residual signal coding parameter of the current frame based on the downmix signal energy and the residual signal energy of each of the M subbands, determining a first parameter based on the downmix signal energy and the residual signal energy of each of the M subbands, wherein the first parameter indicates a value relationship between the downmix signal energy and the residual signal energy of each of the M subbands; determining a second parameter based on the downmix signal energy and the residual signal energy of each of the M subbands, wherein the second parameter indicates a value relationship between a first energy sum and a second energy sum, the first energy sum being a sum of the residual signal energy and the downmix signal energy of the M subbands, the second energy sum being a sum of the residual signal energy and the downmix signal energy of the M subbands in a frequency domain signal of a frame previous to the current frame, and the M subbands of the current frame have the same subband index numbers as the M subbands of the previous frame; determining the residual signal coding parameters for the current frame based on the first parameters, the second parameters, and a long-term smoothing parameter of the previous frame of the current frame; 3. The method of claim 1 or 2, comprising:
8. determining a first parameter based on the downmix signal energy and the residual signal energy of each of the M subbands, determining M energy parameters based on the downmix signal energy and the residual signal energy of each of the M subbands, wherein the M energy parameters respectively indicate the value relationship between the downmix signal energy and the residual signal energy of each of the M subbands, and the M energy parameters correspond one-to-one to the M subbands; determining an energy parameter having a maximum value among the M energy parameters as the first parameter; 8. The method of claim 7, comprising:
9. Among the M energy parameters, the energy parameter of the subband having the subband index number b satisfies the following formula: res_dmx_ratio[b]=res_cod_NRG_S[b] / res_cod_NRG_M[b] 9. The method of claim 8, wherein res_dmx_ratio[b] represents the energy parameter of the subband having subband index number b, where b is greater than or equal to 0 and less than or equal to a preset maximum subband index number, res_cod_NRG_S[b] represents the residual signal energy of the subband having subband index number b, and res_cod_NRG_M[b] represents the downmix signal energy of the subband having subband index number b.
10. The residual signal coding parameter of the current frame is a long-term smoothing parameter of the current frame, and the long-term smoothing parameter of the current frame satisfies the following equation: res_dmx_ratio_lt=res_dmx_ratio・α+res_dmx_ratio_lt_prev・(1−α) where res_dmx_ratio_lt represents the long-term smoothing parameter of the current frame, res_dmx_ratio represents the first parameter, and res_dmx_ratio_lt_prev represents the long-term smoothing parameter of the previous frame of the current frame, and 0<α<1; When the second parameter is greater than a preset third threshold, the value of α when the first parameter is less than a preset second threshold is greater than the value of α when the first parameter is equal to or greater than the preset second threshold, the second threshold is greater than or equal to 0.6 and the third threshold is greater than or equal to 2.7 and less than or equal to 3.7; or When the second parameter is greater than a fifth preset threshold, the value of α when the first parameter is greater than a fourth preset threshold is greater than the value of α when the first parameter is equal to or less than the fourth preset threshold, the fourth threshold is greater than or equal to 0.9, and the fifth threshold is greater than or equal to 0.71; or 10. The method of claim 4, wherein, when the second parameter is greater than or equal to a fifth threshold and less than or equal to a third threshold, the value of α is less than the value of α when the first parameter is less than the second threshold and the second parameter is greater than the third threshold, the second threshold being greater than or equal to 0.6, the third threshold being greater than or equal to 2.7 and less than or equal to 3.7, and the fifth threshold being greater than or equal to 0.
71.
11. when it is decided to encode the residual signal of the M subbands, encoding the downmix signal of the M subbands and the residual signal; or encoding a downmix signal of the M subbands when it is determined not to encode the residual signal of the M subbands.
11. The method of any one of claims 1 to 10, further comprising:
12. A stereo signal encoding apparatus comprising: a memory configured to store a program; a processor configured to execute the program stored in the memory, wherein when the program in the memory is executed, the processor is configured to determine a residual signal coding parameter of a current frame of a stereo signal based on a downmix signal energy and a residual signal energy of each of M subbands of the current frame, the residual signal coding parameter of the current frame being used to indicate whether to code a residual signal of the M subbands, the M subbands being at least a portion of N subbands, N being a positive integer greater than 1, and M≦N, M being a positive integer; and A stereo signal encoding device comprising:
13. The processor: comparing the residual signal coding parameter with a preset first threshold, the first threshold being greater than 0 and less than 1.0; determining not to code the residual signal of the M subbands if the residual signal coding parameter of the current frame is less than or equal to the first threshold; or determining to encode the residual signal of the M subbands if the residual signal coding parameter of the current frame is greater than the first threshold; The apparatus of claim 12 further configured to:
14. The processor: determining the residual signal coding parameters of the current frame based on the downmix signal energy, the residual signal energy, and a side gain for each of the M subbands; 14. The apparatus of claim 12 or 13, further configured to:
15. The processor: determining a first parameter based on the downmix signal energy, the residual signal energy, and the side gain for each of the M subbands, wherein the first parameter indicates a value relationship between the downmix signal energy and the residual signal energy for each of the M subbands; determine a second parameter based on the downmix signal energy and the residual signal energy of each of the M subbands, the second parameter indicating a value relationship between a first energy sum and a second energy sum, the first energy sum being a sum of the residual signal energy and the downmix signal energy of the M subbands, the second energy sum being a sum of the residual signal energy and the downmix signal energy of the M subbands in a frequency domain signal of a frame previous to the current frame, and the M subbands of the current frame having the same subband index numbers as the M subbands of the previous frame; determining the residual signal coding parameters of the current frame based on the first parameter, the second parameter, and a long-term smoothing parameter of the previous frame of the current frame; 15. The apparatus of claim 14, further configured to:
16. The processor: determining M energy parameters based on the downmix signal energy, the residual signal energy, and the side gain for each of the M subbands, wherein the M energy parameters respectively indicate the value relationship between the downmix signal energy and the residual signal energy for each of the M subbands, and the M energy parameters correspond one-to-one to the M subbands; determining the energy parameter having the maximum value among the M energy parameters as the first parameter; 16. The apparatus of claim 15, further configured to:
17. Among the M energy parameters determined by the processor, the energy parameter of a subband having a subband index number b satisfies the following formula: res_dmx_ratio[b]=res_cod_NRG_S[b] / (res_cod_NRG_S[b]+(1-g(b))・(1-g(b))res_cod_NRG_M[b]+1) 17. The apparatus of claim 16, wherein res_dmx_ratio[b] represents the energy parameter of the subband having subband index number b, where b is greater than or equal to 0 and less than or equal to a preset maximum subband index number, res_cod_NRG_S[b] represents the residual signal energy of the subband having subband index number b, res_cod_NRG_M[b] represents the downmix signal energy of the subband having subband index number b, and g(b) represents a function of side gain side_gain[b] of the subband having subband index number b.
18. The processor: determining a first parameter based on the downmix signal energy and the residual signal energy of each of the M subbands, wherein the first parameter indicates a value relationship between the downmix signal energy and the residual signal energy of each of the M subbands; determine a second parameter based on the downmix signal energy and the residual signal energy of each of the M subbands, the second parameter indicating a value relationship between a first energy sum and a second energy sum, the first energy sum being a sum of the residual signal energy and the downmix signal energy of the M subbands, the second energy sum being a sum of the residual signal energy and the downmix signal energy of the M subbands in a frequency domain signal of a frame previous to the current frame, and the M subbands of the current frame having the same subband index numbers as the M subbands of the previous frame; determining the residual signal coding parameters of the current frame based on the first parameter, the second parameter, and a long-term smoothing parameter of the previous frame of the current frame; 14. The apparatus of claim 12 or 13, further configured to:
19. The processor: determining M energy parameters based on the downmix signal energy and the residual signal energy of each of the M subbands, the M energy parameters respectively indicating the value relationship between the downmix signal energy and the residual signal energy of each of the M subbands, the M energy parameters having a one-to-one correspondence with the M subbands; determining the energy parameter having the maximum value among the M energy parameters as the first parameter; 20. The apparatus of claim 18, further configured to:
20. Among the M energy parameters determined by the processor, the energy parameter of a subband having a subband index number b satisfies the following formula: res_dmx_ratio[b]=res_cod_NRG_S[b] / res_cod_NRG_M[b] 20. The apparatus of claim 19, wherein res_dmx_ratio[b] represents the energy parameter of the subband having subband index number b, where b is greater than or equal to 0 and less than or equal to a preset maximum subband index number, res_cod_NRG_S[b] represents the residual signal energy of the subband having subband index number b, and res_cod_NRG_M[b] represents the downmix signal energy of the subband having subband index number b.
21. The residual signal coding parameter of the current frame is a long-term smoothing parameter of the current frame, and the long-term smoothing parameter of the current frame satisfies the following equation: res_dmx_ratio_lt=res_dmx_ratio・α+res_dmx_ratio_lt_prev・(1−α) where res_dmx_ratio_lt represents the long-term smoothing parameter of the current frame, res_dmx_ratio represents the first parameter, and res_dmx_ratio_lt_prev represents the long-term smoothing parameter of the previous frame of the current frame, and 0<α<1; When the second parameter is greater than a preset third threshold, the value of α when the first parameter is less than a preset second threshold is greater than the value of α when the first parameter is equal to or greater than the preset second threshold, the second threshold is greater than or equal to 0.6 and the third threshold is greater than or equal to 2.7 and less than or equal to 3.7; or When the second parameter is greater than a fifth preset threshold, the value of α when the first parameter is greater than a fourth preset threshold is greater than the value of α when the first parameter is equal to or less than the fourth preset threshold, the fourth threshold is greater than or equal to 0.9, and the fifth threshold is greater than or equal to 0.71; or 21. The apparatus of claim 15, wherein, when the second parameter is greater than or equal to a fifth preset threshold and less than or equal to a third preset threshold, the value of α is less than the value of α when the first parameter is less than a second preset threshold and the second parameter is greater than the third preset threshold, the second threshold being greater than or equal to 0.6 and less than or equal to 2.7 and less than or equal to 3.7, and the fifth threshold being greater than or equal to 0.
71.
22. The processor: when it is determined to encode the residual signal of the M subbands, encoding the downmix signal of the M subbands and the residual signal; or When it is determined not to encode the residual signal of the M subbands, encoding the downmix signal of the M subbands.
22. The apparatus of claim 12, further configured to: