Audio signal encoding method, audio signal decoding method, device, and storage medium
By introducing target coding mode selection and neural network models into audio signal encoding, combined with traditional coding algorithms, the problems of sound quality and compression rate on extremely low bandwidth networks in existing technologies are solved, achieving efficient audio encoding and decoding that is adaptable to devices with different decoding capabilities.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- CHINA MOBILE COMM LTD RES INST
- Filing Date
- 2025-10-14
- Publication Date
- 2026-04-23
AI Technical Summary
Existing audio signal encoding and decoding technologies struggle to achieve higher compression rates while maintaining sound quality on extremely low bandwidth networks. Furthermore, traditional neural network encoding methods have inconsistent structures and parameters, requiring frequent upgrades to encoders and decoders, making it difficult to achieve good results in extremely low-bitrate encoding scenarios.
A target coding mode selection mechanism is adopted, which combines traditional coding algorithms and neural networks. The coding method is selected according to the decoding capability of the communication peer, and signal enhancement and prediction are performed through a pre-trained neural network model to ensure compatibility and improve coding efficiency.
It achieves improved audio encoding efficiency by introducing neural networks while maintaining compatibility with traditional encoding and decoding methods, ensuring that audio bitstream parsing and complete reconstruction are completed on peer devices with different decoding capabilities, thereby improving user experience.
Smart Images

Figure CN2025127510_23042026_PF_FP_ABST
Abstract
Description
Audio signal encoding methods, decoding methods, devices, and storage media
[0001] Cross-references to related applications
[0002] This disclosure claims priority to Chinese Patent Application No. 202411443211.2, filed in China on October 16, 2024, the entire contents of which are incorporated herein by reference. Technical Field
[0003] This disclosure relates to the field of speech encoding and decoding technology, specifically to an audio signal encoding method, decoding method, device, and storage medium. Background Technology
[0004] Encoding and decoding audio signals (speech signals) are essential foundational steps in communication. With the development of communication technology and the widespread adoption of digital devices, the demand for audio signal transmission and storage is increasing daily. Audio signal encoding and decoding technologies, as key technologies for achieving efficient transmission and high-quality reproduction, have become fundamental in multiple fields such as communication, broadcasting, and storage. Traditional analog signal transmission methods suffer from problems such as high bandwidth consumption, susceptibility to interference, and difficulty in long-distance transmission. To address these issues, audio signal encoding technology has emerged. It uses compression algorithms to convert the original analog audio signal into a digital signal and effectively compress it to reduce the bandwidth and storage space required for transmission.
[0005] During the encoding process, the encoder analyzes the characteristics of the audio signal, such as frequency distribution and amplitude variations, and uses appropriate algorithms to encode the audio signal. Common encoding algorithms include, but are not limited to, Linear Predictive Coding (LPC), transform coding such as Moving Pictures Experts Group-Advanced Audio Coding (MPEG-AAC), Subband Coding (SBC), Adaptive Multi-Rate (AMR), and G-series audio codec algorithms. These algorithms achieve effective data compression by removing redundant information from the audio signal. The decoding process is the reverse of encoding. After receiving the encoded audio data, the decoder restores the audio signal with near-original sound quality according to the encoding algorithm. This process is crucial for ensuring the clarity and intelligibility of voice communication.
[0006] However, existing encoding and decoding technologies still face challenges in processing complex speech signals, particularly in improving encoding efficiency. For example, achieving higher compression rates while maintaining sound quality on extremely low-bandwidth networks. Therefore, researching and developing more efficient and intelligent speech and audio encoding and decoding technologies to meet ever-increasing communication and storage demands is an important direction for current technological development. Summary of the Invention
[0007] At least one embodiment of this disclosure provides an audio signal encoding method, decoding method, device, and storage medium for achieving efficient and intelligent audio encoding or decoding processing.
[0008] To solve the above-mentioned technical problems, this disclosure is implemented as follows:
[0009] In a first aspect, embodiments of this disclosure provide an audio signal encoding method, applied to a first device, comprising:
[0010] Obtain the first audio signal of frame t;
[0011] Based on the target encoding pattern, determine the first target signal to be encoded;
[0012] The first target signal is encoded to obtain a first encoded signal, and the first encoded signal is sent to the second device.
[0013] The target encoding mode includes a first mode and a second mode; in the first mode, the first target signal is the first audio signal; in the second mode, the first target signal is the second audio signal, the second audio signal is the one with less information in the first audio signal and the first residual signal, and the first residual signal is the residual between the first audio signal and the predicted signal of the predicted frame t.
[0014] Optionally, the above methods also include:
[0015] The first encoded signal is decoded to obtain the first decoded signal;
[0016] The signal to be enhanced is enhanced to obtain the first enhanced signal;
[0017] Based on the first enhanced signal, the predicted signal for the (t+1)th frame is predicted;
[0018] Wherein, when the second audio signal is the first audio signal, the signal to be enhanced is the first decoded signal; when the second audio signal is the first residual signal, the signal to be enhanced is the first summed signal, which is obtained by adding the first decoded signal to the prediction signal of the t-th frame.
[0019] Optionally, if the target encoding mode is the second mode, the method further includes:
[0020] Send a first message, which indicates whether the first coded signal is obtained based on the residual signal encoding.
[0021] Optionally, enhancing the signal to be enhanced to obtain a first enhanced signal includes:
[0022] The signal to be enhanced is enhanced using a neural network model to obtain a first enhanced signal.
[0023] Optionally, the above methods also include:
[0024] The first coding quality is calculated based on the difference between the first audio signal and the first enhanced signal;
[0025] Based on the first encoding quality, the neural network model is optimized.
[0026] Optionally, the above methods also include:
[0027] Determine the target encoding pattern;
[0028] Wherein, if the second device does not support decoding the second audio signal, the target encoding mode is the first mode; if the second device supports decoding the second audio signal, the target encoding mode is the second mode.
[0029] Optionally, the above methods also include:
[0030] Obtain capability indication information of the second device, the capability indication information being used to indicate whether the second device supports decoding the second audio signal.
[0031] Secondly, embodiments of this disclosure provide an audio signal decoding method, applied to a second device, comprising:
[0032] Receive the first encoded signal and first information of the t-th frame, wherein the first information is used to indicate whether the first encoded signal is encoded based on the residual signal;
[0033] The first encoded signal is decoded to obtain the first decoded signal;
[0034] The signal to be enhanced is enhanced to obtain the first enhanced signal of the t-th frame;
[0035] Wherein, if the first information indicates that the first encoded signal is obtained based on the residual signal encoding, the signal to be enhanced is the first summed signal, which is obtained by adding the first decoded signal to the prediction signal of the t-th frame; if the first information indicates that the first encoded signal is not obtained based on the residual signal encoding, the signal to be enhanced is the first decoded signal.
[0036] Optionally, the above methods also include:
[0037] Based on the first enhanced signal, the predicted signal for the (t+1)th frame is predicted.
[0038] Optionally, enhancing the signal to be enhanced to obtain the first enhanced signal of the t-th frame includes:
[0039] The signal to be enhanced is enhanced using a neural network model to obtain a first enhanced signal.
[0040] Thirdly, embodiments of this disclosure provide a first device, including:
[0041] The first acquisition module is used to acquire the first audio signal of the t-th frame;
[0042] The first determining module is used to determine the first target signal to be encoded according to the target encoding pattern;
[0043] An encoding module is used to encode the first target signal to obtain a first encoded signal;
[0044] The transmitting module is used to transmit the first encoded signal to the second device;
[0045] The target encoding mode includes a first mode and a second mode; in the first mode, the first target signal is the first audio signal; in the second mode, the first target signal is the second audio signal, the second audio signal is the one with less information in the first audio signal and the first residual signal, and the first residual signal is the residual between the first audio signal and the predicted signal of the predicted frame t.
[0046] Optionally, the above-mentioned equipment also includes:
[0047] A decoding module is used to decode the first encoded signal to obtain a first decoded signal;
[0048] The enhancement module is used to enhance the signal to be enhanced to obtain the first enhanced signal;
[0049] The prediction module is used to predict the prediction signal of the (t+1)th frame based on the first enhanced signal.
[0050] Wherein, when the second audio signal is the first audio signal, the signal to be enhanced is the first decoded signal; when the second audio signal is the first residual signal, the signal to be enhanced is the first summed signal, which is obtained by adding the first decoded signal to the prediction signal of the t-th frame.
[0051] Optionally, the sending module is further configured to send first information when the target encoding mode is the second mode, wherein the first information is used to indicate whether the first encoded signal is encoded based on the residual signal.
[0052] Optionally, the enhancement module is further configured to enhance the signal to be enhanced using a neural network model to obtain a first enhanced signal.
[0053] Optionally, the above-mentioned equipment also includes:
[0054] The calculation module is used to calculate the first coding quality based on the difference between the first audio signal and the first enhanced signal;
[0055] An optimization module is used to optimize the neural network model based on the first encoding quality.
[0056] Optionally, the above-mentioned equipment also includes:
[0057] The second determining module is used to determine the target encoding mode;
[0058] Wherein, if the second device does not support decoding the second audio signal, the target encoding mode is the first mode; if the second device supports decoding the second audio signal, the target encoding mode is the second mode.
[0059] Optionally, the above-mentioned equipment also includes:
[0060] The second acquisition module is used to acquire capability indication information of the second device, the capability indication information being used to indicate whether the second device supports decoding the second audio signal.
[0061] Fourthly, embodiments of this disclosure provide a second device, comprising:
[0062] The receiving module is configured to receive a first encoded signal and first information of the t-th frame, wherein the first information is used to indicate whether the first encoded signal is encoded based on the residual signal;
[0063] A decoding module is used to decode the first encoded signal to obtain a first decoded signal;
[0064] The enhancement module is used to enhance the signal to be enhanced, so as to obtain the first enhanced signal of the t-th frame;
[0065] Wherein, if the first information indicates that the first encoded signal is obtained based on the residual signal encoding, the signal to be enhanced is the first summed signal, which is obtained by adding the first decoded signal to the prediction signal of the t-th frame; if the first information indicates that the first encoded signal is not obtained based on the residual signal encoding, the signal to be enhanced is the first decoded signal.
[0066] Optionally, the above-mentioned equipment also includes:
[0067] The prediction module is used to predict the prediction signal of the (t+1)th frame based on the first enhanced signal.
[0068] Optionally, the enhancement module is further configured to enhance the signal to be enhanced using a neural network model to obtain a first enhanced signal.
[0069] Fifthly, embodiments of this disclosure provide a first device, including: a processor, a memory, and a program stored in the memory and executable on the processor, wherein the program, when executed by the processor, implements the steps of the method described in the first aspect.
[0070] In a sixth aspect, embodiments of this disclosure provide a second device, including: a processor, a memory, and a program stored in the memory and executable on the processor, wherein the program, when executed by the processor, implements the steps of the method described in the second aspect.
[0071] In a seventh aspect, embodiments of this disclosure provide a computer-readable storage medium storing a program that, when executed by a processor, implements the steps of the method as described in either the first or second aspect.
[0072] Eighthly, embodiments of this disclosure provide a computer program product including computer instructions that, when executed by a processor, implement the steps of the method as described in either the first or second aspect.
[0073] Compared with related technologies, the audio signal encoding method, decoding method, device and storage medium provided in the embodiments of this disclosure enable the first device to select a suitable encoding method according to the decoding capability of the communication peer, so as to adapt to the decoding capability of the communication peer. While being compatible with traditional encoding and decoding methods, it also further improves the audio encoding efficiency by introducing a neural network. Attached Figure Description
[0074] Various other advantages and benefits will become apparent to those skilled in the art upon reading the following detailed description of preferred embodiments. The accompanying drawings are for illustrative purposes only and are not intended to limit the scope of this disclosure. Furthermore, the same reference numerals denote the same parts throughout the drawings. In the drawings:
[0075] Figure 1 is a schematic diagram of an application scenario according to an embodiment of this disclosure;
[0076] Figure 2 is a schematic diagram of the basic methods of traditional speech coding in related technologies;
[0077] Figure 3 is a schematic diagram of the basic method of speech coding after the introduction of related technologies into neural networks;
[0078] Figure 4 is a schematic diagram of the audio signal encoding and decoding methods of this disclosure applied to the first device and the second device;
[0079] Figure 5 is a flowchart of an audio signal encoding method according to an embodiment of the present disclosure;
[0080] Figure 6 is a flowchart of an audio signal decoding method according to an embodiment of the present disclosure;
[0081] Figure 7 is a schematic diagram of the structure of a first device according to an embodiment of the present disclosure;
[0082] Figure 8 is a schematic diagram of the structure of a second device according to an embodiment of the present disclosure.
[0083] Figure 9 is a schematic diagram of the structure of a first device according to another embodiment of the present disclosure;
[0084] Figure 10 is a schematic diagram of the structure of a second device according to another embodiment of this disclosure. Detailed Implementation
[0085] The terms "first," "second," etc., used in this disclosure are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such terms can be used interchangeably where appropriate so that embodiments of this disclosure can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first" and "second" are generally of the same class, not limited in number; for example, the first object can be one or more. Furthermore, "or" in this disclosure indicates at least one of the connected objects. For example, "A or B" covers three scenarios: Scenario 1: including A but not B; Scenario 2: including B but not A; Scenario 3: including both A and B. The character " / " generally indicates that the preceding and following objects are in an "or" relationship.
[0086] The term "instruction" in this disclosure can be either a direct instruction (or explicit instruction) or an indirect instruction (or implicit instruction). A direct instruction can be understood as one in which the sender explicitly informs the receiver of specific information, the operation to be performed, or the requested result, etc.; an indirect instruction can be understood as one in which the receiver determines the corresponding information based on the instruction sent by the sender, or makes a judgment and determines the operation to be performed or the requested result, etc., based on the judgment result.
[0087] It is worth noting that the technologies described in this disclosure are not limited to Long Term Evolution (LTE) / LTE-Advanced (LTE-A) systems, but can also be used in other wireless communication systems, such as Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Frequency Division Multiple Access (FDMA), Orthogonal Frequency Division Multiple Access (OFDMA), Single-carrier Frequency-Division Multiple Access (SC-FDMA), or other systems. The terms "system" and "network" in this disclosure are often used interchangeably, and the described technologies can be used with the systems and radio technologies mentioned above, as well as with other systems and radio technologies. The following description describes New Radio (NR) systems for illustrative purposes, and the term NR is used in most of the following description; however, these technologies can also be applied to systems other than NR systems, such as 6th Generation (6G) communication systems.
[0088] Figure 1 shows a block diagram of a wireless communication system applicable to embodiments of the present disclosure. The wireless communication system includes a terminal 11 and a network-side device 12. The terminal 11 can be a mobile phone, tablet computer, laptop computer, notebook computer, personal digital assistant (PDA), handheld computer, netbook, ultra-mobile personal computer (UMPC), mobile internet device (MID), augmented reality (AR), virtual reality (VR) device, robot, wearable device, flight vehicle, vehicle user equipment (VUE), shipboard equipment, pedestrian user equipment (PUE), smart home (home devices with wireless communication capabilities, such as refrigerators, televisions, washing machines, or furniture), game console, personal computer (PC), ATM, or self-service machine, etc. Wearable devices include: smartwatches, smart bracelets, smart headphones, smart glasses, smart jewelry (smart bracelets, smart chains, smart rings, smart necklaces, smart anklets, smart anklets, etc.), smart wristbands, smart clothing, etc. Among these, in-vehicle devices can also be referred to as in-vehicle terminals, in-vehicle controllers, in-vehicle modules, in-vehicle components, in-vehicle chips, or in-vehicle units, etc. It should be noted that the specific type of terminal 11 is not limited in this disclosure embodiment. Network-side equipment 12 may include access network equipment or core network equipment, wherein access network equipment may also be referred to as Radio Access Network (RAN) equipment, radio access network function, or radio access network unit. Access network equipment may include base stations, Wireless Local Area Network (WLAN) access points (APs), or Wireless Fidelity (WiFi) nodes, etc.In this context, a base station may be referred to as a Node B (NB), Evolved Node B (eNB), Next Generation Node B (gNB), New Radio Node B (NR Node B), Access Point, Relay Base Station (RBS), Serving Base Station (SBS), Base Transceiver Station (BTS), Radio Base Station, Radio Transceiver, Basic Service Set (BSS), Extended Service Set (ESS), Home Node B (HNB), Home Evolved Node B, Transmission Reception Point (TRP), or any other suitable term in the relevant field, as long as the same technical effect is achieved. The base station is not limited to any specific technical terminology. It should be noted that in this disclosure, only a base station in an NR system is used as an example for description, and the specific type of base station is not limited.
[0089] Core network equipment may include, but is not limited to, at least one of the following: core network node, core network function, Mobility Management Entity (MME), Access and Mobility Management Function (AMF), Session Management Function (SMF), User Plane Function (UPF), Policy Control Function (PCF), Policy and Charging Rules Function (PCRF), Edge Application Server Discovery Function (EASDF), Unified Data Management (UDM), Unified Data Repository (UDR), Home Subscriber Server (HSS), Centralized network configuration (CNC), Network Repository Function (NRF), Network Exposure Function (NEF), Local NEF (or L-NEF), Binding Support Function (BSF), and Application Function. Functions, AFs, etc. It should be noted that this disclosure only uses the core network equipment in the NR system as an example for introduction, and does not limit the specific type of core network equipment.
[0090] As shown in Figure 2, the basic methods of traditional speech coding in related technologies can be divided into waveform coding and parametric coding. Waveform coding is a digital speech signal formed by sampling, quantizing, and encoding the waveform signal of analog speech in the time domain. Parametric coding is based on the articulation mechanism of human speech, finding characteristic parameters that represent speech, and encoding these characteristic parameters. There are also hybrid coding methods that combine waveform coding and parametric coding.
[0091] To further improve coding efficiency, neural networks have been gradually introduced into speech and audio coding. Speech or audio signals are fed into neural networks / artificial intelligence (AI) for encoding, and the sound signal is then recovered through the neural network during decoding. The basic coding method is shown in Figure 3.
[0092] Neural network-based encoding methods in related technologies typically suffer from the following problems:
[0093] 1. The structure of neural networks is not fixed. Currently, various companies and research institutions have proposed their own neural network architectures, which differ in the number of network layers, the type of neural computation units in each layer, and the number of neural units in each layer.
[0094] 2. The parameters of neural networks are not fixed. Currently, various companies and research institutions are continuously training models, and the parameters of the same network are updated rapidly. Therefore, encoders and decoders, as well as encoded speech and audio files, all face the problem of being upgraded at any time.
[0095] 3. Traditional speech coding methods have achieved good results in speech and audio coding, and can achieve good coding efficiency in most scenarios. However, further improvements are needed in scenarios such as extremely low bitrate coding. AI-based encoders are still in their infancy, and since they are incompatible with traditional encoders, further data accumulation is required to achieve better coding performance.
[0096] To address at least one of the above problems, embodiments of this disclosure provide an audio signal encoding method and a decoding method, which can reduce or avoid the occurrence of the above situations, improve communication efficiency, and enhance user experience.
[0097] This disclosure provides an audio signal encoding and decoding method that, while being compatible with traditional encoding and decoding methods, introduces a neural network to further improve the efficiency of audio encoding. Furthermore, even when the decoding end does not support a neural network / neural network architecture, or when the neural network parameters are different, this disclosure can still complete the parsing of the audio bitstream and output a complete reconstructed sound signal.
[0098] Figure 4 is a schematic diagram illustrating the application of the audio signal encoding and decoding methods of this disclosure to a first device and a second device, wherein the first device can be a terminal and the second device can be a network-side device, or the first device can be a network-side device and the second device can be a terminal. The terminal / network-side device can simultaneously serve as both the encoding and decoding end of the audio signal. For example, in a practical two-way communication scenario, one device needs to both send audio encoded signals to the other device and receive and decode the audio encoded signals sent by the other device.
[0099] The methods of this disclosure embodiment will be described below from the perspectives of the encoding and decoding ends of audio encoding. It is understood that the encoding end device (such as a terminal or network-side device) may also have related functions or modules of the decoding end. Similarly, the decoding end device (such as a terminal or network-side device) may also have related functions or modules of the encoding end.
[0100] Figure 5 is a flowchart illustrating an example of the audio signal encoding method described in this disclosure when applied to a first device.
[0101] The first device may be a terminal or a network-side device, and this embodiment does not specifically limit it. As shown in Figure 5, the method includes:
[0102] Step 51: Obtain the first audio signal of frame t.
[0103] Here, the audio signal of frame t (referred to as the first audio signal for ease of description) refers to the source audio signal of frame t. The first device performs row encoding based on the first audio signal and sends the encoded signal of frame t (referred to as the first encoded signal for ease of description) to the second device. The second device is the peer device in the communication between the first device and the second device. The source audio signals of other frames are processed in the same way as shown in Figure 5.
[0104] Step 52: Determine the first target signal to be encoded according to the target encoding pattern.
[0105] Here, the target encoding mode includes a first mode and a second mode. In the first mode, the first target signal is the first audio signal; in the second mode, the first target signal is the second audio signal, which is the one with less information content between the first audio signal and the first residual signal, and the first residual signal is the residual between the first audio signal and the predicted signal of the predicted frame t.
[0106] In this paper, the first mode is sometimes referred to as the traditional encoding mode, and the second mode is sometimes referred to as the AI encoding mode. In the first mode, the first audio signal can be directly encoded using traditional encoding algorithms, including but not limited to LPC, transform coding (such as MPEG-AAC), SBC, AMR, and G-series audio codecs. In the second mode, the audio signal with less information in the first audio signal and the first residual signal is encoded, and the encoding algorithm used can also be one of the aforementioned traditional encoding algorithms. Furthermore, the predicted signal of the t-th frame is predicted based on the audio signal of the (t-1)-th frame obtained through neural network enhancement. When t equals 1, the signal of the (t-1)-th frame can be a zero signal. The neural network is pre-trained, and the training method will be explained later. Alternatively, the neural network can be further trained online.
[0107] Step 53: Encode the first target signal to obtain a first encoded signal, and send the first encoded signal to the second device.
[0108] Here, in this embodiment of the disclosure, the encoder can use the aforementioned conventional encoding algorithm or other proprietary algorithm to encode the first target signal determined in step 52, thereby obtaining the encoded signal of the t-th frame (i.e., the first encoded signal). The first encoded signal can be transmitted to the second device via wired or wireless means. For example, it can be transmitted to the second device via a communication network.
[0109] Through the above steps, the first device in this embodiment can choose to directly encode the first audio signal, or encode the one with less information between the first audio signal and the first residual signal. This allows the first device to select a suitable encoding method based on the decoding capability of the communication peer (second device), adapting to the peer's decoding ability. While maintaining compatibility with traditional encoding and decoding methods, it further improves audio encoding efficiency by introducing a neural network. For example, if the communication peer does not support a neural network / neural network architecture, or if the neural network parameters are different, the first device can choose the first mode; if the communication peer supports a neural network / neural network architecture, or if the neural network parameters are the same, the first device can choose the second mode. Thus, even if the communication peer does not support the decoding capability corresponding to the AI encoding mode, it can still ensure that the communication peer can complete the parsing of the audio bitstream and obtain the complete reconstructed sound signal.
[0110] In addition, when the target encoding mode is the second mode, the first device may also send first information to the second device when sending the first encoding signal to the second device. The first information is used to indicate whether the first encoding signal is encoded based on the residual signal.
[0111] In this embodiment of the disclosure, when the target encoding mode is either the first mode or the second mode, the first device can reconstruct and enhance the first encoded signal to obtain the prediction signal of the (t+1)th frame.
[0112] Specifically, the first encoded signal can be decoded using a decoder (such as a traditional decoder) to obtain a first decoded signal (i.e., the audio signal reconstructed by the decoder). The decoding algorithm can be set according to the encoding algorithm used. Then, the signal to be enhanced is enhanced to obtain a first enhanced signal. Wherein, if the second audio signal is the first audio signal, the signal to be enhanced is the first decoded signal; if the second audio signal is the first residual signal, the signal to be enhanced is a first summed signal, which is obtained by adding the first decoded signal to the predicted signal of the t-th frame. For example, a pre-trained neural network model can be used to enhance the signal to be enhanced to obtain the first enhanced signal of the t-th frame. Then, based on the first enhanced signal, the predicted signal of the (t+1)-th frame is predicted. When predicting the audio signal of the (t+1)-th frame, various existing prediction algorithms can be used in this embodiment to predict the predicted signal of the (t+1)-th frame based on historical enhanced signals, including the first enhanced signal of the t-th frame. This embodiment does not limit the specific prediction algorithm.
[0113] In this embodiment of the disclosure, the neural network model can be pre-trained. During the training process, the coding quality can be calculated based on the difference between the enhanced signal of the i-th frame obtained by the neural network model and the i-th frame audio signal of the source audio signal. Then, based on the coding quality, the neural network model is trained until a preset training termination condition is met, thereby obtaining the neural network model.
[0114] Furthermore, embodiments of this disclosure can also perform online training on the neural network model. For example, during the transmission of the first audio signal, the first device can calculate a first coding quality based on the difference between the first audio signal and the first enhanced signal; then, based on the first coding quality, the neural network model can be optimized, thereby fine-tuning the neural network model.
[0115] To adapt to the decoding capabilities of the communication peer (second device), the first device can select a suitable encoding mode based on the second device's decoding capabilities. Before step 52 above, the first device can further determine a target encoding mode. Wherein, if the second device does not support decoding the second audio signal, the target encoding mode is the first mode; if the second device supports decoding the second audio signal, the target encoding mode is the second mode. It should be noted that here, the ability to decode the second audio signal is used as an example to represent the decoding capability of the second device. It can be understood that supporting the decoding of the second audio signal means supporting the decoding of the bitstream obtained by the AI encoding mode mentioned above; more specifically, it means supporting the decoding of the encoded signal obtained by encoding based on the one with less information in the original audio signal and the residual signal. The residual signal is the residual between the original audio signal and the predicted audio signal.
[0116] In this embodiment, the first device can obtain capability indication information of the second device, which indicates whether the second device supports decoding the second audio signal. Then, the first device determines the target encoding mode based on the capability indication information of the second device. The capability indication information of the second device can be pre-configured on the first device side, obtained from the second device, or obtained from the third device side. For example, during or after establishing a communication link with the second device, the first device sends capability request information to the second device, and the second device sends its own capability indication information to the first device based on the capability request information. Alternatively, both the first and second devices send their own capability indication information to the third device, and the third device stores the capability indication information of each device. Thus, when the first device needs to communicate with the second device, it can request the capability information of the second device from the third device and receive the capability indication information of the second device sent by the third device.
[0117] Figure 6 is a flowchart illustrating an example of the audio signal decoding method described in this disclosure when applied to a second device. The second device can be a network-side device or a terminal; this disclosure does not specifically limit its application. As shown in Figure 6, the method includes:
[0118] Step 61: Receive the first encoded signal and first information of the t-th frame, wherein the first information is used to indicate whether the first encoded signal is encoded based on the residual signal.
[0119] Step 62: Decode the first encoded signal to obtain the first decoded signal.
[0120] Here, based on the encoding algorithm of the first device, the second device uses the corresponding decoding algorithm to decode, thereby obtaining the first decoded signal.
[0121] Step 63: Enhance the signal to be enhanced to obtain the first enhanced signal of the t-th frame.
[0122] Here, when the first information indicates that the first encoded signal is obtained based on residual signal encoding, the signal to be enhanced is a first summed signal, which is obtained by adding the first decoded signal to the prediction signal of the t-th frame; when the first information indicates that the first encoded signal is not obtained based on residual signal encoding, the signal to be enhanced is a first decoded signal. The prediction signal of the t-th frame is predicted by the second device based on historical enhancement signals, including the enhancement signal of the (t-1)-th frame. The second device can use various existing prediction algorithms for prediction. The prediction algorithm used by the second device is the same as that used by the first device, and this embodiment does not limit the specific prediction algorithm.
[0123] In addition, the second device can use a neural network model to enhance the signal to be enhanced, thereby obtaining a first enhanced signal. Typically, the neural network model used by the second device is the same as that used by the first device.
[0124] Through the above steps, the embodiments of this disclosure can achieve the decoding of the encoded signal sent by the first device.
[0125] Furthermore, the second device can also predict the predicted signal for frame t+1 based on the first enhanced signal. For example, the second device predicts the predicted signal for frame t+1 based on historical enhanced signals, including the first enhanced signal for frame t.
[0126] The following section, referring to Figure 4 above, will further illustrate the above method through a more specific example of speech or audio signal encoding and decoding. The spectral range of speech signals is typically narrower than that of audio signals.
[0127] The encoding process for the voice or audio signal on the first device side includes:
[0128] Step 1: Receive the speech or audio signal S at the current time t. t The received signal is simultaneously sent to the mode selection module, the prediction error calculation module, and the coding quality calculation module.
[0129] Here, in Step 1, S t The signal duration can be 10ms, 20ms, or other durations. tThe signal can be sampled at 8K, 16K, 32K, and other sampling rates. t The signal bit depth can be 8 bits, 16 bits, and other bit depths. t The number of signal channels can be 1, 2, or other numbers.
[0130] Step 2, the prediction error calculation module calculates the input speech or audio signal S at time t. t The speech or audio signal P predicted at time t t To obtain the residual R of the speech or audio signal at time t. t .
[0131] Step 3: The mode selection module first completes the selection between traditional encoding mode and AI encoding mode. If traditional encoding mode is selected, the mode selection module inputs S. t If it is AI encoding mode, the mode selection module calculates S. t and R t The information content is considered, and the output with less information content is selected, denoted as the AI original signal mode and the AI residual mode, respectively. There are several methods for calculating the information content, for example: one method is to calculate S... t and R t The signal energy should be selected based on its low energy; one approach is to calculate S. t and R t The distribution of the spectrum should be narrow. This disclosure does not specifically limit this aspect.
[0132] Step 4: Encode the input signal into speech or audio format using an existing standard encoding method. The output is the encoded bitstream C. s C s This is a bitstream that fully conforms to traditional encoding formats.
[0133] Step 5, C s It is sent to the decoder for decoding. In traditional encoding mode, C s The bitstream can be decoded into a complete bitstream on any standard decoder; in AI encoding mode, C s The bitstream is decoded on any standard decoder to obtain a bitstream that is a mixture of residual signal and / or original signal (e.g., the i-th frame is the residual signal, the (i+1)-th frame is the original signal, etc.).
[0134] Step 6: For the traditional encoding mode or the AI original signal mode, the decoded signal is denoted as SE. t The code is fed into a neural network model for encoding enhancement, and the output is SO. t .
[0135] Step 7, for the AI residual mode, the decoded residual signal REt The signal is fed into the residual summing module and compared with the predicted speech or audio signal P at time t. t Sum the results to obtain PE. t .
[0136] Step 8, Neural Network Model for PE t Perform encoding enhancement, output as SO t .
[0137] Here, in Step 8, the neural network model can be a convolutional neural network, a long short-term memory network (LSTM), a Transformer, or a network such as WaveNet.
[0138] Step 9, Next Frame Prediction Module, input is the currently reconstructed SO t The output is the predicted signal P for the next time step, i.e., time t+1. t+1 .
[0139] Step 10, Encoding quality calculation module, for the original signal S t, Reconstructed SO t Quality calculations can be performed using Perceptual Evaluation of Speech Quality (PESQ), Perceptual Objective Listening Quality Assessment (POLQA), and other quality calculation algorithms. The input to the coding quality calculation module is the speech or audio coding quality at the current time t, denoted as Q. t .
[0140] Step 11, Encoding Quality Q t The input is fed into the neural network training module as the error calculation value for neural network training, and is used for training or optimizing the neural network model.
[0141] In Step 11, there are various error calculation strategies for training neural network models, such as cross-entropy and minimum mean square error.
[0142] The above steps are performed during the coding process.
[0143] The audio signal decoding process on the second device side includes:
[0144] Step 1, receive C s It is sent to the decoder for decoding. In traditional encoding mode, C sThe bitstream can be decoded into a complete bitstream on any standard decoder; in AI encoding mode, C s The bitstream is decoded on any standard decoder to obtain a mixed bitstream of residual and original data.
[0145] If the decoder is a traditional decoder, then its decoded signal SE t Output directly to complete the entire decoding process.
[0146] Step 2: If the decoder supports AI mode, then for either the traditional encoding mode or the AI original signal mode, the decoded signal SE... t The code is fed into a neural network model for encoding enhancement, and the output is SO. t .
[0147] Step 3, for the AI residual mode, the decoded residual signal RE t The signal is fed into the residual summing module and compared with the predicted speech or audio signal P at time t. t Sum the results to obtain PE. t .
[0148] Step 4, for PE t Perform encoding enhancement, output as SO t SO t This refers to the decoded audio signal output in AI mode.
[0149] Step 5, the next frame prediction module, the input is the currently reconstructed SO t The output is the predicted signal P for the next time step, i.e., time t+1. t+1 .
[0150] The above steps are performed during the decoding process.
[0151] This example also provides a speech and audio coding device (i.e., the first device) as shown in Figure 4, including a conventional codec, a mode selection module, a prediction error module, a next frame prediction module, a residual summing module, a coding quality calculation module, a neural network model, a conventional decoder, and a coding quality-based optimization and network training strategy module.
[0152] This example also provides a speech and audio decoding device (i.e., a second device) as shown in Figure 4, including a conventional decoder, a next frame prediction module, a residual summing module, a neural network model, etc.
[0153] As can be seen from the examples above, the embodiments of this disclosure, while fully compatible with traditional encoding and decoding methods, introduce neural networks to further improve the efficiency of speech and audio encoding. Even when the decoding end does not support neural networks, or when the neural network architecture and parameters are different, it can still complete the parsing of speech or audio bitstreams and output a complete reconstructed sound signal. The above-described solutions of the embodiments of this disclosure are compatible with a large number of existing devices and can support massive amounts of existing speech and audio content.
[0154] Based on the above-described solutions of the embodiments of this disclosure, efficient encoding techniques can be employed to significantly reduce the network bandwidth required for audio data transmission between devices, thereby reducing the operating costs of network operators. For example, advanced audio compression algorithms can reduce data traffic while maintaining sound quality, which means mobile network operators can serve more users without increasing additional infrastructure investment. Secondly, with the rapid development of the mobile internet, users' demand for high-quality audio content is constantly growing. High-quality voice and audio encoding technologies can provide a better user experience, enhance user stickiness, and thus improve user satisfaction and loyalty. This not only helps attract new users but also helps retain existing users and increase user lifetime value. In addition, advancements in encoding technology also support the development of new business models, such as online music and video streaming services. Finally, with the promotion of 5G, voice and audio encoding technologies have broad application prospects in fields such as the Internet of Things and autonomous driving, and the solutions of the embodiments of this disclosure can be widely applied to the above scenarios.
[0155] Please refer to Figure 7. This embodiment of the disclosure also provides a first device, including:
[0156] The first acquisition module 701 is used to acquire the first audio signal of the t-th frame;
[0157] The first determining module 702 is used to determine the first target signal to be encoded according to the target encoding mode;
[0158] Encoding module 703 is used to encode the first target signal to obtain a first encoded signal;
[0159] The transmitting module 704 is used to transmit the first encoded signal to the second device;
[0160] The target encoding mode includes a first mode and a second mode; in the first mode, the first target signal is the first audio signal; in the second mode, the first target signal is the second audio signal, the second audio signal is the one with less information in the first audio signal and the first residual signal, and the first residual signal is the residual between the first audio signal and the predicted signal of the predicted frame t.
[0161] Through the above modules, the embodiments of this disclosure can achieve efficient and intelligent audio encoding or decoding processing.
[0162] Optionally, the above-mentioned equipment also includes:
[0163] A decoding module is used to decode the first encoded signal to obtain a first decoded signal;
[0164] The enhancement module is used to enhance the signal to be enhanced to obtain the first enhanced signal;
[0165] The prediction module is used to predict the prediction signal of the (t+1)th frame based on the first enhanced signal.
[0166] Wherein, when the second audio signal is the first audio signal, the signal to be enhanced is the first decoded signal; when the second audio signal is the first residual signal, the signal to be enhanced is the first summed signal, which is obtained by adding the first decoded signal to the prediction signal of the t-th frame.
[0167] Optionally, the sending module is further configured to send first information when the target encoding mode is the second mode, wherein the first information is used to indicate whether the first encoded signal is encoded based on the residual signal.
[0168] Optionally, the enhancement module is further configured to enhance the signal to be enhanced using a neural network model to obtain a first enhanced signal.
[0169] Optionally, the above-mentioned equipment also includes:
[0170] The calculation module is used to calculate the first coding quality based on the difference between the first audio signal and the first enhanced signal;
[0171] An optimization module is used to optimize the neural network model based on the first encoding quality.
[0172] Optionally, the above-mentioned equipment also includes:
[0173] The second determining module is used to determine the target encoding mode;
[0174] Wherein, if the second device does not support decoding the second audio signal, the target encoding mode is the first mode; if the second device supports decoding the second audio signal, the target encoding mode is the second mode.
[0175] Optionally, the above-mentioned equipment also includes:
[0176] The second acquisition module is used to acquire capability indication information of the second device, the capability indication information being used to indicate whether the second device supports decoding the second audio signal.
[0177] It should be noted that the device in this embodiment corresponds to the method applied to the first device side described above. The implementation methods in each of the above embodiments are applicable to the embodiments of this device and can achieve the same technical effect. The device provided in this disclosure can implement all the method steps implemented in the above method embodiments and can achieve the same technical effect. Therefore, the parts that are the same as those in the method embodiments and the beneficial effects will not be described in detail here.
[0178] Please refer to Figure 8. This embodiment of the disclosure also provides a second device, including:
[0179] The receiving module 801 is used to receive the first encoded signal and first information of the t-th frame, wherein the first information is used to indicate whether the first encoded signal is encoded based on the residual signal;
[0180] Decoding module 802 is used to decode the first encoded signal to obtain a first decoded signal;
[0181] Enhancement module 803 is used to enhance the signal to be enhanced to obtain the first enhanced signal of the t-th frame;
[0182] Wherein, if the first information indicates that the first encoded signal is obtained based on the residual signal encoding, the signal to be enhanced is the first summed signal, which is obtained by adding the first decoded signal to the prediction signal of the t-th frame; if the first information indicates that the first encoded signal is not obtained based on the residual signal encoding, the signal to be enhanced is the first decoded signal.
[0183] Through the above modules, the embodiments of this disclosure can achieve efficient and intelligent audio encoding or decoding processing.
[0184] Optionally, the above-mentioned equipment also includes:
[0185] The prediction module is used to predict the prediction signal of the (t+1)th frame based on the first enhanced signal.
[0186] Optionally, the enhancement module is further configured to enhance the signal to be enhanced using a neural network model to obtain a first enhanced signal.
[0187] It should be noted that the device in this embodiment corresponds to the method applied to the second device side described above. The implementation methods in each of the above embodiments are applicable to the embodiments of this device and can achieve the same technical effect. The device provided in this disclosure can implement all the method steps implemented in the above method embodiments and can achieve the same technical effect. Therefore, the parts that are the same as those in the method embodiments and the beneficial effects will not be described in detail here.
[0188] Another embodiment of the present disclosure provides a first device, as shown in FIG9, which includes a transceiver 910, a processor 900, a memory 920, and a program or instructions stored in the memory 920 and executable on the processor 900. When the processor 900 executes the program or instructions, it implements various processes of the above audio signal encoding method embodiment and can achieve the same technical effect. To avoid repetition, it will not be described again here.
[0189] The transceiver 910 is used to receive and send data under the control of the processor 900.
[0190] In Figure 9, the bus architecture can include any number of interconnected buses and bridges, specifically linking various circuits of one or more processors represented by processor 900 and memory represented by memory 920. The bus architecture can also link various other circuits such as peripheral devices, voltage regulators, and power management circuits, which are well known in the art and therefore will not be described further herein. The bus interface provides an interface. Transceiver 910 can be multiple elements, including transmitters and receivers, providing a unit for communicating with various other devices over a transmission medium. For different user equipment, user interface 930 can also be an interface capable of connecting external or internal devices, including but not limited to keypads, displays, speakers, microphones, joysticks, etc.
[0191] The processor 900 is responsible for managing the bus architecture and general processing, while the memory 920 can store the data used by the processor 900 during operation.
[0192] Another embodiment of this disclosure provides a second device, as shown in FIG10, including a transceiver 1010, a processor 1000, a memory 1020, and a program or instructions stored in the memory 1020 and executable on the processor 1000; when the processor 1000 executes the program or instructions, it implements the various processes of the above-described audio signal decoding method embodiment and can achieve the same technical effect. To avoid repetition, it will not be described again here.
[0193] The transceiver 1010 is used to receive and send data under the control of the processor 1000.
[0194] In Figure 10, the bus architecture may include any number of interconnected buses and bridges, specifically linking various circuits of one or more processors represented by processor 1000 and memory represented by memory 1020. The bus architecture may also link various other circuits such as peripheral devices, voltage regulators, and power management circuits, which are well known in the art and therefore will not be described further herein. The bus interface provides an interface. Transceiver 1010 may be multiple elements, including transmitters and receivers, providing units for communicating with various other devices over a transmission medium. Processor 1000 is responsible for managing the bus architecture and general processing, and memory 1020 may store data used by processor 1000 during operation.
[0195] This disclosure also provides a computer-readable storage medium storing a computer program. When executed by a processor, the computer program implements the various processes of the above-described audio signal encoding and decoding methods, achieving the same technical effects. To avoid repetition, these will not be described again here. The computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0196] This disclosure also provides a computer program product, including computer instructions. When the computer instructions are executed by a processor, they implement the various processes of the above-described audio signal encoding and decoding methods, and achieve the same technical effects. To avoid repetition, these will not be described again here.
[0197] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
[0198] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this disclosure, in essence, or the part that contributes to the related technology, can be embodied in the form of a software product. The computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of this disclosure.
[0199] The embodiments of this disclosure have been described above with reference to the accompanying drawings. However, this disclosure is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this disclosure without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this disclosure.
Claims
1. An audio signal encoding method, applied to a first device, the method comprising: Obtain the first audio signal of frame t; Based on the target encoding pattern, determine the first target signal to be encoded; The first target signal is encoded to obtain a first encoded signal, and the first encoded signal is sent to the second device. The target encoding mode includes a first mode and a second mode; in the first mode, the first target signal is the first audio signal; in the second mode, the first target signal is the second audio signal, the second audio signal is the one with less information in the first audio signal and the first residual signal, and the first residual signal is the residual between the first audio signal and the predicted signal of the predicted frame t.
2. The method according to claim 1, further comprising: The first encoded signal is decoded to obtain the first decoded signal; The signal to be enhanced is enhanced to obtain the first enhanced signal; Based on the first enhanced signal, the predicted signal for the (t+1)th frame is predicted; Wherein, when the second audio signal is the first audio signal, the signal to be enhanced is the first decoded signal; when the second audio signal is the first residual signal, the signal to be enhanced is the first summed signal, which is obtained by adding the first decoded signal to the prediction signal of the t-th frame.
3. The method of claim 2, wherein, When the target encoding mode is the second mode, the method further includes: Send a first message, which indicates whether the first coded signal is obtained based on the residual signal encoding.
4. The method of claim 2, wherein, The enhancement of the signal to be enhanced to obtain a first enhanced signal includes: The signal to be enhanced is enhanced using a neural network model to obtain a first enhanced signal.
5. The method according to claim 4, further comprising: The first coding quality is calculated based on the difference between the first audio signal and the first enhanced signal; Based on the first encoding quality, the neural network model is optimized.
6. The method according to claim 1, further comprising: Determine the target encoding pattern; Wherein, if the second device does not support decoding the second audio signal, the target encoding mode is the first mode; if the second device supports decoding the second audio signal, the target encoding mode is the second mode.
7. The method according to claim 6, further comprising: Obtain capability indication information of the second device, the capability indication information being used to indicate whether the second device supports decoding the second audio signal.
8. A method for decoding an audio signal, applied to a second device, the method comprising: Receive the first encoded signal and first information of the t-th frame, wherein the first information is used to indicate whether the first encoded signal is encoded based on the residual signal; The first encoded signal is decoded to obtain the first decoded signal; The signal to be enhanced is enhanced to obtain the first enhanced signal of the t-th frame; Wherein, if the first information indicates that the first encoded signal is obtained based on the residual signal encoding, the signal to be enhanced is the first summed signal, which is obtained by adding the first decoded signal to the prediction signal of the t-th frame; if the first information indicates that the first encoded signal is not obtained based on the residual signal encoding, the signal to be enhanced is the first decoded signal.
9. The method according to claim 8, further comprising: Based on the first enhanced signal, the predicted signal for the (t+1)th frame is predicted.
10. The method of claim 8, wherein, The enhancement of the signal to be enhanced to obtain the first enhanced signal of the t-th frame includes: The signal to be enhanced is enhanced using a neural network model to obtain a first enhanced signal.
11. A first device, comprising: The first acquisition module is used to acquire the first audio signal of the t-th frame; The first determining module is used to determine the first target signal to be encoded according to the target encoding pattern; An encoding module is used to encode the first target signal to obtain a first encoded signal; The transmitting module is used to transmit the first encoded signal to the second device; The target encoding mode includes a first mode and a second mode; in the first mode, the first target signal is the first audio signal; in the second mode, the first target signal is the second audio signal, the second audio signal is the one with less information in the first audio signal and the first residual signal, and the first residual signal is the residual between the first audio signal and the predicted signal of the predicted frame t.
12. The device according to claim 11, further comprising: A decoding module is used to decode the first encoded signal to obtain a first decoded signal; The enhancement module is used to enhance the signal to be enhanced to obtain the first enhanced signal; The prediction module is used to predict the prediction signal of the (t+1)th frame based on the first enhanced signal. Wherein, when the second audio signal is the first audio signal, the signal to be enhanced is the first decoded signal; when the second audio signal is the first residual signal, the signal to be enhanced is the first summed signal, which is obtained by adding the first decoded signal to the prediction signal of the t-th frame.
13. The device according to claim 12, wherein, The sending module is further configured to send first information when the target encoding mode is the second mode, wherein the first information is used to indicate whether the first encoded signal is encoded based on the residual signal.
14. The apparatus of claim 12, wherein, The enhancement module is further configured to enhance the signal to be enhanced using a neural network model to obtain a first enhanced signal.
15. The apparatus of claim 14, further comprising: The calculation module is used to calculate the first coding quality based on the difference between the first audio signal and the first enhanced signal; An optimization module is used to optimize the neural network model based on the first encoding quality.
16. The apparatus of claim 11, further comprising: The second determining module is used to determine the target encoding mode; Wherein, if the second device does not support decoding the second audio signal, the target encoding mode is the first mode; if the second device supports decoding the second audio signal, the target encoding mode is the second mode.
17. The apparatus of claim 16, further comprising: The second acquisition module is used to acquire capability indication information of the second device, the capability indication information being used to indicate whether the second device supports decoding the second audio signal.
18. A second device, comprising: The receiving module is configured to receive a first encoded signal and first information of the t-th frame, wherein the first information is used to indicate whether the first encoded signal is encoded based on the residual signal; A decoding module is used to decode the first encoded signal to obtain a first decoded signal; The enhancement module is used to enhance the signal to be enhanced, so as to obtain the first enhanced signal of the t-th frame; Wherein, if the first information indicates that the first encoded signal is obtained based on the residual signal encoding, the signal to be enhanced is the first summed signal, which is obtained by adding the first decoded signal to the prediction signal of the t-th frame; if the first information indicates that the first encoded signal is not obtained based on the residual signal encoding, the signal to be enhanced is the first decoded signal.
19. The apparatus of claim 18, further comprising: The prediction module is used to predict the prediction signal of the (t+1)th frame based on the first enhanced signal.
20. The apparatus of claim 18, wherein, The enhancement module is further configured to enhance the signal to be enhanced using a neural network model to obtain a first enhanced signal.
21. A first device comprising: Transceiver, processor, memory, and programs or instructions stored in the memory and executable on the processor; When the processor executes the program or instructions, it implements the steps of the device as described in any one of claims 1 to 7.
22. A second device comprising: Transceiver, processor, memory, and programs or instructions stored in the memory and executable on the processor; When the processor executes the program or instructions, it implements the steps of the device as described in any one of claims 8 to 10.
23. A computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the steps of the method as claimed in any one of claims 1 to 7, or implements the steps of the method as claimed in any one of claims 8 to 10.
24. A computer program product comprising computer instructions that, when executed by a processor, implement the steps of the method as claimed in any one of claims 1 to 7, or implement the steps of the method as claimed in any one of claims 8 to 10.
Citation Information
Patent Citations
Audio encoder, audio decoder and audio processor having a dynamically variable warping characteristic
CN101501759A
Method for encoding audio signals, apparatus for encoding audio signals, method for decoding audio signals and apparatus for decoding audio signals
CN105264595A
Audio signal coding method, decoding method, equipment and storage medium
CN119323964A
Multi-mode speech encoder and decoder
CN1275228A
Apparatus and method for audio coding
US20060015329A1