Audio signal processing method and device

Through the encoding method of basic frames and extended frames, the problem of wasted bandwidth resources in audio signal transmission between multiple electronic devices is solved, and flexible encoding of different audio applications is realized to meet different audio quality and delay requirements.

CN114945981BActive Publication Date: 2025-08-08HUAWEI TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202080092744.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-06-24
Publication Date
2025-08-08
Estimated Expiration
2040-06-24

AI Technical Summary

Technical Problem

When transmitting audio signals between multiple electronic devices, different audio applications have different requirements for compression and encoding of audio signals, resulting in repeated transmission and waste of bandwidth resources.

Method used

The encoding method of basic frames and extended frames is adopted to perform different encoding processing on the audio signals. The basic frames are used for audio applications with low delays and average quality, and the extended frames are used for audio applications with long delays and high quality, and audio signals of different quality are restored through joint decoding.

Benefits of technology

It effectively avoids repeated transmission of the same audio signal on the encoding side, reduces bandwidth resource waste, meets the needs of different audio applications, and improves encoding rate and system efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114945981B_ABST
    Figure CN114945981B_ABST
Patent Text Reader

Abstract

The present application provides an audio signal processing method and device, which relates to the field of multimedia processing technology, and solves the problem of repeated transmission and bandwidth resource waste caused by different audio applications having different requirements for audio signal compression encoding when transmitting audio signals between multiple electronic devices in the prior art. The method includes: a first device samples and quantizes the acquired first audio signal to obtain a second audio signal; encodes the second audio signal in a first encoding method in units of a first time length to obtain a basic frame; encodes the second audio signal in a second encoding method in units of a second time length to obtain an extended frame, wherein the second time length is greater than the first time length, and the first encoding method and the second encoding method respectively encode different signals carried in the second audio signal, and / or respectively encode the second audio signal to different encoding degrees; and sends the basic frame and the extended frame to the second device.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of multimedia processing technology, and in particular to an audio signal processing method and device. Background Art

[0002] With the increasing use of electronic devices, collaborative audio signal processing between multiple electronic devices will become a key technological development trend in future audio signal processing. When transmitting audio signals between multiple electronic devices, the transmitting electronic device can sample, quantize, and encode the collected audio signals before compressing and transmitting them to the receiving electronic device. However, the various applications on the receiving electronic device may have different requirements for the latency and quality of the audio signals, requiring the transmitting electronic device to perform different compression and encoding processes on the audio signals.

[0003] like Figure 1 A possible application scenario is shown, where the mobile phone sends the collected audio signal to a smart headset. The smart headset has different audio applications. For example, audio application 1 is a voice enhancement application that has high real-time requirements for the received audio signal and general transmission quality requirements for the audio signal; audio application 2 is a three-dimensional sound field acquisition application that has high transmission quality requirements for the received audio signal, but low requirements for the audio signal delay. According to the processing method of the existing technology, the mobile phone needs to perform different compression encoding processes on the same audio signal and transmit multiple audio signals to the smart headset. The transmission delay and quality of different audio signals are different, but the content of different audio signals is the same audio signal collected by the mobile phone. Therefore, it will cause repeated transmission of audio signals, resulting in the occupation and waste of bandwidth resources. Summary of the Invention

[0004] The present application provides an audio signal processing method and apparatus, which solves the problem of repeated transmission and bandwidth resource waste caused by different audio applications having different requirements for audio signal compression encoding when transmitting audio signals between multiple electronic devices in the prior art.

[0005] In a first aspect, a method for processing an audio signal is provided, the method comprising: performing sampling and quantization processing on an acquired first audio signal by a first device to obtain a second audio signal; encoding the second audio signal using a first encoding method in units of a first duration to obtain a basic frame, and encoding the second audio signal using a second encoding method in units of a second duration to obtain an extended frame, wherein the second duration is greater than the first duration, and the first encoding method and the second encoding method respectively encode different signals carried in the second audio signal, and / or respectively encode the second audio signal to different degrees of encoding; and sending the basic frame and the extended frame to a second device.

[0006] In the above technical solution, the transmitting end of the audio signal can encode and compress the same audio signal to obtain two encoded frames of different frame lengths, including a basic frame and an extended frame. The extended frame can be obtained by re-encoding the portion of the basic frame that does not encode the second audio signal, or by re-encoding the portion of the basic frame that is not encoded precisely. The receiving end can then decode the basic frame to obtain one audio signal, and jointly decode the basic frame and the extended frame to obtain another audio signal. The two recovered audio signals have different time delays and audio quality, thereby meeting the needs of the above-mentioned different audio applications, avoiding the problem of repeated transmission of the same audio signal after encoding on the encoding side, which wastes bandwidth resources, and reducing system overhead.

[0007] In a possible design, the second duration is N times the first duration, where N is a natural number greater than or equal to 2.

[0008] In the possible implementation described above, when the first device encodes the second audio signal, the interval between basic frames is a first duration, and the interval between extended frames is N times the first duration. That is, for every N basic frames encoded, one extended frame is encoded. This allows the encoding side to obtain encoded frames with different delays, which are used by the decoding side to recover audio signals with different delays based on the encoded frames with different delays. This improves the encoding rate, resolves bandwidth waste, and reduces system overhead.

[0009] In one possible design, encoding the second audio signal using a first encoding method in units of a first duration to obtain basic frames specifically includes: downsampling the second audio signal to obtain a low-frequency signal carried in the second audio signal; and encoding the low-frequency signal using a time domain encoding method to obtain multiple basic frames with the first duration as a frame length.

[0010] In the first possible implementation described above, the encoding side can encode the low-frequency signal included in the second audio signal using time-domain coding to obtain basic frames. Because time-domain coding can encode audio signals into digital signals with low latency, it is suitable for encoding basic frames with low latency that only include the low-frequency portion of the original audio signal. The decoding side can then recover an audio signal with high real-time performance and average audio quality from these basic frames for application in corresponding audio applications.

[0011] In one possible design, encoding the second audio signal using a second encoding method in units of a second time length to obtain an extended frame specifically includes: performing a frequency domain transform on the second audio signal to obtain frequency domain coefficients corresponding to the second audio signal; grouping multiple frequency domain coefficients of a high-frequency part of the frequency domain coefficients corresponding to the second audio signal in an average order from low frequency to high frequency to obtain group envelope values of multiple high-frequency groups, wherein the group envelope value is an average value of the multiple high-frequency frequency domain coefficients in each group; and encoding according to the group envelope values to obtain multiple extended frames with the second time length as a frame length.

[0012] In the first possible implementation method described above, corresponding to the basic frame obtained by the above encoding, the encoding side can also encode the high-frequency signal included in the second audio signal according to the frequency domain encoding method to obtain an extended frame to encode the high-frequency signal part of the basic frame that is not encoded. Therefore, the decoding side can recover an audio signal with weaker real-time performance but better audio quality including the low-frequency and high-frequency parts of the original audio signal based on the above basic frame combined with the extended frame, so as to be applied to the corresponding audio application. The above implementation method can meet the needs of various audio applications, improve the encoding rate, and solve the problem of bandwidth resource waste through basic frame encoding and extended frame encoding.

[0013] In one possible design, encoding the second audio signal using a first encoding method in units of a first time length to obtain basic frames specifically includes: performing a frequency domain transform on the second audio signal to obtain multiple frequency domain coefficients of a low-frequency signal and multiple frequency domain coefficients of a high-frequency signal corresponding to the second audio signal; grouping the multiple frequency domain coefficients of the high-frequency signal in order from low frequency to high frequency to obtain group envelope values of multiple high-frequency groups, wherein the group envelope value is an average value of the multiple high-frequency frequency domain coefficients in each group; and encoding the multiple frequency domain coefficients of the low-frequency signal and the group envelope value of the high-frequency signal to obtain multiple basic frames with the first time length as a frame length.

[0014] In the second possible implementation, the encoding side can encode the low-frequency signal and high-frequency signal included in the second audio signal using frequency domain coding, where multiple frequency domain coefficients of the low-frequency signal are encoded, while only the group envelope value of the high-frequency signal is encoded for the high-frequency signal, to obtain a basic frame. This basic frame encoding method encodes the low-frequency portion with high quality while encoding the high-frequency portion with lower quality. The decoding side can recover an audio signal with strong real-time performance and general audio quality based on the basic frame for application in corresponding audio applications.

[0015] In one possible design, the second audio signal is encoded using a second encoding method in units of a second time length to obtain an extended frame, specifically including: encoding the difference between multiple frequency domain coefficients of the high-frequency signal and the corresponding group envelope value in units of the second time length to obtain multiple extended frames with the second time length as the frame length.

[0016] In the second possible implementation method described above, the encoding side can further encode the high-frequency portion of the signal with lower encoding quality in the basic frame based on the basic frame of the second method described above, that is, it can encode the difference between multiple frequency domain coefficients of the high-frequency signal and the corresponding group envelope value. This extended encoding method further encodes the high-frequency portion with high quality. Therefore, the decoding side can recover an audio signal with average real-time performance and high audio quality based on the combined decoding of the basic frame and the extended frame, so as to be applied to the corresponding audio application. The above implementation method obtains encoded frames with different delays and encoding qualities through basic frame encoding and extended frame encoding, thereby improving the encoding rate and reducing system overhead.

[0017] In addition, there is a possible implementation method three, in which the encoding side can obtain a basic frame according to the time domain encoding method of the above-mentioned method one, and obtain a first extended frame according to the encoding method of the extended frame in the above-mentioned method one, and then obtain a second extended frame according to the encoding method of the extended frame in the above-mentioned method two. Through this encoding method, a basic frame with strong real-time performance, containing only low-frequency signals, and low encoding quality can be obtained; a first extended frame with strong real-time performance, containing low-frequency and high-frequency signals but high-frequency signals, and low encoding quality can be obtained; and a second extended frame with weak real-time performance, containing low-frequency and high-frequency signals, and high-frequency signal encoding quality can be obtained. As a result, the layers of the encoded frames are richer, and the decoding side can perform joint decoding and recovery based on the above-mentioned basic frame, the first extended frame, and the second extended frame to obtain audio signals of different qualities, so as to meet the needs of different audio applications, improve the flexibility and encoding rate of audio encoding, and reduce system overhead.

[0018] In one possible design, encoding the second audio signal using a second encoding method in units of a second time length to obtain an extended frame specifically includes: performing a frequency domain transform on the second audio signal to obtain multiple frequency domain coefficients of a low-frequency signal and multiple frequency domain coefficients of a high-frequency signal corresponding to the second audio signal; averaging and grouping the multiple frequency domain coefficients of the low-frequency signal and the multiple frequency domain coefficients of the high-frequency signal in order from low frequency to high frequency to obtain corresponding group envelope values, wherein the group envelope value is an average of the multiple frequency domain coefficients in each group; and encoding according to the group envelope values to obtain multiple extended frames with the second time length as a frame length.

[0019] In the fourth possible implementation, corresponding to the basic frame obtained by encoding in the first embodiment, the encoding side may encode the group envelope values of the low-frequency frequency-domain coefficients and the group envelope values of the high-frequency frequency-domain coefficients corresponding to the second audio signal using frequency-domain encoding to obtain an extended frame. Therefore, even if the basic frame is lost, the decoding side can still decode the extended frame to recover the audio signal, thereby improving the reliability of audio encoding transmission and enhancing the user experience.

[0020] In a possible design, performing frequency domain transformation on the second audio signal specifically includes: obtaining MDCT frequency domain component coefficients corresponding to the second audio signal according to an improved discrete cosine transform (MDCT) algorithm.

[0021] In a second aspect, a method for processing an audio signal is provided, the method comprising: a second device receiving a basic frame and an extended frame sent from a first device, wherein a frame length of the extended frame is greater than a frame length of the basic frame, and the extended frame is obtained by re-encoding audio signals corresponding to multiple basic frames; decoding the basic frame to obtain a basic audio signal; or jointly decoding the basic frame and the extended frame to obtain an extended audio signal.

[0022] In a possible design, decoding the basic frame to obtain the basic audio signal specifically includes: decoding the basic frame according to a time domain coding and decoding method to obtain the basic audio signal.

[0023] In one possible design, a basic frame and an extended frame are jointly decoded to obtain an extended audio signal, specifically including: if the extended frame includes group envelope values of multiple high-frequency signals, obtaining multiple frequency domain coefficients of the high-frequency signal based on the group envelope values of the multiple high-frequency signals, where the frequency domain coefficients of the high-frequency signal are the group envelope values corresponding to the frequency domain coefficients; upsampling the basic audio signal to obtain a third audio signal; performing a frequency domain transform on the third audio signal frame by frame to obtain multiple frequency domain coefficients of a low-frequency signal corresponding to the third audio signal; and performing an inverse frequency domain transform based on the multiple frequency domain coefficients of the high-frequency signal and the multiple frequency domain coefficients of the low-frequency signal to obtain the extended audio signal.

[0024] In one possible design, a basic frame is decoded to obtain a basic audio signal, specifically including: if the basic frame includes multiple frequency domain coefficients of a low-frequency signal and multiple group envelope values of a high-frequency signal, then multiple frequency domain coefficients of the low-frequency signal and multiple frequency domain coefficients of the high-frequency signal are obtained according to the basic frame, wherein the multiple frequency domain coefficients of the high-frequency signal are group envelope values corresponding to the frequency domain coefficients; and frequency domain inverse transform is performed according to the multiple frequency domain coefficients of the low-frequency signal and the multiple frequency domain coefficients of the high-frequency signal to obtain the basic audio signal.

[0025] In one possible design, the basic frame and the extended frame are jointly decoded to obtain an extended audio signal, specifically including: if the extended frame includes the difference between multiple frequency domain coefficients of the high-frequency signal and the corresponding group envelope values, then the multiple frequency domain coefficients of the high-frequency signal are obtained according to the multiple group envelope values of the high-frequency signal and the difference between the multiple frequency domain coefficients of the high-frequency signal and the corresponding group envelope values; and the frequency domain inverse transform is performed according to the multiple frequency domain coefficients of the low-frequency signal and the multiple frequency domain coefficients of the high-frequency signal to obtain the extended audio signal.

[0026] In one possible design, a basic frame and an extended frame are jointly decoded to obtain an extended audio signal, specifically including: if the extended frame includes multiple group envelope values of a low-frequency signal and multiple group envelope values of a high-frequency signal, then multiple frequency domain coefficients of the low-frequency signal are obtained based on the multiple group envelope values of the low-frequency signal, and multiple frequency domain coefficients of the high-frequency signal are obtained based on the multiple group envelope values of the high-frequency signal; wherein the multiple frequency domain coefficients of the low-frequency signal are determined by performing a frequency domain transformation on the basic audio signal obtained from the basic frame, or the multiple frequency domain coefficients of the low-frequency signal are determined based on the multiple group envelope values of the low-frequency signal in the extended frame, and the multiple frequency domain coefficients of the low-frequency signal are the group envelope values corresponding to the frequency domain coefficients; and performing an inverse frequency domain transformation on the multiple frequency domain coefficients of the low-frequency signal and the multiple frequency domain coefficients of the high-frequency signal to obtain the extended audio signal.

[0027] In a possible design, frequency domain inverse transformation is performed based on the frequency domain coefficients, specifically including: obtaining an audio analog signal corresponding to the frequency domain coefficients based on an improved inverse discrete cosine transform algorithm.

[0028] In a possible design, the group envelope value includes an average value of multiple frequency domain coefficients in each group obtained by grouping the multiple frequency domain coefficients in order from low frequency to high frequency.

[0029] According to a third aspect, an audio signal processing device is provided, comprising: a preprocessing module for sampling and quantizing an acquired first audio signal to obtain a second audio signal; an encoding module for encoding the second audio signal using a first encoding method in units of a first duration to obtain basic frames, and encoding the second audio signal using a second encoding method in units of a second duration to obtain extended frames, wherein the second duration is greater than the first duration, and the first encoding method and the second encoding method respectively encode different signals carried in the second audio signal and / or respectively encode the second audio signal to different degrees of encoding; and a sending module for sending the basic frames and the extended frames to a second device.

[0030] In a possible design, the second duration is N times the first duration, where N is a natural number greater than or equal to 2.

[0031] In one possible design, the encoding module is specifically used to: downsample the second audio signal to obtain a low-frequency signal carried in the second audio signal; encode the low-frequency signal according to a time domain coding method to obtain multiple basic frames with the first time length as the frame length.

[0032] In one possible design, the encoding module is specifically used to: perform a frequency domain transform on the second audio signal to obtain frequency domain coefficients corresponding to the second audio signal; group multiple frequency domain coefficients of the high-frequency part of the frequency domain coefficients corresponding to the second audio signal in an average order from low frequency to high frequency to obtain group envelope values of multiple high-frequency groups, wherein the group envelope value is the average value of the multiple high-frequency frequency domain coefficients in each group; and encode according to the group envelope value to obtain multiple extended frames with a second time length as a frame length.

[0033] In one possible design, the encoding module is specifically used to: perform frequency domain transformation on the second audio signal to obtain multiple frequency domain coefficients of the low-frequency signal and multiple frequency domain coefficients of the high-frequency signal corresponding to the second audio signal; group the multiple frequency domain coefficients of the high-frequency signal in order from low frequency to high frequency to obtain group envelope values of multiple high-frequency groups, wherein the group envelope value is the average value of the multiple high-frequency frequency domain coefficients in each group; and encode the multiple frequency domain coefficients of the low-frequency signal and the group envelope value of the high-frequency signal to obtain multiple basic frames with a first time length as the frame length.

[0034] In one possible design, the encoding module is specifically used to encode the differences between multiple frequency domain coefficients of the high-frequency signal and the corresponding group envelope values in units of the second time length to obtain multiple extended frames with the second time length as the frame length.

[0035] In one possible design, the encoding module is specifically used to: perform a frequency domain transform on the second audio signal to obtain multiple frequency domain coefficients of a low-frequency signal and multiple frequency domain coefficients of a high-frequency signal corresponding to the second audio signal; averagely group the multiple frequency domain coefficients of the low-frequency signal and the multiple frequency domain coefficients of the high-frequency signal in order from low frequency to high frequency to obtain corresponding group envelope values, wherein the group envelope value is the average value of the multiple frequency domain coefficients in each group; and encode according to the group envelope values to obtain multiple extended frames with a frame length of the second time length.

[0036] In a possible design, the frequency domain transformation specifically includes: improving the discrete cosine transform (MDCT) algorithm.

[0037] In a fourth aspect, an audio signal processing device is provided, comprising: a receiving module for receiving basic frames and extended frames sent from a first device, wherein the frame length of the extended frame is greater than the frame length of the basic frame, and the extended frame is obtained by re-encoding audio signals corresponding to multiple basic frames; a decoding module for decoding the basic frame to obtain a basic audio signal; or, jointly decoding the basic frame and the extended frame to obtain an extended audio signal.

[0038] In a possible design, the decoding module is specifically configured to decode the basic frame according to a time domain coding and decoding method to obtain a basic audio signal.

[0039] In one possible design, the decoding module is specifically used to: if the extended frame includes group envelope values of multiple high-frequency signals, obtain multiple frequency domain coefficients of the high-frequency signal based on the group envelope values of the multiple high-frequency signals, and the frequency domain coefficients of the high-frequency signal are the group envelope values corresponding to the frequency domain coefficients; upsample the basic audio signal to obtain a third audio signal; perform frequency domain transformation on the third audio signal frame by frame to obtain multiple frequency domain coefficients of a low-frequency signal corresponding to the third audio signal; and perform frequency domain inverse transformation based on the multiple frequency domain coefficients of the high-frequency signal and the multiple frequency domain coefficients of the low-frequency signal to obtain the extended audio signal.

[0040] In one possible design, the decoding module is specifically used to: if the basic frame includes multiple frequency domain coefficients of the low-frequency signal and multiple group envelope values of the high-frequency signal, then obtain multiple frequency domain coefficients of the low-frequency signal and multiple frequency domain coefficients of the high-frequency signal according to the basic frame, wherein the multiple frequency domain coefficients of the high-frequency signal are the group envelope values corresponding to the frequency domain coefficients; perform frequency domain inverse transform according to the multiple frequency domain coefficients of the low-frequency signal and the multiple frequency domain coefficients of the high-frequency signal to obtain the basic audio signal.

[0041] In one possible design, the decoding module is specifically used to: if the extended frame includes the difference between multiple frequency domain coefficients of the high-frequency signal and the corresponding group envelope values, then obtain multiple frequency domain coefficients of the high-frequency signal based on the multiple group envelope values of the high-frequency signal and the difference between the multiple frequency domain coefficients of the high-frequency signal and the corresponding group envelope values; perform frequency domain inverse transform based on the multiple frequency domain coefficients of the low-frequency signal and the multiple frequency domain coefficients of the high-frequency signal to obtain the extended audio signal.

[0042] In one possible design, the decoding module is specifically used to: if the extended frame includes multiple group envelope values of the low-frequency signal and multiple group envelope values of the high-frequency signal, then obtain multiple frequency domain coefficients of the low-frequency signal according to the multiple group envelope values of the low-frequency signal, and obtain multiple frequency domain coefficients of the high-frequency signal according to the multiple group envelope values of the high-frequency signal; wherein the multiple frequency domain coefficients of the low-frequency signal are determined by performing a frequency domain transformation on a basic audio signal obtained from the basic frame, or the frequency domain coefficients of the multiple low-frequency signals are determined based on the multiple group envelope values of the low-frequency signal in the extended frame, and the multiple frequency domain coefficients of the low-frequency signal are the group envelope values corresponding to the frequency domain coefficients; perform an inverse frequency domain transformation on the multiple frequency domain coefficients of the low-frequency signal and the multiple frequency domain coefficients of the high-frequency signal to obtain the extended audio signal.

[0043] In a possible design approach, the frequency domain inverse transformation specifically includes: improving the inverse discrete cosine transform algorithm.

[0044] In a possible design, the group envelope value includes an average value of multiple frequency domain coefficients in each group obtained by grouping the multiple frequency domain coefficients in order from low frequency to high frequency.

[0045] In a fifth aspect, an electronic device is provided, comprising: a processor and a transmission interface; a memory for storing instructions executable by the processor; wherein the processor is configured to execute the instructions so that the electronic device implements the audio signal processing method as described in the first aspect and any one of the first aspects above.

[0046] In a sixth aspect, an electronic device is provided, comprising: a processor and a transmission interface; a memory for storing instructions executable by the processor; wherein the processor is configured to execute the instructions so that the electronic device implements the audio signal processing method as described in the second aspect and any one of the second aspects above.

[0047] In a seventh aspect, a computer-readable storage medium is provided. When the instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to perform the audio signal processing method as described in the first aspect and any one of the first aspects above.

[0048] In an eighth aspect, a computer program product is provided. When the computer program product is run on a computer, the computer is caused to execute the audio signal processing method as described in any one of the first and second aspects above.

[0049] In a ninth aspect, a computer-readable storage medium is provided. When the instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to perform the audio signal processing method as described in the second aspect and any one of the second aspects above.

[0050] In a tenth aspect, a computer program product is provided. When the computer program product is run on a computer, the computer is caused to execute the audio signal processing method as described in any one of the second aspect and the second aspect above.

[0051] It can be understood that any of the audio signal processing devices, electronic devices, computer-readable storage media, and computer program products provided above can be used to execute the corresponding methods provided above. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects in the corresponding methods provided above, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] Figure 1 A schematic diagram of an application scenario of an audio signal processing method provided in an embodiment of the present application;

[0053] Figure 2 A flowchart of an audio signal processing method provided in an embodiment of the present application;

[0054] Figure 3 A schematic diagram of a processing process of an audio signal processing method provided in an embodiment of the present application;

[0055] Figure 4 A schematic diagram of an audio signal encoding frame provided in an embodiment of the present application;

[0056] Figure 5 A flowchart of another audio signal processing method provided in an embodiment of the present application;

[0057] Figure 6 A schematic diagram of an audio signal processing device provided in an embodiment of the present application;

[0058] Figure 7 A schematic diagram of another audio signal processing device provided in an embodiment of the present application;

[0059] Figure 8 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0060] In the following, the terms "first" and "second" are used for descriptive purposes only and should not be understood to indicate or imply relative importance or implicitly specify the number of the technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the features. In the description of this embodiment, unless otherwise specified, "plurality" means two or more.

[0061] It should be noted that, in this application, words such as "exemplary" or "for example" are used to indicate examples, illustrations, or descriptions. Any embodiment or design described in this application as "exemplary" or "for example" should not be construed as being preferred or advantageous over other embodiments or designs. Rather, the use of words such as "exemplary" or "for example" is intended to present the relevant concepts in a concrete manner.

[0062] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0063] First, a brief introduction to the implementation environment and application scenarios of the embodiments of the present application is given.

[0064] The present application provides an audio signal processing method and apparatus that can be applied to the transmission of audio signals between multiple electronic devices. This method can flexibly encode and decode audio signals based on basic and extended frames, addressing the varying audio signal processing requirements of different applications. This allows for audio processing with varying latency or quality requirements. This solves the existing problem of repeated transmission and wasted bandwidth resources caused by different audio applications requiring different real-time transmission and quality when transmitting the same audio signal between multiple electronic devices.

[0065] like Figure 1 As shown, the audio signal processing method provided in the embodiments of the present application can be applied to an electronic device capable of audio signal processing, and includes at least two electronic devices, and data can be transmitted between the two electronic devices. For example, the audio signal can be transmitted via a wired network, a wireless local area network, near field communication (NFC), or Bluetooth.

[0066] Specifically, the electronic device may be a mobile phone, smart speaker, smart headset, tablet computer, desktop computer, laptop computer, handheld computer, notebook computer, vehicle-mounted device, ultra-mobile personal computer (UMPC), netbook, cellular phone, personal digital assistant (PDA), augmented reality (AR) or virtual reality (VR) device, etc. The embodiments of the present disclosure do not impose any special restrictions on the specific form of the electronic device. For example, Figure 1As shown, electronic device 1 may be a mobile phone, and electronic device 2 may be a smart headset.

[0067] The embodiment of the present application provides an audio signal processing method, which is applied to a first device and a second device. Figure 2 As shown, the method may include:

[0068] S201: The first device samples and quantizes the acquired first audio signal to obtain a second audio signal.

[0069] The first audio signal may be an audio signal collected by the first device, or may be an audio signal stored locally in the first device or from other devices or equipment.

[0070] If the first device responds to the audio request of the second device and needs to send the first audio signal to the second device, it is necessary to sample and quantize the first audio signal to obtain a digital signal to save transmission bandwidth. The basic processing process can be referred to Figure 3 As shown in FIG, the first audio signal is sampled and quantized to obtain the second audio signal s(n), where n corresponds to different audio sampling points, arranged in time sequence. If the audio signal is sampled based on a frequency of 16 kHz, it means that 16×10 samples are sampled per second. 3 sampling points, then the time interval between each two sampling points is 0.0625ms.

[0071] Next, the quantized values corresponding to the audio signal sampling points are encoded into binary digital signals and can be transmitted. The quantized values of the sampling points can be represented by different quantization precisions, such as 16 bits, 24 bits, or 32 bits.

[0072] S202: The first device encodes the second audio signal frame by frame using a first encoding method in units of a first duration to obtain basic frames, and encodes the second audio signal frame by frame using a second encoding method in units of a second duration to obtain extended frames.

[0073] The second duration is greater than the first duration, and therefore the frame length of the extended frame is greater than the frame length of the basic frame.

[0074] During encoding and compression, a second audio signal of a fixed duration can be used as an interval. After each frame of the second audio signal is collected and quantized, the second audio signal can be compressed and encoded, and then transmitted frame by frame after encoding. In this application, the second audio signal is encoded according to different time intervals, that is, different frame lengths, to generate two or more encoding frames, including basic frames and extended frames.

[0075] It should be noted that, based on the above-mentioned encoding principles and sampling rates of audio signals, it can be known that, relative to the original audio signals in nature, the current audio encoding technology can only be infinitely close to the original audio signals. That is, the encoding and decoding rules of the audio signals determine that the digital encoding and decoding methods will have a certain degree of distortion on the audio signals and cannot completely restore the original audio signals. The encoding method involved in this application is a lossy encoding technology.

[0076] Therefore, the basic frame or extended frame in the embodiment of the present application may encode only a portion of the first audio signal, not the entire first audio signal. Specifically, the extended frame may be obtained by re-encoding the second audio signal segments corresponding to multiple basic frames. The extended frame may further encode the audio signal that was not encoded or was encoded with insufficient accuracy in the basic frame.

[0077] Specifically, the first encoding method and the second encoding method can respectively encode different signals carried in the second audio signal. For example, the low-frequency signal portion carried in the second audio signal is encoded using the first encoding method to obtain a base frame, and the high-frequency signal portion carried in the second audio signal is encoded using the second encoding method to obtain an extended frame.

[0078] Alternatively, the first encoding method and the second encoding method may each encode the second audio signal to different degrees of encoding quality, resulting in encoded frames of lower and higher encoding quality, which are then transmitted to the decoding side for decoding. Thus, the decoding side can recover different audio signals based on the basic frames or extended frames. Compared to the original audio signal, the audio signal recovered based on the extended frames combined with the basic frames has less distortion and therefore has better encoding quality.

[0079] It can be seen that, generally speaking, the longer the frame length used to encode the second audio signal, the higher the compression rate of the first audio signal and the longer the signal transmission delay; at the same bit rate, the better the encoding quality of the audio signal. The encoding quality of an audio signal refers to the degree to which the audio signal recovered after decoding is restored relative to the original audio signal before encoding and compression. In other words, the longer the frame length used to encode the second audio signal, the higher the degree of signal restoration and lower the distortion rate of the audio signal obtained after decoding relative to the original audio signal.

[0080] In an embodiment of the present application, a basic frame may be a low-latency and / or low-quality encoding of the current second audio signal. The first device may transmit the basic frames individually to the second device frame by frame. In this way, after receiving the basic frames frame by frame, the second device may decode the received basic frames according to a preset decoding method to obtain an audio signal, which can be applied to audio applications with low latency requirements or relatively low audio quality requirements.

[0081] The extended frame can be used to encode the current second audio signal with a higher delay and / or higher quality. The frame length of the extended frame is greater than the frame length of the basic frame. The extended frame encoding transmits enhanced information for multiple basic frame audio signals, and further encodes data that is not included in the basic frame or is incompletely encoded in the audio signal. In this way, after receiving the extended frame frame by frame, the second device side can jointly decode it with the basic frame to obtain an audio signal with higher audio quality, which can be applied to audio applications that do not require high real-time performance but have relatively high requirements for audio quality.

[0082] In one embodiment, the first device may encode the second audio signal in units of a first duration to obtain a basic frame; and the first device may encode the second audio signal in units of a second duration to obtain an extended frame. The second duration may be N times the first duration, where N is a natural number greater than or equal to 2. The first duration is the length of a basic frame, i.e., the time interval between two basic frames, and the second duration is the length of an extended frame, i.e., the time interval between two extended frames.

[0083] by Figure 4 For example, t1, t2, t3, t4, t5, t6, t7, and t8 represent basic frames of audio coding. The algorithmic delay of the basic frame is about Δt, that is, the time interval between two basic frames is Δt. T1 and T2 represent extended frames of audio coding. Figure 4 In this example, an extended frame is compressed every four basic frames. The algorithmic delay of the extended frame is ΔT, that is, the time interval between two extended frames is ΔT, where ΔT = 4 × Δt, or N = 4. The basic frame or extended frame contains digitized audio sample data.

[0084] For example, the delay Δt can be 0.5ms or 5ms. The delay Δt and ΔT depend on the design of the coding structure and the actual application requirements. For example, if the sampling frequency is 16kHz and the frame length of the basic frame is 5ms, the number of audio sampling points contained in each basic frame is 80.

[0085] S203: The first device sends the basic frame and the extended frame to the second device.

[0086] The first device can encode the basic frame and send it frame by frame to the second device. The first device can also encode the extended frame and send it frame by frame to the second device. This allows the second device to decode the basic frame or extended frame after receiving it and recover the audio signal for different audio applications.

[0087] According to the above encoding method provided in the embodiment of the present application, the second device receives a digital signal sent from the first device, the digital signal includes a basic frame or an extended frame, and the second device can decode it according to the preset encoding and decoding method to restore the audio signal. Figure 5 As shown, the specific process may include:

[0088] S501: A second device receives a basic frame and an extended frame sent from a first device, wherein the frame length of the extended frame is greater than that of the basic frame, and the extended frame is obtained by re-encoding audio signals corresponding to multiple basic frames.

[0089] S502: The second device decodes the basic frame to obtain a basic audio signal, or jointly decodes the basic frame and the extended frame to obtain an extended audio signal.

[0090] The second device decodes the received basic frame or extended frame according to a preset coding rule, that is, the second device decodes the digital signal to obtain an analog signal to meet the audio signal requirements of different audio applications on the second device.

[0091] Furthermore, after receiving the basic frame, the second device performs frame decoding based on the basic frame to obtain the corresponding basic audio signal s1(n). After receiving the extended frame, the second device performs comprehensive decoding based on the extended frame and the basic frame to obtain the corresponding extended audio signal s2(n).

[0092] The audio content of the basic audio signal s1(n) and the second audio signal s2(n) is the same, but the transmission delay and audio quality of the basic audio signal s1(n) and the extended audio signal s2(n) are different. The audio quality of the basic audio signal s1(n) is slightly worse than that of the extended audio signal s2(n), and the transmission delay of the basic audio signal s1(n) is lower than that of the extended audio signal s2(n).

[0093] Through the above-mentioned implementation of the present application, the same set of encoding schemes can be used to transmit audio applications with different delay requirements between the encoding side and the decoding side, that is, the encoding side only obtains one audio signal, but can encode basic frames and extended frames respectively according to different delay requirements, so that the decoding side can decode different audio signals based on these two encoded frames to meet the needs of different audio applications. Among them, the audio signal decoded based on the basic frame has a lower delay, but the audio signal quality is poor. The audio signal decoded based on the extended frame combined with the basic frame has a longer delay, but the audio signal quality is better, and the distortion of the original audio signal is small. Therefore, the decoding side can recover more than two audio signals based on different basic frames and extended frames, and only encode one audio signal during encoding. This encoding method reduces redundant information, avoids the problem of repeated transmission and bandwidth resource waste after encoding the same audio signal on the encoding side, and greatly reduces system overhead.

[0094] Next, the encoding and decoding methods and processes in the technical solution of the present application are described in detail by listing several preferred encoding and decoding implementation methods, such as Method 1, Method 2, Method 3, and Method 4. The following several implementation methods are not all possible implementation methods of the present application, but are only exemplary implementation methods.

[0095] Method 1:

[0096] 1. Encoding process on the encoding side:

[0097] In one possible implementation, the first device may use a time-domain encoding method with a relatively low latency to obtain a basic frame, that is, only encode the low-frequency portion of the second audio signal. The first device may use a frequency-domain encoding method with a relatively high latency to obtain an extended frame, and the extended frame only includes the high-frequency portion of the second audio signal.

[0098] For example, consider two different audio applications on a second device. One is for device calibration and positioning, requiring real-time audio signals with a transmission latency of no more than 1ms. However, the audio quality requirements are low, and the audio signal can contain only low-frequency signals without high-frequency signals. The other is for speech enhancement, requiring less real-time audio signals with a transmission latency of no more than 6ms. However, the audio quality requirements are high, requiring both high- and low-frequency signals.

[0099] In the above step S202, the encoding of the basic frame by the first device may specifically include:

[0100] (1) The first device downsamples the second audio signal to obtain a low-frequency signal included in the second audio signal.

[0101] Here, downsampling means sampling a sample value sequence once every several sample values to obtain a new sequence. For example, if the sampling rate of the first audio signal is 16kHz, the bandwidth of the second audio signal obtained by quantization can be half of the sampling rate, that is, the bandwidth can be 8kHz. If the second audio signal includes a frequency band of 0 to 8kHz, the low-frequency signal s L (n) is the 0~4kHz part, the high frequency signal s H (n) is the 4k~8kHz part. Then the second audio signal is downsampled by one time to obtain the low-frequency signal s included in the second audio signal. L (n) is the audio signal from 0 to 4 kHz.

[0102] (2) Encode the low-frequency signal in a first time unit according to a time domain coding method to obtain a plurality of basic frames.

[0103] Time-domain coding encodes the waveform of the audio signal. Typical examples of time-domain coding include the International Telecommunication Union (ITU)'s G.726, G.723.1, and G.728 coding standards. These standards widely utilize code-excited linear prediction (CEL) technology, which is modeled after the human vocal tract and utilizes the inherent characteristics of the glottis and vocal tract to remove redundant information from the audio signal. This significantly reduces the required audio encoding bit rate while maintaining high audio quality.

[0104] For example, the first device may L (n) Encode using G.726 encoding, assemble into basic frames with the first duration as the interval, and the frame length of the basic frame is the first duration. For example, the first duration can be 0.5ms, and the s of each 0.5ms duration are sequentially L (n) signal is encoded, and the resulting digital signal is a basic frame. Among them, G.726 is a voice codec algorithm that can encode audio signals into digital signals with low latency.

[0105] Furthermore, in the above-mentioned step S202, the encoding of the extended frame by the first device may specifically include:

[0106] (1) Performing a frequency domain transformation on the second audio signal in units of a second time length to obtain frequency domain coefficients corresponding to the second audio signal.

[0107] Frequency-domain coding utilizes the human ear's perception of sound to encode audio signals in the frequency domain. It prioritizes encoding frequency bands of interest to humans, while using coarse or no quantization for bands masked by other bands or difficult for humans to perceive. The advantage of frequency-domain coding is that it removes some redundancy based on the characteristics of the human ear. Therefore, it offers comparable encoding quality for various audio signals. For signals like music, the encoding quality is particularly superior to time-domain coding.

[0108] Specifically, a modified discrete cosine transform (MDCT) may be performed on the second audio signal to obtain MDCT frequency domain coefficients corresponding to the second audio signal. MDCT transform is an algorithm that transforms a signal from the time domain to the frequency domain, and the obtained coefficients represent the frequency domain components of each frequency point.

[0109] The transformation formula for transforming the time domain signal s(n) to the MDCT frequency domain coefficient S(k) is as follows:

[0110] The MDCT coefficient S(k) is obtained, where S(k) is the frequency domain part of the second audio signal.

[0111] For example, if the second duration is 5ms, that is, the frame length for encoding the extended frame is 5ms, and the sampling rate is 16kHz, then s(n) includes 80 sampling points, that is, N=80, and the value range of the sampling point n is 0 to 79. MDCT transform is performed on each 5ms duration s(n) signal one by one to obtain the corresponding MDCT coefficient, and the value range of k can be 0 to 79. The frequency domain coefficient k starts from 0, representing from low frequency to high frequency. The low-frequency frequency domain coefficients are S(0) to S(39) from low to high, and the high-frequency frequency domain coefficients are S(40) to S(79) from low to high.

[0112] (2) The frequency domain coefficients of the high frequency part of the frequency domain coefficients corresponding to the second audio signal are evenly grouped in order from low frequency to high frequency, and group envelope values of the multiple high frequency groups are obtained, and encoded according to the envelope coding method.

[0113] For example, the 40 high-frequency frequency domain coefficients S(40) to S(79) are evenly divided into 8 groups, each of which includes five high-frequency frequency domain coefficients. The specific groupings are as follows:

[0114] Group 1 contains high-frequency domain coefficients: S(40) to S(44);

[0115] Group 2 contains high-frequency domain coefficients: S(45) to S(49);

[0116] Group 3 contains high-frequency domain coefficients: S(50) to S(54);

[0117] Group 4 contains high-frequency frequency domain coefficients: S(55) to S(59);

[0118] Group 5 contains high-frequency domain coefficients: S(69) to S(64);

[0119] Group 6 contains high-frequency frequency domain coefficients: S(65) to S(69);

[0120] Group 7 contains high-frequency domain coefficients: S(70) to S(74);

[0121] Group 8 contains high frequency domain coefficients: S(75) to S(79).

[0122] Next, group envelope values for the plurality of high-frequency groups are obtained, where the group envelope value is an average of the plurality of high-frequency frequency-domain coefficients in each group. The first device may obtain the group envelope value for each group of the high-frequency portion of the second audio signal, and then encode the group envelope value to obtain a plurality of extended frames having a second duration as a frame length.

[0123] Exemplarily, the calculation of the group envelope value may be:

[0124] Group 1 envelope value: S HE (0)=[S(40)+S(41)+S(42)+S(43)+S(44)] / 5;

[0125] Group 2 envelope value: S HE (1)=[S(45)+S(46)+S(47)+S(48)+S(49)] / 5;

[0126] Group 3 envelope value: S HE (2)=[S(50)+S(51)+S(52)+S(53)+S(54)] / 5;

[0127] Group 4 envelope value: S HE (3)=[S(55)+S(56)+S(57)+S(58)+S(59)] / 5;

[0128] Group 5 envelope value: S HE (4)=[S(60)+S(61)+S(62)+S(63)+S(64)] / 5;

[0129] Group 6 envelope value: S HE (5)=[S(65)+S(66)+S(67)+S(68)+S(69)] / 5;

[0130] Group 7 envelope value: S HE(6)=[S(70)+S(71)+S(72)+S(73)+S(74)] / 5;

[0131] Group 8 envelope value: S HE (7)=[S(75)+S(76)+S(77)+S(78)+S(79)] / 5.

[0132] With the second time length as the frame length, the first device can digitally encode the group envelope values of the plurality of high frequency groups obtained above and send them frame by frame to the second device. For example, every 5 ms, the first device will digitally encode the S HE (0)~S HE (7) Encode and assemble into an extended frame and send it to the second device.

[0133] 2. Decoding process on the decoding side:

[0134] Based on the above encoding method, the second device receives a basic frame at regular intervals, and then decodes the basic frame according to the time domain decoding method to obtain a first audio signal. The first audio signal only contains a low-frequency part relative to the original audio signal on the encoding side.

[0135] The second device receives an extended frame at regular intervals. The extended frame contains only the high-frequency portion of the original audio signal. The second device decodes the extended frame and the basic frame to obtain a second audio signal. The second audio signal includes both the low-frequency portion and the high-frequency portion.

[0136] Taking the above embodiment as an example, the second device can receive a basic frame every 0.5 ms and then decode the basic frame according to the G.726 decoding method to obtain the basic audio signal s1(n). This basic audio signal s1(n) only has a low-frequency component, but has a low latency of 0.5 ms. Therefore, this audio signal can be used in audio applications with low latency requirements, such as device calibration and positioning.

[0137] If the extended frame received by the second device includes multiple group envelope values of the high-frequency signal, multiple high-frequency frequency domain coefficients of the high-frequency signal are obtained based on the multiple group envelope values of the high-frequency signal, that is, the frequency domain coefficients of the high-frequency signal are the group envelope values corresponding to the high-frequency frequency domain coefficients. In addition, the basic audio signal is upsampled to obtain a third audio signal, and the third audio signal is frequency-domain transformed frame by frame to obtain multiple low-frequency frequency domain coefficients of the low-frequency signal corresponding to the third audio signal. The audio signal recovered by the second device based on the multiple high-frequency frequency domain coefficients and the multiple low-frequency frequency domain coefficients is then the extended audio signal.

[0138] For example, the second device may receive an extended frame every 5 ms, and obtain the group envelope value S of the high frequency part of the audio signal from the extended frame. HE (0)~SHE (7). Then, multiple high-frequency frequency domain coefficients can be obtained according to the group envelope value, that is, the high-frequency frequency domain coefficient of the audio signal is equal to the group envelope value of the corresponding high-frequency frequency domain coefficient group, that is:

[0139] S(40)=S(41)=S(42)=S(43)=S(44)=S HE (0);

[0140] S(45)=S(46)=S(47)=S(48)=S(49)=S HE (1);

[0141] S(50)=S(51)=S(52)=S(53)=S(54)=S HE (2);

[0142] S(55)=S(56)=S(57)=S(58)=S(59)=S HE (3);

[0143] S(60)=S(61)=S(62)=S(63)=S(64)=S HE (4);

[0144] S(65)=S(66)=S(67)=S(68)=S(69)=S HE (5);

[0145] S(70)=S(71)=S(72)=S(73)=S(74)=S HE (6);

[0146] S(75)=S(76)=S(77)=S(78)=S(79)=S HE (7), we can get S(40) to S(79).

[0147] Take the audio signal restored from the basic frames received in the second time segment, for example, the audio signal s1(n) obtained by decoding multiple basic frames within the above 5ms, and upsample the audio signal s1(n) to obtain the third audio signal s′ L (n). The upsampling process is to insert one or more zero points between two adjacent points in the original signal. For example, after upsampling the above audio signal s1(n), a third audio signal s′ with an 8k bandwidth and a sampling rate of 16kHz can be obtained. L (n), but the third audio signal s′ L The high frequency part of (n) is still 0.

[0148] For the low-frequency audio signal s′ L(n) Perform MDCT transformation and obtain the frequency domain coefficient S′ according to the following formula L (k):

[0149]

[0150] Among them, corresponding to a 5ms delay, the audio signal segment with a sampling rate of 16kHz has 80 sampling points, that is, N=80 in the above formula. L The low-frequency coefficients of (k) are integrated with the high-frequency coefficients S(40) to S(79) obtained from the extended frame in the above steps to obtain the complete MDCT coefficients S(k) of the audio frame. L (k), k=0~39.

[0151] By performing an inverse transform of the improved discrete cosine transform on S(k), an extended audio signal s2(n) can be obtained. The extended audio signal s2(n) includes both high-frequency components and low-frequency components. The specific formula for the inverse transform of the improved discrete cosine transform is as follows:

[0152]

[0153] In the audio signal decoded according to the above embodiment, the audio signal s1(n) decoded from the basic frame contains only low-frequency components, resulting in lower decoding quality. However, this audio signal has a lower latency and can be used in audio services that require less audio quality but lower latency. The audio signal s2(n) decoded from the extended frame and the basic frame contains both high- and low-frequency components, resulting in higher decoding quality, but with a longer latency. Therefore, it can be used in audio services that require higher audio quality but lower real-time audio transmission requirements.

[0154] The above-mentioned implementation mode of the present application transmits one audio application through the same set of encoding and decoding solutions, and the different audio signals obtained by decoding can be applied to different audio applications respectively, thereby avoiding repeated encoding and decoding and transmission processes, which can greatly avoid the waste of bandwidth resources and reduce system overhead.

[0155] Furthermore, when encoding and decoding according to the above embodiment, if a basic frame received by the decoding device is lost or not received, and the audio signal cannot be recovered based on the basic frame decoding, the decoding device can decode based on the extended frame. When performing an inverse frequency domain transform, the low-frequency frequency domain coefficients are 0, and an inverse frequency domain transform is performed only based on the frequency domain coefficients of the high-frequency portion, thereby recovering the audio signal. The audio signal only contains the high-frequency portion.

[0156] Method 2:

[0157] 1. Encoding process on the encoding side:

[0158] In one possible implementation, the first device may use a time-domain encoding method with a relatively low latency to obtain a basic frame, that is, only encode the low-frequency portion of the second audio signal. The first device may use a frequency-domain encoding method with a relatively high latency to obtain an extended frame, and the extended frame only includes the high-frequency portion of the second audio signal.

[0159] For example, the second device has two different audio applications. One is a speech enhancement application that requires a real-time audio signal with a low signal latency of no more than 6ms, and both high and low frequency components. The other is a three-dimensional (3D) sound field acquisition application that requires high audio signal quality and can tolerate longer signal latency.

[0160] In the above step S202, the encoding of the basic frame by the first device may specifically include:

[0161] (1) The first device performs frequency domain transformation on the second audio signal with the first time length as the frame length to obtain frequency domain coefficients, that is, obtains multiple low-frequency frequency domain coefficients of the low-frequency signal and multiple high-frequency frequency domain coefficients of the high-frequency signal corresponding to the second audio signal.

[0162] (2) The multiple frequency domain coefficients of the high-frequency signal are grouped in order from low frequency to high frequency to obtain group envelope values of the multiple high-frequency groups, wherein the group envelope value is the average value of the multiple high-frequency frequency domain coefficients in each group.

[0163] (3) Encoding is performed according to the multiple frequency domain coefficients of the low-frequency signal and the group envelope value of the high-frequency signal to obtain multiple basic frames with the first time length as the frame length.

[0164] For example, to meet the real-time requirements of the speech enhancement application on the second device, the first duration can be 5ms. If the sampling rate is 16kHz, the first device can perform MDCT transformation on the audio signal s(n) of every 5ms to obtain the MDCT coefficient S(k), where the value range of k can be 0 to 79. The high-frequency frequency domain coefficients S(40) to S(79) are divided into 8 groups in order, each group including 5 high-frequency frequency domain coefficients, and the group envelope values S of the multiple high-frequency groups are obtained. HE (0)~S HE (7) The first device combines the multiple frequency domain coefficients S(0) to S(39) of the low frequency signal and the group envelope value S of the high frequency signal. HE (0)~S HE (7) Encode to obtain a basic frame.

[0165] Furthermore, in the above-mentioned step S202, the encoding of the extended frame by the first device may specifically include:

[0166] The first device encodes the differences between the multiple frequency domain coefficients of the high-frequency signal and the corresponding group envelope values in units of the second time length to obtain multiple extended frames with the second time length as the frame length.

[0167] Exemplarily, the first device may calculate the difference between each high-frequency frequency domain coefficient of the high-frequency part after the basic frame is encoded and the group envelope value of the corresponding high-frequency group every 20ms. Specifically, the group envelope coefficient difference SD may be obtained by subtracting the multiple high-frequency frequency domain coefficients from the group envelope value corresponding to the high-frequency frequency domain coefficient. HE (k), where k = 40 to 79. The calculation method can be as follows: SD HE (40) = S(40) - S HE (0);

[0168] SD HE (41) = S (41) - S HE (0);

[0169]

[0170] SD HE (45) = S(45) - S HE (1);

[0171] SD HE (46) = S(45) - S HE (1);

[0172]

[0173] SD HE (78) = S (78) - S HE (7)

[0174] SD HE (79) = S (79) - S HE (7).

[0175] The first device can collect the envelope coefficient differences SD of these groups every 20ms. HE (40)~SD HE (79) assemble into an extended frame and transmit it to the second device. The first device can use these groups of envelope coefficient differences SD HE (40)~SD HE (79) It can be directly encapsulated for transmission, or it can be encoded and transmitted using differential quantization.

[0176] 2. Decoding process on the decoding side:

[0177] Based on the above encoding method, the second device receives a basic frame every first time length. If the basic frame includes multiple frequency domain coefficients of the low-frequency signal and multiple group envelope values of the high-frequency signal, the second device obtains multiple frequency domain coefficients of the high-frequency signal according to the multiple group envelope values of the high-frequency signal in the basic frame, and then performs an inverse frequency domain transform according to the multiple frequency domain coefficients of the low-frequency signal and the multiple frequency domain coefficients of the high-frequency signal to obtain the first audio signal.

[0178] The second device receives an extended frame every second time period. If the extended frame includes differences between multiple frequency domain coefficients of the high-frequency signal and corresponding group envelope values, the second device may combine the group envelope values of the high-frequency signal in the basic frame to obtain multiple frequency domain coefficients of the high-frequency signal, and then perform an inverse frequency domain transform based on the multiple frequency domain coefficients of the low-frequency signal and the multiple frequency domain coefficients of the high-frequency signal to obtain a second audio signal. The second audio signal includes both a low-frequency component and a high-frequency component.

[0179] Taking the above embodiment as an example, the second device can receive a basic frame every 5 ms. The second device first obtains the low-frequency frequency domain coefficients of S(k) based on the basic frame, that is, obtains S(0) to S(39). The second device then obtains the high-frequency coefficients based on the high-frequency group envelope value in the basic frame, that is, each high-frequency frequency domain coefficient can be equal to its corresponding group envelope value, that is:

[0180] S(40)=S(41)=S(42)=S(43)=S(44)=S HE (0);

[0181] S(45)=S(46)=S(47)=S(48)=S(49)=S HE (1);

[0182] S(50)=S(51)=S(52)=S(53)=S(54)=S HE (2);

[0183] S(55)=S(56)=S(57)=S(58)=S(59)=S HE (3);

[0184] S(60)=S(61)=S(62)=S(63)=S(64)=S HE (4);

[0185] S(65)=S(66)=S(67)=S(68)=S(69)=S HE (5);

[0186] S(70)=S(71)=S(72)=S(73)=S(74)=S HE (6);

[0187] S(75)=S(76)=S(77)=S(78)=S(79)=S HE (7), we get S(40) to S(79).

[0188] The low-frequency frequency domain coefficients S(0) to S(39) obtained by decoding the basic frame and the high-frequency frequency domain coefficients S(40) to S(79) with defects in the high-frequency portion are combined. The obtained S(0) to S(79) are subjected to an inverse MDCT transform to obtain the basic audio signal s1(n). This basic audio signal s1(n) has a low time delay and includes both the high-frequency portion and the low-frequency portion of the original audio signal. However, since the high-frequency portion is only a high-frequency signal restored using the group envelope value, that is, the values of multiple frequency bands are the same, the signal quality of the high-frequency portion is slightly poor, which is equivalent to reducing the frequency domain resolution of the high-frequency portion.

[0189] The second device can receive the extended frame every 20ms, and obtains the group envelope coefficient difference SD of the high frequency part of the audio signal from the extended frame. HE (40)~SD HE (79). Then according to SD HE (40)~SD HE (79) The frequency domain coefficients of the high frequency part in each basic frame are obtained, that is, each high frequency domain coefficient is obtained by adding the group envelope coefficient difference to the spectrum envelope as shown below:

[0190] S(40)=SD HE (40)+S HE (0);

[0191] S(41)=SD HE (41)+S HE (0);

[0192]

[0193] S(45)=SD HE (45)+S HE (1);

[0194] S(46)=SD HE (46)+S HE (1);

[0195]

[0196] S(78)=SD HE (78)+S HE (7);

[0197] S(79)=SD HE (79)+S HE (7), that is, the complete high-frequency part S(40) to S(79) of the spectrum can be obtained.

[0198] Based on the frequency domain coefficients S(0) to S(39) of the low-frequency part obtained by decoding the basic frame, an inverse MDCT transform is performed on the obtained S(0) to S(79) to obtain an extended audio signal s2(n). The extended audio signal s2(n) includes both the high-frequency part and the low-frequency part of the original audio signal, and the high-frequency part is a high-frequency signal restored by combining the group envelope value with the group envelope coefficient difference. Therefore, the restoration quality of the extended audio signal s2(n) is higher than that of the basic audio signal s1(n), but the time delay of the extended audio signal s2(n) is longer. In terms of the real-time performance of signal transmission, the basic audio signal s1(n) is superior to the extended audio signal s2(n).

[0199] Method 3:

[0200] 1. Encoding process on the encoding side:

[0201] In a possible implementation, when the first device needs to meet three or more different audio application requirements on the second device, the first device may encode a basic frame and two or more extended frames.

[0202] Specifically, the first device may employ a relatively low-latency, low-quality time-domain encoding method to obtain a basic frame, i.e., only encode the low-frequency portion of the second audio signal. The first device may employ a relatively high-latency, low-quality frequency-domain encoding method to obtain a first extended frame, wherein the first extended frame only encodes the frequency-domain group envelope value of the high-frequency portion of the second audio signal. The first device may employ a relatively high-latency, high-quality frequency-domain encoding method to obtain a second extended frame, wherein the second extended frame includes the high-frequency portion of the second audio signal.

[0203] For example, there are three different audio applications on the second device. One is a device calibration and positioning application. The requirement for processing audio signals is strong real-time performance. The signal sending delay interval needs to be no more than 1ms. The audio signal can only contain low-frequency signals but not high-frequency signals. The second is a speech enhancement application. The requirement for processing audio signals by this application is strong real-time performance. The signal sending delay does not exceed 6ms. The audio quality requirement is high. Both high-frequency and low-frequency signals in the audio signal are required. The third is a 3D sound field acquisition application. The real-time requirements for processing audio signals by this application are not high, but the audio quality requirements are high.

[0204] In the above step S202, the encoding of the basic frame by the first device may specifically refer to the encoding method of the basic frame in the above method 1, which may include:

[0205] (1) downsampling the second audio signal to obtain a low-frequency signal included in the second audio signal;

[0206] (2) Encode the low-frequency signal according to a time domain coding method to obtain a plurality of basic frames with the first time length as the frame length.

[0207] For example, the first device may L (n) Encoding is performed using the G.726 encoding method, and basic frames are assembled with the first duration as the interval. For example, the first duration may be 0.5 ms, which meets the requirements of the first audio application mentioned above.

[0208] Furthermore, in the above step S202, the encoding of the first extended frame by the first device may refer to the encoding process of the extended frame in the above method 1, including:

[0209] (1) performing a frequency domain transform on the second audio signal using the second time length as the frame length to obtain a frequency domain coefficient corresponding to the second audio signal;

[0210] (2) The frequency domain coefficients of the high frequency part of the frequency domain coefficients corresponding to the second audio signal are evenly grouped in order from low frequency to high frequency, and group envelope values of the multiple high frequency groups are obtained, and encoded according to the envelope coding method.

[0211] For example, the first device can perform MDCT transformation on s(n) to obtain MDCT frequency domain coefficients. For example, if the frame length is 5ms, the sampling rate is 16kHz, and s(n) includes 80 sampling points, S(0) to S(79) can be obtained. The 40 high-frequency component coefficients S(40) to S(79) are evenly divided into 8 groups, each high-frequency group has five high-frequency component coefficients, and the group envelope value S of each high-frequency group is obtained. HE (0)~S HE (7), wherein the group envelope value is the average value of the multiple high frequency domain coefficients in each group. The first device can obtain the group envelope values S of the multiple high frequency groups obtained above. HE (0)~S HE (7) Perform digital encoding, and every 5ms the first device will convert the S HE (0)~S HE (7) Encode and assemble into an extended frame and send it to the second device.

[0212] In combination with the above, the encoding of the second extended frame in the above step S202 may refer to the encoding process of the extended frame in the above method 2, including:

[0213] The first device encodes the differences between the multiple frequency domain coefficients of the high-frequency signal and the corresponding group envelope values in units of the third time to obtain multiple extended frames with the third time as the frame length.

[0214] Exemplarily, the first device may calculate the difference between each high-frequency frequency domain coefficient of the high-frequency part after encoding the first extended frame and the group envelope value of the corresponding high-frequency group every 20 ms. Specifically, the plurality of high-frequency frequency domain coefficients may be subtracted from the group envelope value corresponding to the high-frequency frequency domain coefficient to obtain the group envelope coefficient difference SD of the plurality of high-frequency frequency domain coefficients. HE (40)~SD HE (79) Then the first device can calculate the envelope coefficient difference SD of these groups every 20ms. HE (40)~SD HE (79) Assemble into a second extended frame and transmit to the second device.

[0215] 2. Decoding process on the decoding side:

[0216] Based on the above encoding method, the second device receives a basic frame every first time length, and then decodes the basic frame according to the time domain decoding method to obtain a basic audio signal, which only contains a low-frequency part relative to the original audio signal on the encoding side.

[0217] The second device receives a first extended frame every second time period. If the first extended frame includes multiple group envelope values of high-frequency signals, the second device obtains multiple frequency domain coefficients of the high-frequency signal based on the multiple group envelope values of the high-frequency signals, where the frequency domain coefficients of the high-frequency signal are the group envelope values corresponding to the frequency domain coefficients. At the same time, the first audio signal obtained by decoding the basic frame is upsampled to obtain a third audio signal. The third audio signal is frequency-domain transformed frame by frame to obtain multiple frequency domain coefficients of the low-frequency signal corresponding to the third audio signal. Then, an inverse frequency domain transform is performed based on the multiple frequency domain coefficients of the high-frequency signal and the multiple frequency domain coefficients of the low-frequency signal to obtain a first extended audio signal. The first extended audio signal includes a low-frequency signal and a high-frequency signal, but the high-frequency quality is slightly weaker and the delay of the first extended audio signal is longer. Therefore, it can be used for the application of the second audio service mentioned above.

[0218] The second device receives a second extended frame every third time interval. If the second extended frame includes differences between multiple frequency domain coefficients of the high frequency signal and corresponding group envelope values, the second device may combine the group envelope values of the high frequency signal in the first extended frame to obtain multiple frequency domain coefficients of the high frequency signal, and then perform an inverse frequency domain transform based on the multiple frequency domain coefficients of the low frequency signal and the multiple frequency domain coefficients of the high frequency signal to obtain a second extended audio signal. The second extended audio signal includes both a low frequency portion and a high frequency portion.

[0219] For example, in conjunction with the above embodiment, the second device can receive a basic frame every 0.5 ms and then decode the basic frame according to the G.726 decoding method to obtain a basic audio signal s1(n). This basic audio signal s1(n) only has a low-frequency component, but has a low latency of 0.5 ms. Therefore, this audio signal can be used in audio applications with low latency requirements, such as the aforementioned device calibration and positioning applications.

[0220] The second device can receive a first extended frame every 5 ms, and obtain the group envelope value S of the high frequency part of the audio signal from the first extended frame. HE (0)~S HE (7), the second device can obtain multiple high-frequency frequency domain coefficients S(40) to S(79) according to the group envelope value. The second device decodes the multiple basic frames received within 5ms to obtain the audio signal s L (n) Perform upsampling to obtain the audio signal s′ L (n), for s′ L (n) is subjected to MDCT transformation to obtain low-frequency frequency domain coefficients S(0) to S(39). S(0) to S(79) are subjected to inverse MDCT transformation to obtain a first extended audio signal s2(n). The first extended audio signal s2(n) includes both high-frequency and low-frequency components, wherein the high-frequency component has a slightly weaker quality.

[0221] The second device can receive a second extended frame every 20ms, and obtain the group envelope coefficient difference SD of the high frequency part of the audio signal from the second extended frame. HE (40)~SD HE (79). Then according to SD HE (40)~SD HE (79), combined with the group envelope value S of the high frequency part of the audio signal obtained in the first extended frame HE (0)~S HE (7), and obtain the frequency domain coefficients S(40) to S(79) of each high-frequency part. Performing an inverse MDCT transform on S(0) to S(79) yields the second extended audio signal s3(n) of the 20ms time period. The second extended audio signal s3(n) includes both the high-frequency part and the low-frequency part. The high-frequency part of the second extended audio signal s3(n) is slightly better than that of the first extended audio signal s2(n).

[0222] Through the above implementation methods, the present application provides more possible audio coding structures, which can be applicable to three or more audio applications with different requirements, thereby saving transmission bandwidth and improving system performance.

[0223] Method 4:

[0224] 1. Encoding process on the encoding side:

[0225] In one possible implementation, the first device may use a relatively low-latency, low-quality time-domain coding method to obtain a base frame, that is, only encode the low-frequency portion of the second audio signal. The first device may use a relatively high-latency, low-quality frequency-domain coding method to obtain an extended frame, only encoding the frequency-domain group envelope values of the low-frequency portion and the frequency-domain group envelope values of the high-frequency portion of the second audio signal.

[0226] In the above step S202, the encoding of the basic frame by the first device may specifically refer to the encoding method of the basic frame in the above method 1, which may include:

[0227] (1) Down-sampling the second audio signal to obtain a low-frequency signal included in the second audio signal.

[0228] (2) Encode the low-frequency signal according to a time domain coding method to obtain a plurality of basic frames with the first time length as the frame length.

[0229] For example, the first device may L (n) Encoding is performed using the G.726 encoding method, and basic frames are assembled with a first duration as an interval. For example, the first duration may be 0.5 ms.

[0230] Furthermore, in the above step S202, the first device may encode the extended frame by referring to the encoding process of the extended frame in the above method 1, including:

[0231] (1) Performing a frequency domain transformation on the second audio signal in units of a second time length to obtain frequency domain coefficients corresponding to the second audio signal.

[0232] (2) A plurality of frequency domain coefficients of a high-frequency portion of the frequency domain coefficients corresponding to the second audio signal are averagely grouped in order from low frequency to high frequency to obtain group envelope values of a plurality of high-frequency groups, and a plurality of frequency domain coefficients of a low-frequency portion are averagely grouped in order from low frequency to high frequency to obtain group envelope values of a plurality of low-frequency groups, and the signals are encoded according to the envelope coding method.

[0233] For example, the first device can perform MDCT transformation on s(n) to obtain MDCT frequency domain coefficients. For example, if the frame length is 5ms, the sampling rate is 16kHz, and s(n) includes 80 sampling points, S(0) to S(79) can be obtained. The 40 low-frequency component coefficients S(0) to S(39) are evenly divided into 8 groups, each high-frequency group has five low-frequency component coefficients, and the group envelope value S of each low-frequency group is obtained. LE (0)~S LE(7). In addition, the 40 high-frequency component coefficients S(40) to S(79) are evenly divided into 8 groups, each high-frequency group has five high-frequency component coefficients, and the group envelope value S of each high-frequency group is obtained. HE (0)~S HE (7), wherein the group envelope value is the average value of the multiple high-frequency frequency domain coefficients in each group. The first device can obtain the group envelope values S of the multiple low-frequency groups obtained above. LE (0)~S LE (7) Digital encoding is performed, and the group envelope values S of multiple high-frequency groups are HE (0)~S HE (7) Perform digital encoding, and every 5ms the first device will convert the S LE (0)~S LE (7) and S HE (0)~S HE (7) Encode and assemble into an extended frame and send it to the second device.

[0234] 2. Decoding process on the decoding side:

[0235] Based on the above encoding method, the second device receives a basic frame every first time length, and then decodes the basic frame according to the time domain decoding method to obtain a basic audio signal. The first audio signal only contains a low-frequency part relative to the original audio signal on the encoding side.

[0236] The second device receives an extended frame every second time period. If the extended frame includes multiple group envelope values of the low-frequency signal and multiple group envelope values of the high-frequency signal, multiple frequency domain coefficients of the low-frequency signal are obtained based on the multiple group envelope values of the low-frequency signal, and multiple frequency domain coefficients of the high-frequency signal are obtained based on the multiple group envelope values of the high-frequency signal. If the second device receives multiple basic frames normally, the multiple frequency domain coefficients of the low-frequency signal can be determined by performing a frequency domain transform on the first audio signal obtained from the basic frames. If the second device does not receive multiple basic frames normally, the second device can determine the multiple frequency domain coefficients of the low-frequency signal based on the multiple group envelope values of the low-frequency signal in the extended frame, wherein the frequency domain coefficients of the multiple low-frequency signals are the group envelope values corresponding to the frequency domain coefficients. The second device can perform an inverse frequency domain transform on the multiple frequency domain coefficients of the low-frequency signal and the multiple frequency domain coefficients of the high-frequency signal to obtain the extended audio signal.

[0237] For example, if the second device receives basic frames normally, for example, one basic frame every 0.5 ms, and then decodes the basic frames according to the G.726 decoding method to obtain the basic audio signal s1(n), the basic audio signal s1(n) only has a low-frequency portion, but has a low latency of 0.5 ms.

[0238] The second device can receive an extended frame every 5ms, and obtains the group envelope value S of the high frequency part of the audio signal from the extended frame. HE (0)~S HE (7), then multiple high-frequency frequency domain coefficients S(40) to S(79) can be obtained according to the group envelope value. The second device decodes the multiple extended frames received within 5ms to obtain the audio signal s L (n) Perform upsampling to obtain the audio signal s′ L (n), for s′ L (n) is transformed by MDCT to obtain low-frequency frequency domain coefficients S(0) to S(39). S(0) to S(79) are inversely transformed by MDCT to obtain the extended audio signal s2(n). The extended audio signal s2(n) includes both high-frequency and low-frequency components, with the high-frequency component having a slightly lower quality.

[0239] For example, if the second device does not receive the basic frame normally, for example, the basic frame is lost or the basic frame received is verified to be erroneous, the second device decodes the low-frequency portion of the extended frame and obtains the group envelope value S LE (0)~S LE (7) obtaining a plurality of low-frequency frequency domain coefficients S(0) to S(39), wherein the plurality of low-frequency frequency domain coefficients are equal to the group envelope value of the corresponding low-frequency frequency domain coefficient group. The second device obtains the group envelope value S of the high-frequency part according to the extended frame decoding. HE (0)~S HE (7) A plurality of high-frequency frequency domain coefficients S(40) to S(79) are obtained, wherein the plurality of high-frequency frequency domain coefficients are equal to the group envelope value of the corresponding high-frequency frequency domain coefficient group. The second device performs an inverse MDCT transform on S(0) to S(79) obtained by decoding the plurality of extended frames received within 5 ms, thereby obtaining an extended audio signal s2(n), which includes both a high-frequency portion and a low-frequency portion.

[0240] According to the above embodiment, when the basic frame cannot be decoded normally to restore the audio signal, the decoding side device can still perform decoding based on the extended frame to restore the entire audio signal.

[0241] In summary, the above-mentioned implementation method provided by this application can transmit one audio application through the same set of codec schemes. Different audio signals obtained by decoding the basic frame or extended frame can be applied to different audio applications respectively, thereby avoiding repeated encoding, decoding and transmission processes, greatly avoiding the waste of bandwidth resources, and reducing system overhead. In addition, when the basic frame on the decoding side is lost and the audio signal cannot be recovered based on the basic frame decoding, the decoding side device can decode based on the extended frame, further improving the reliability of audio transmission.

[0242] In another possible implementation, before the audio signal is encoded and decoded for transmission, the encoding side device may communicate with the decoding side device in advance based on the audio application's encoding requirements for the transmitted audio signal, and negotiate a specific encoding and decoding method. For example, based on the first audio application on the second device requiring a low-latency, low-quality audio signal, the second device sends an audio signal request message to the first device carrying the configuration information, which is used to indicate the encoding method corresponding to the audio signal request. Alternatively, when the first device sends an encoded frame to the second device, the encoding method of the encoded frame can be indicated by an agreed bit. For example, the first device sends a basic frame of an audio signal to the second device, and the basic frame includes two pre-configured bits, such as 01, which can represent encoding method two. It can be seen that the above-mentioned encoding and decoding configuration methods are only illustrative and are not limited to the above-mentioned two methods. The embodiments of the present application do not specifically limit this.

[0243] This application also provides an audio processing device, such as Figure 6 The device 600 may include a preprocessing module 601, an encoding module 602 and a sending module 603.

[0244] The pre-processing module 601 may be configured to perform sampling and quantization processing on the acquired first audio signal to obtain a second audio signal.

[0245] The encoding module 602 may be configured to encode the second audio signal using a first encoding method in units of a first duration to obtain basic frames, and to encode the second audio signal using a second encoding method in units of a second duration to obtain extended frames, wherein the second duration is greater than the first duration, and the first encoding method and the second encoding method respectively encode different signals carried in the second audio signal and / or respectively encode the second audio signal at different encoding levels.

[0246] The sending module 603 may be configured to send the basic frame and the extended frame to the second device.

[0247] In a possible design, the second duration is N times the first duration, where N is a natural number greater than or equal to 2.

[0248] In one possible design, the encoding module 602 may be specifically configured to: downsample the second audio signal to obtain a low-frequency signal carried in the second audio signal; and encode the low-frequency signal according to a time domain coding method to obtain a plurality of basic frames having a first duration as a frame length.

[0249] In one possible design, the encoding module 602 may be specifically configured to: perform a frequency domain transform on the second audio signal to obtain frequency domain coefficients corresponding to the second audio signal; group multiple frequency domain coefficients of a high-frequency portion of the frequency domain coefficients corresponding to the second audio signal in an average order from low frequency to high frequency to obtain group envelope values of multiple high-frequency groups, where the group envelope value is an average value of the multiple high-frequency frequency domain coefficients in each group; and perform encoding based on the group envelope value to obtain multiple extended frames having a frame length of the second duration.

[0250] In one possible design, the encoding module 602 can be specifically used to: perform frequency domain transformation on the second audio signal to obtain multiple frequency domain coefficients of the low-frequency signal and multiple frequency domain coefficients of the high-frequency signal corresponding to the second audio signal; group the multiple frequency domain coefficients of the high-frequency signal in order from low frequency to high frequency to obtain group envelope values of multiple high-frequency groups, wherein the group envelope value is the average value of the multiple high-frequency frequency domain coefficients in each group; and encode the multiple frequency domain coefficients of the low-frequency signal and the group envelope value of the high-frequency signal to obtain multiple basic frames with a first time length as a frame length.

[0251] In one possible design, the encoding module 602 can be specifically used to encode the differences between multiple frequency domain coefficients of the high-frequency signal and the corresponding group envelope values in units of the second time length to obtain multiple extended frames with the second time length as the frame length.

[0252] In one possible design, the encoding module 602 may be specifically configured to: perform a frequency domain transform on the second audio signal to obtain multiple frequency domain coefficients of a low-frequency signal and multiple frequency domain coefficients of a high-frequency signal corresponding to the second audio signal; group the multiple frequency domain coefficients of the low-frequency signal and the multiple frequency domain coefficients of the high-frequency signal to obtain corresponding group envelope values, where the group envelope value is an average value of the multiple frequency domain coefficients in each group; and perform encoding based on the group envelope values to obtain multiple extended frames having a frame length of the second duration.

[0253] In a possible design, the frequency domain transformation in the above embodiment may specifically be a modified discrete cosine transform (MDCT) algorithm.

[0254] The present application also provides an audio signal processing device, such as Figure 7 As shown, the device 700 includes a receiving module 701 and a decoding module 702 .

[0255] The receiving module 701 can be used to receive basic frames and extended frames sent from the first device, wherein the frame length of the extended frame is greater than the frame length of the basic frame, and the extended frame is obtained by re-encoding audio signals corresponding to multiple basic frames.

[0256] The decoding module 702 may be configured to decode the basic frame to obtain a basic audio signal; or to jointly decode the basic frame and the extended frame to obtain an extended audio signal.

[0257] In a possible design, the decoding module 702 may be specifically configured to decode the basic frame according to a time domain coding and decoding method to obtain a basic audio signal.

[0258] In one possible design, the decoding module 702 may be specifically configured to: if the extended frame includes group envelope values of multiple high-frequency signals, obtain multiple frequency domain coefficients of the high-frequency signal based on the group envelope values of the multiple high-frequency signals, where the frequency domain coefficients of the high-frequency signal are the group envelope values corresponding to the frequency domain coefficients; upsample the basic audio signal to obtain a third audio signal; perform frequency domain transformation on the third audio signal frame by frame to obtain multiple frequency domain coefficients of a low-frequency signal corresponding to the third audio signal; and perform inverse frequency domain transformation based on the multiple frequency domain coefficients of the high-frequency signal and the multiple frequency domain coefficients of the low-frequency signal to obtain the extended audio signal.

[0259] In one possible design, the decoding module 702 can be specifically used to: if the basic frame includes multiple frequency domain coefficients of the low-frequency signal and multiple group envelope values of the high-frequency signal, then obtain multiple frequency domain coefficients of the low-frequency signal and multiple frequency domain coefficients of the high-frequency signal according to the basic frame, wherein the multiple frequency domain coefficients of the high-frequency signal are the group envelope values corresponding to the frequency domain coefficients; perform frequency domain inverse transform based on the multiple frequency domain coefficients of the low-frequency signal and the multiple frequency domain coefficients of the high-frequency signal to obtain the basic audio signal.

[0260] In one possible design, the decoding module 702 can be specifically used to: if the extended frame includes the difference between multiple frequency domain coefficients of the high-frequency signal and the corresponding group envelope values, then obtain multiple frequency domain coefficients of the high-frequency signal based on the multiple group envelope values of the high-frequency signal and the difference between the multiple frequency domain coefficients of the high-frequency signal and the corresponding group envelope values; perform frequency domain inverse transformation based on the multiple frequency domain coefficients of the low-frequency signal and the multiple frequency domain coefficients of the high-frequency signal to obtain the extended audio signal.

[0261] In one possible design, the decoding module 702 can be specifically used to: if the extended frame includes multiple group envelope values of the low-frequency signal and multiple group envelope values of the high-frequency signal, then obtain multiple frequency domain coefficients of the low-frequency signal according to the multiple group envelope values of the low-frequency signal, and obtain multiple frequency domain coefficients of the high-frequency signal according to the multiple group envelope values of the high-frequency signal; wherein the multiple frequency domain coefficients of the low-frequency signal are determined by performing a frequency domain transformation on a basic audio signal obtained from the basic frame, or the frequency domain coefficients of the multiple low-frequency signals are determined based on the multiple group envelope values of the low-frequency signal in the extended frame, and the multiple frequency domain coefficients of the low-frequency signal are the group envelope values corresponding to the frequency domain coefficients; and perform an inverse frequency domain transformation on the multiple frequency domain coefficients of the low-frequency signal and the multiple frequency domain coefficients of the high-frequency signal to obtain the extended audio signal.

[0262] In a possible design, the frequency domain inverse transformation in the above embodiment may specifically be an improved inverse discrete cosine transform algorithm.

[0263] In a possible design, the group envelope value includes an average value of multiple frequency domain coefficients in each group obtained by grouping the multiple frequency domain coefficients in order from low frequency to high frequency.

[0264] It is understood that when the above-mentioned audio signal processing device is an electronic device, the above-mentioned sending module can be a transmitter, which can include an antenna and a radio frequency circuit, etc., and the preprocessing module, encoding module, and decoding module can be processors, such as baseband chips, etc. When the above-mentioned audio signal processing device is a component having the functions of the first device or the second device, the sending module can be a radio frequency unit, and the preprocessing module, encoding module, and decoding module can be processors. When the above-mentioned audio signal processing device is a chip system, the sending module can be the output interface of the chip system, and the preprocessing module, encoding module, and decoding module can be the processor of the chip system, such as a central processing unit (CPU).

[0265] It should be noted that the specific execution process and embodiments of the above-mentioned device 600 can refer to the steps and related descriptions executed by the first device in the above-mentioned method embodiment, and the specific execution process and embodiments of the above-mentioned device 700 can refer to the steps and related descriptions executed by the second device in the above-mentioned method embodiment. The technical problems solved and the technical effects brought about can also refer to the contents described in the above-mentioned embodiments, and will not be repeated here one by one.

[0266] In this embodiment, the audio signal processing device is presented in the form of various functional modules divided in an integrated manner. Here, "module" can refer to a specific circuit, a processor and memory that executes one or more software or firmware programs, an integrated logic circuit, and / or other devices that can provide the above functions. In a simple embodiment, those skilled in the art can imagine that the audio signal processing device can be implemented as follows Figure 8 The form shown.

[0267] Figure 8 This is a schematic diagram of the structure of an exemplary electronic device 800 shown in an embodiment of the present application. The electronic device 800 may be the first device or the second device in the above embodiment, and is used to execute the test method of the smart camera in the above embodiment. Figure 8 As shown, the electronic device 800 may include at least one processor 801 , a communication line 802 and a memory 803 .

[0268] The processor 801 may be a general-purpose central processing unit (CPU), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits.

[0269] The communication line 802 may include a path for transmitting information between the above components, and the communication line may be, for example, a bus.

[0270] The memory 803 can be a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM) or other types of dynamic storage devices that can store information and instructions, or an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disc storage, optical disc storage (including compressed optical disc, laser disc, optical disc, digital versatile disc, Blu-ray disc, etc.), a magnetic disk storage medium or other magnetic storage device, or any other medium that can be used to carry or store the desired program code in the form of an instruction or data structure and can be accessed by a computer, but is not limited thereto. The memory can be independent and connected to the processor via a communication line 802. The memory can also be integrated with the processor. The memory provided in the embodiment of the present application is generally a non-volatile memory. Among them, the memory 803 is used to store the computer program instructions involved in the scheme of the embodiment of the present application, and is controlled by the processor 801 to execute. The processor 801 is used to execute computer program instructions stored in the memory 803, thereby implementing the method provided in the embodiment of the present application.

[0271] Optionally, the computer program instructions in the embodiments of the present application may also be referred to as application code, which is not specifically limited in the embodiments of the present application.

[0272] In a specific implementation, as an embodiment, the processor 801 may include one or more CPUs, such as Figure 8 CPU0 and CPU1 in.

[0273] In a specific implementation, as an embodiment, the electronic device 800 may include multiple processors, such as Figure 8801 and processor 807 in FIG. These processors may be single-CPU processors or multi-CPU processors. A processor herein may refer to one or more devices, circuits, and / or processing cores for processing data (e.g., computer program instructions).

[0274] In a specific implementation, as an embodiment, the electronic device 800 may further include a communication interface 804. The electronic device may send and receive data or communicate with other devices or communication networks through the communication interface 804. The communication interface 804 may be, for example, an Ethernet interface, a radio access network interface (RAN), a wireless local area network interface (WLAN), or a USB interface.

[0275] In a specific implementation, as an embodiment, the electronic device 800 may further include an output device 805 and an input device 806. The output device 805 communicates with the processor 801 and can display information in a variety of ways. For example, the output device 805 can be a liquid crystal display (LCD), a light emitting diode (LED) display device, a cathode ray tube (CRT) display device, or a projector. The input device 806 communicates with the processor 801 and can receive user input in a variety of ways. For example, the input device 806 can be a mouse, a keyboard, a touch screen device, or a sensor device.

[0276] In a specific implementation, the electronic device 800 can be a desktop computer, a portable computer, a network server, a personal digital assistant (PDA), a mobile phone, a tablet computer, a wireless terminal device, an embedded device, a smart camera or a Figure 8 The embodiment of the present application does not limit the type of the electronic device 800. For example, to implement the method of the second device in the above embodiment, the electronic device 800 needs to be equipped with a smart camera.

[0277] In some embodiments, Figure 8 The processor 801 in the electronic device 800 can call the computer program instructions stored in the memory 803 to enable the electronic device 800 to execute the method in the above method embodiment.

[0278] For example, Figure 6 or Figure 7 The functions / implementation processes of each processing module in Figure 8The processor 801 in the embodiment calls the computer program instructions stored in the memory 803 to implement. For example, Figure 7 The functions / implementation processes of the pre-processing module 601 and the encoding module 602 can be realized by Figure 8 The processor 801 in the memory 803 calls the computer execution instructions stored therein to implement the above. Figure 7 The functions / implementation processes of the receiving module 701 and the decoding module 702 can be realized by Figure 8 The processor 801 in the memory 803 calls the computer execution instructions stored therein to implement the above.

[0279] In an exemplary embodiment, a computer-readable storage medium including instructions is also provided. The instructions can be executed by the processor 801 of the electronic device 800 to implement the smart camera testing method of the above embodiment. Therefore, the technical effects that can be achieved can be referred to the above method embodiment and will not be repeated here.

[0280] In the above embodiments, all or part of the embodiments may be implemented using software, hardware, firmware, or any combination thereof. When implemented using a software program, all or part of the embodiments may be implemented in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions according to the embodiments of the present application are generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device.

[0281] Those skilled in the art will readily appreciate other embodiments of the present invention after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present invention that follow the general principles of this application and include common knowledge or customary techniques in the art that are not disclosed herein.

[0282] Finally, it should be noted that the above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any changes or substitutions within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.

Claims

1. A method for processing an audio signal, characterized in that: The method comprises: The first device samples and quantizes the acquired first audio signal to obtain a second audio signal; Encoding the second audio signal in a first encoding mode using a first time length as a unit to obtain basic frames; Encoding the second audio signal using a second encoding method in units of a second duration to obtain an extended frame, comprising: performing a frequency domain transform on the second audio signal to obtain frequency domain coefficients corresponding to the second audio signal; grouping multiple frequency domain coefficients of a high-frequency portion of the frequency domain coefficients corresponding to the second audio signal in order from low frequency to high frequency to obtain group envelope values of multiple high-frequency groups, wherein the group envelope value is an average of multiple high-frequency frequency domain coefficients in each group; and encoding based on the group envelope values to obtain multiple extended frames having the second duration as a frame length; wherein the second duration is greater than the first duration, and the first encoding method and the second encoding method respectively encode different signals carried in the second audio signal, and / or respectively encode the second audio signal to different encoding degrees. The basic frame and the extended frame are sent to a second device.

2. The method according to claim 1, characterized in that The second duration is N times the first duration, where N is a natural number greater than or equal to 2.

3. The method according to claim 1 or 2, characterized in that The encoding of the second audio signal in a first encoding mode using the first time length as a unit to obtain a basic frame specifically includes: downsampling the second audio signal to obtain a low-frequency signal carried in the second audio signal; The low-frequency signal is encoded according to a time domain encoding method to obtain a plurality of basic frames with the first time length as a frame length.

4. The method according to claim 1 or 2, characterized in that Performing a frequency domain transform on the second audio signal specifically includes: According to the improved discrete cosine transform (MDCT) algorithm, MDCT frequency domain component coefficients corresponding to the second audio signal are obtained.

5. A method for processing an audio signal, characterized in that: The method comprises: The first device samples and quantizes the acquired first audio signal to obtain a second audio signal; Encoding the second audio signal in a first encoding mode using a first time length as a unit to obtain basic frames; Encoding the second audio signal using a second encoding method in units of a second duration to obtain extended frames, including: performing a frequency domain transform on the second audio signal to obtain multiple frequency domain coefficients of a low-frequency signal and multiple frequency domain coefficients of a high-frequency signal corresponding to the second audio signal; averaging and grouping the multiple frequency domain coefficients of the low-frequency signal and the multiple frequency domain coefficients of the high-frequency signal in order from low frequency to high frequency to obtain corresponding group envelope values, wherein the group envelope value is an average of the multiple frequency domain coefficients in each group; and encoding based on the group envelope values to obtain multiple extended frames having the second duration as a frame length; wherein the second duration is greater than the first duration, and the first encoding method and the second encoding method respectively encode different signals carried in the second audio signal, and / or respectively encode the second audio signal to different encoding degrees. The basic frame and the extended frame are sent to a second device.

6. The method according to claim 5, characterized in that The second duration is N times the first duration, where N is a natural number greater than or equal to 2.

7. The method according to claim 5 or 6, characterized in that The encoding of the second audio signal in a first encoding mode using the first time length as a unit to obtain a basic frame specifically includes: downsampling the second audio signal to obtain a low-frequency signal carried in the second audio signal; The low-frequency signal is encoded according to a time domain encoding method to obtain a plurality of basic frames with the first time length as a frame length.

8. The method according to claim 5 or 6, characterized in that Performing a frequency domain transform on the second audio signal specifically includes: According to the improved discrete cosine transform (MDCT) algorithm, MDCT frequency domain component coefficients corresponding to the second audio signal are obtained.

9. An audio signal processing method, characterized in that: The method comprises: The first device samples and quantizes the acquired first audio signal to obtain a second audio signal; The second audio signal is encoded using a first encoding method in units of a first time length to obtain basic frames, the second audio signal is frequency-domain transformed to obtain multiple frequency-domain coefficients of a low-frequency signal and multiple frequency-domain coefficients of a high-frequency signal corresponding to the second audio signal; the multiple frequency-domain coefficients of the high-frequency signal are averaged and grouped in order from low frequency to high frequency to obtain group envelope values of multiple high-frequency groups, wherein the group envelope value is an average value of the multiple high-frequency frequency-domain coefficients in each group; and the multiple frequency-domain coefficients of the low-frequency signal and the group envelope value of the high-frequency signal are encoded to obtain multiple basic frames having the first time length as a frame length. encoding the second audio signal using a second encoding method in units of a second duration to obtain an extended frame, wherein the second duration is greater than the first duration, and the first encoding method and the second encoding method respectively encode different signals carried in the second audio signal and / or respectively encode the second audio signal with different encoding degrees; The basic frame and the extended frame are sent to a second device.

10. The method according to claim 9, characterized in that The second duration is N times the first duration, where N is a natural number greater than or equal to 2.

11. The method according to claim 9, characterized in that The encoding of the second audio signal in a second encoding manner using the second time length as a unit to obtain an extended frame specifically includes: Taking the second time length as a unit, the differences between the multiple frequency domain coefficients of the high-frequency signal and the corresponding group envelope values are encoded to obtain the multiple extended frames with the second time length as the frame length.

12. The method according to any one of claims 9 to 11, characterized in that: Performing a frequency domain transform on the second audio signal specifically includes: According to the improved discrete cosine transform (MDCT) algorithm, MDCT frequency domain component coefficients corresponding to the second audio signal are obtained.

13. An audio signal processing method, characterized in that: The method comprises: The second device receives a basic frame and an extended frame sent from the first device, wherein the frame length of the extended frame is greater than the frame length of the basic frame, and the extended frame is obtained by re-encoding audio signals corresponding to multiple basic frames; Jointly decoding the basic frame and the extended frame to obtain an extended audio signal specifically includes: If the extended frame includes a plurality of group envelope values of the high-frequency signal, a plurality of frequency domain coefficients of the high-frequency signal are obtained according to the group envelope values of the plurality of high-frequency signals, and the frequency domain coefficients of the high-frequency signal are the group envelope values corresponding to the frequency domain coefficients; Upsampling the basic audio signal to obtain a third audio signal; performing a frequency domain transform on the third audio signal frame by frame to obtain a plurality of frequency domain coefficients of a low-frequency signal corresponding to the third audio signal; Performing frequency domain inverse transformation according to the multiple frequency domain coefficients of the high frequency signal and the multiple frequency domain coefficients of the low frequency signal to obtain the extended audio signal.

14. The method according to claim 13, characterized in that The method further comprises: The basic frame is decoded to obtain the basic audio signal.

15. The method according to claim 14, characterized in that The decoding of the basic frame to obtain the basic audio signal specifically includes: The basic frame is decoded according to a time domain coding and decoding method to obtain the basic audio signal.

16. The method according to claim 13, characterized in that The jointly decoding the basic frame and the extended frame to obtain the extended audio signal specifically includes: If the extended frame includes differences between a plurality of frequency domain coefficients of the high frequency signal and corresponding group envelope values, obtaining a plurality of frequency domain coefficients of the high frequency signal according to the plurality of group envelope values of the high frequency signal and the differences between the plurality of frequency domain coefficients of the high frequency signal and the corresponding group envelope values; Performing frequency domain inverse transformation according to the multiple frequency domain coefficients of the low-frequency signal and the multiple frequency domain coefficients of the high-frequency signal to obtain the extended audio signal.

17. The method according to any one of claims 13 to 16, characterized in that: Perform frequency domain inverse transform based on frequency domain coefficients, specifically including: An audio analog signal corresponding to the frequency domain coefficient is obtained according to an improved inverse discrete cosine transform algorithm.

18. The method according to any one of claims 13 to 16, characterized in that: The group of envelope values includes an average value of multiple frequency domain coefficients in each group obtained by grouping the multiple frequency domain coefficients in order from low frequency to high frequency.

19. An audio signal processing method, characterized in that: The method comprises: The second device receives a basic frame and an extended frame sent from the first device, wherein the frame length of the extended frame is greater than the frame length of the basic frame, and the extended frame is obtained by re-encoding audio signals corresponding to multiple basic frames; Jointly decoding the basic frame and the extended frame to obtain an extended audio signal specifically includes: If the extended frame includes a plurality of group envelope values of a low-frequency signal and a plurality of group envelope values of a high-frequency signal, obtaining a plurality of frequency domain coefficients of the low-frequency signal according to the plurality of group envelope values of the low-frequency signal, and obtaining a plurality of frequency domain coefficients of the high-frequency signal according to the plurality of group envelope values of the high-frequency signal; The multiple frequency domain coefficients of the low-frequency signal are determined by performing a frequency domain transformation on a basic audio signal obtained from the basic frame, or the multiple frequency domain coefficients of the low-frequency signal are determined based on multiple group envelope values of the low-frequency signal in the extended frame, and the multiple frequency domain coefficients of the low-frequency signal are group envelope values corresponding to the frequency domain coefficients. Performing frequency domain inverse transformation according to the multiple frequency domain coefficients of the low-frequency signal and the multiple frequency domain coefficients of the high-frequency signal to obtain the extended audio signal.

20. The method according to claim 19, characterized in that The method further comprises: The basic frame is decoded to obtain the basic audio signal.

21. The method according to claim 20, characterized in that The decoding of the basic frame to obtain the basic audio signal specifically includes: The basic frame is decoded according to a time domain coding and decoding method to obtain the basic audio signal.

22. The method according to any one of claims 19 to 21, characterized in that Perform frequency domain inverse transform based on frequency domain coefficients, specifically including: An audio analog signal corresponding to the frequency domain coefficient is obtained according to an improved inverse discrete cosine transform algorithm.

23. The method according to any one of claims 19 to 21, characterized in that The group of envelope values includes an average value of multiple frequency domain coefficients in each group obtained by grouping the multiple frequency domain coefficients in order from low frequency to high frequency.

24. An audio signal processing method, characterized in that: The method comprises: The second device receives a basic frame and an extended frame sent from the first device, wherein the frame length of the extended frame is greater than the frame length of the basic frame, and the extended frame is obtained by re-encoding audio signals corresponding to multiple basic frames; Decoding the basic frame to obtain a basic audio signal specifically includes: If the basic frame includes a plurality of frequency domain coefficients of a low-frequency signal and a plurality of group envelope values of a high-frequency signal, obtaining the plurality of frequency domain coefficients of the low-frequency signal and the plurality of frequency domain coefficients of the high-frequency signal according to the basic frame, wherein the plurality of frequency domain coefficients of the high-frequency signal are the group envelope values corresponding to the frequency domain coefficients; Performing an inverse frequency domain transform on the multiple frequency domain coefficients of the low-frequency signal and the multiple frequency domain coefficients of the high-frequency signal to obtain the basic audio signal.

25. The method according to claim 24, characterized in that The method further comprises: The basic frame and the extended frame are jointly decoded to obtain an extended audio signal.

26. The method according to claim 25, characterized in that The jointly decoding the basic frame and the extended frame to obtain the extended audio signal specifically includes: If the extended frame includes differences between a plurality of frequency domain coefficients of the high frequency signal and corresponding group envelope values, obtaining a plurality of frequency domain coefficients of the high frequency signal according to the plurality of group envelope values of the high frequency signal and the differences between the plurality of frequency domain coefficients of the high frequency signal and the corresponding group envelope values; Performing frequency domain inverse transformation according to the multiple frequency domain coefficients of the low-frequency signal and the multiple frequency domain coefficients of the high-frequency signal to obtain the extended audio signal.

27. The method according to any one of claims 24 to 26, characterized in that Perform frequency domain inverse transform based on frequency domain coefficients, specifically including: An audio analog signal corresponding to the frequency domain coefficient is obtained according to an improved inverse discrete cosine transform algorithm.

28. The method according to any one of claims 24 to 26, characterized in that The group of envelope values includes an average value of multiple frequency domain coefficients in each group obtained by grouping the multiple frequency domain coefficients in order from low frequency to high frequency.

29. An audio signal processing device, characterized in that: The device comprises: a preprocessing module, configured to perform sampling and quantization processing on the acquired first audio signal to obtain a second audio signal; an encoding module, configured to encode the second audio signal using a first encoding method in units of a first duration to obtain basic frames, and to encode the second audio signal using a second encoding method in units of a second duration to obtain extended frames, wherein the second duration is greater than the first duration, and the first encoding method and the second encoding method respectively encode different signals carried in the second audio signal and / or respectively encode the second audio signal with different encoding levels; a sending module, configured to send the basic frame and the extended frame to a second device; The encoding module is specifically used for: performing a frequency domain transform on the second audio signal to obtain a frequency domain coefficient corresponding to the second audio signal; Averagely grouping multiple frequency domain coefficients of a high frequency portion of the frequency domain coefficients corresponding to the second audio signal in order from low frequency to high frequency to obtain group envelope values of multiple high frequency groups, wherein the group envelope value is an average value of the multiple high frequency domain coefficients in each group; Encoding is performed according to the group of envelope values to obtain a plurality of the extended frames with the second duration as the frame length.

30. The device according to claim 29, characterized in that The second duration is N times the first duration, where N is a natural number greater than or equal to 2.

31. The device according to claim 29 or 30, characterized in that The encoding module is specifically used for: downsampling the second audio signal to obtain a low-frequency signal carried in the second audio signal; The low-frequency signal is encoded according to a time domain encoding method to obtain a plurality of basic frames with the first time length as a frame length.

32. The device according to claim 29 or 30, characterized in that The frequency domain transformation specifically includes: improving the discrete cosine transform MDCT algorithm.

33. An audio signal processing device, characterized in that: The device comprises: a preprocessing module, configured to perform sampling and quantization processing on the acquired first audio signal to obtain a second audio signal; an encoding module, configured to encode the second audio signal using a first encoding method in units of a first duration to obtain basic frames, and to encode the second audio signal using a second encoding method in units of a second duration to obtain extended frames, wherein the second duration is greater than the first duration, and the first encoding method and the second encoding method respectively encode different signals carried in the second audio signal and / or respectively encode the second audio signal with different encoding levels; a sending module, configured to send the basic frame and the extended frame to a second device; The encoding module is specifically used for: performing a frequency domain transform on the second audio signal to obtain a plurality of frequency domain coefficients of a low-frequency signal and a plurality of frequency domain coefficients of a high-frequency signal corresponding to the second audio signal; averaging and grouping the multiple frequency domain coefficients of the low-frequency signal and the multiple frequency domain coefficients of the high-frequency signal in order from low frequency to high frequency to obtain corresponding group envelope values, wherein the group envelope value is the average value of the multiple frequency domain coefficients in each group; Encoding is performed according to the group of envelope values to obtain a plurality of the extended frames with the second duration as the frame length.

34. The device according to claim 33, characterized in that The second duration is N times the first duration, where N is a natural number greater than or equal to 2.

35. The device according to claim 33 or 34, characterized in that The encoding module is specifically used for: downsampling the second audio signal to obtain a low-frequency signal carried in the second audio signal; The low-frequency signal is encoded according to a time domain encoding method to obtain a plurality of basic frames with the first time length as a frame length.

36. The device according to claim 33 or 34, characterized in that The frequency domain transformation specifically includes: improving the discrete cosine transform MDCT algorithm.

37. An audio signal processing device, characterized in that: The device comprises: a preprocessing module, configured to perform sampling and quantization processing on the acquired first audio signal to obtain a second audio signal; an encoding module, configured to encode the second audio signal using a first encoding method in units of a first duration to obtain basic frames, and to encode the second audio signal using a second encoding method in units of a second duration to obtain extended frames, wherein the second duration is greater than the first duration, and the first encoding method and the second encoding method respectively encode different signals carried in the second audio signal and / or respectively encode the second audio signal with different encoding levels; a sending module, configured to send the basic frame and the extended frame to a second device; The encoding module is specifically used for: performing a frequency domain transform on the second audio signal to obtain a plurality of frequency domain coefficients of a low-frequency signal and a plurality of frequency domain coefficients of a high-frequency signal corresponding to the second audio signal; Averagely grouping the multiple frequency domain coefficients of the high-frequency signal in order from low frequency to high frequency to obtain group envelope values of the multiple high-frequency groups, wherein the group envelope value is the average value of the multiple high-frequency frequency domain coefficients in each group; Encoding is performed according to the multiple frequency domain coefficients of the low frequency signal and the group of envelope values of the high frequency signal to obtain multiple basic frames with the first time length as the frame length.

38. The device according to claim 37, characterized in that The second duration is N times the first duration, where N is a natural number greater than or equal to 2.

39. The device according to claim 37, characterized in that The encoding module is specifically used for: Taking the second time length as a unit, the differences between the multiple frequency domain coefficients of the high-frequency signal and the corresponding group envelope values are encoded to obtain the multiple extended frames with the second time length as the frame length.

40. The device according to any one of claims 37 to 39, characterized in that The frequency domain transformation specifically includes: improving the discrete cosine transform MDCT algorithm.

41. An audio signal processing device, characterized in that The device comprises: a receiving module, configured to receive a basic frame and an extended frame sent from a first device, wherein the frame length of the extended frame is greater than the frame length of the basic frame, and the extended frame is obtained by re-encoding audio signals corresponding to multiple basic frames; A decoding module, configured to jointly decode the basic frame and the extended frame to obtain an extended audio signal; The decoding module is specifically used for: If the extended frame includes a plurality of group envelope values of the high-frequency signal, a plurality of frequency domain coefficients of the high-frequency signal are obtained according to the group envelope values of the plurality of high-frequency signals, and the frequency domain coefficients of the high-frequency signal are the group envelope values corresponding to the frequency domain coefficients; Upsampling the basic audio signal to obtain a third audio signal; performing a frequency domain transform on the third audio signal frame by frame to obtain a plurality of frequency domain coefficients of a low-frequency signal corresponding to the third audio signal; Performing frequency domain inverse transformation according to the multiple frequency domain coefficients of the high frequency signal and the multiple frequency domain coefficients of the low frequency signal to obtain the extended audio signal.

42. The device according to claim 41, characterized in that The decoding module is further configured to: The basic frame is decoded to obtain the basic audio signal.

43. The device according to claim 42, characterized in that The decoding module is specifically used for: The basic frame is decoded according to a time domain coding and decoding method to obtain the basic audio signal.

44. The device according to claim 41, characterized in that The decoding module is specifically used for: If the extended frame includes differences between a plurality of frequency domain coefficients of the high frequency signal and corresponding group envelope values, obtaining a plurality of frequency domain coefficients of the high frequency signal according to the plurality of group envelope values of the high frequency signal and the differences between the plurality of frequency domain coefficients of the high frequency signal and the corresponding group envelope values; Performing frequency domain inverse transformation according to the multiple frequency domain coefficients of the low-frequency signal and the multiple frequency domain coefficients of the high-frequency signal to obtain the extended audio signal.

45. The device according to any one of claims 41 to 44, characterized in that The frequency domain inverse transformation specifically includes: improving the inverse discrete cosine transform algorithm.

46. The device according to any one of claims 41 to 44, characterized in that The group of envelope values includes an average value of multiple frequency domain coefficients in each group obtained by grouping the multiple frequency domain coefficients in order from low frequency to high frequency.

47. An audio signal processing device, characterized in that The device comprises: a receiving module, configured to receive a basic frame and an extended frame sent from a first device, wherein the frame length of the extended frame is greater than the frame length of the basic frame, and the extended frame is obtained by re-encoding audio signals corresponding to multiple basic frames; A decoding module, configured to jointly decode the basic frame and the extended frame to obtain an extended audio signal; The decoding module is specifically used for: If the extended frame includes a plurality of group envelope values of a low-frequency signal and a plurality of group envelope values of a high-frequency signal, obtaining a plurality of frequency domain coefficients of the low-frequency signal according to the plurality of group envelope values of the low-frequency signal, and obtaining a plurality of frequency domain coefficients of the high-frequency signal according to the plurality of group envelope values of the high-frequency signal; The multiple frequency domain coefficients of the low-frequency signal are determined by performing a frequency domain transformation on a basic audio signal obtained from the basic frame, or the multiple frequency domain coefficients of the low-frequency signal are determined based on multiple group envelope values of the low-frequency signal in the extended frame, and the multiple frequency domain coefficients of the low-frequency signal are group envelope values corresponding to the frequency domain coefficients; Performing frequency domain inverse transformation according to the multiple frequency domain coefficients of the low-frequency signal and the multiple frequency domain coefficients of the high-frequency signal to obtain the extended audio signal.

48. The device according to claim 47, characterized in that The decoding module is further configured to: The basic frame is decoded to obtain the basic audio signal.

49. The device according to claim 48, characterized in that The decoding module is specifically used for: The basic frame is decoded according to a time domain coding and decoding method to obtain the basic audio signal.

50. The device according to any one of claims 47 to 49, characterized in that The frequency domain inverse transformation specifically includes: improving the inverse discrete cosine transform algorithm.

51. The device according to any one of claims 47 to 49, characterized in that The group of envelope values includes an average value of multiple frequency domain coefficients in each group obtained by grouping the multiple frequency domain coefficients in order from low frequency to high frequency.

52. An audio signal processing device, characterized in that The device comprises: a receiving module, configured to receive a basic frame and an extended frame sent from a first device, wherein the frame length of the extended frame is greater than the frame length of the basic frame, and the extended frame is obtained by re-encoding audio signals corresponding to multiple basic frames; A decoding module, configured to decode the basic frame to obtain a basic audio signal; The decoding module is specifically used for: If the basic frame includes a plurality of frequency domain coefficients of a low-frequency signal and a plurality of group envelope values of a high-frequency signal, obtaining the plurality of frequency domain coefficients of the low-frequency signal and the plurality of frequency domain coefficients of the high-frequency signal according to the basic frame, wherein the plurality of frequency domain coefficients of the high-frequency signal are the group envelope values corresponding to the frequency domain coefficients; Performing an inverse frequency domain transform on the multiple frequency domain coefficients of the low-frequency signal and the multiple frequency domain coefficients of the high-frequency signal to obtain the basic audio signal.

53. The device according to claim 52, characterized in that The decoding module is further configured to: The basic frame and the extended frame are jointly decoded to obtain an extended audio signal.

54. The device according to claim 53, characterized in that The decoding module is specifically used for: If the extended frame includes differences between a plurality of frequency domain coefficients of the high frequency signal and corresponding group envelope values, obtaining a plurality of frequency domain coefficients of the high frequency signal according to the plurality of group envelope values of the high frequency signal and the differences between the plurality of frequency domain coefficients of the high frequency signal and the corresponding group envelope values; Performing frequency domain inverse transformation according to the multiple frequency domain coefficients of the low-frequency signal and the multiple frequency domain coefficients of the high-frequency signal to obtain the extended audio signal.

55. The device according to any one of claims 52 to 54, characterized in that The frequency domain inverse transformation specifically includes: improving the inverse discrete cosine transform algorithm.

56. The device according to any one of claims 52 to 54, characterized in that The group of envelope values includes an average value of multiple frequency domain coefficients in each group obtained by grouping the multiple frequency domain coefficients in order from low frequency to high frequency.

57. An electronic device, characterized in that: The electronic device comprises: processor and transmission interface; a memory for storing instructions executable by the processor; The processor is configured to execute the instructions so that the electronic device implements the audio signal processing method according to any one of claims 1 to 12.

58. An electronic device, characterized in that: The electronic device comprises: processor and transmission interface; a memory for storing instructions executable by the processor; The processor is configured to execute the instructions so that the electronic device implements the audio signal processing method according to any one of claims 13 to 28.

59. A computer-readable storage medium, characterized in that When the instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to perform the audio signal processing method according to any one of claims 1 to 28.

60. A computer program product, characterized in that When the computer program product is run on a computer, the computer is enabled to perform the audio signal processing method according to any one of claims 1 to 28.

Citation Information

Patent Citations

  • Sound encoding apparatus and sound encoding method

    CN101425294A