Audio Processing Method, Apparatus, Storage Medium, Device, and Product
By receiving the spectrum characteristic parameters of high-frequency signals and the low-frequency encoding data of low-frequency signals, using prediction network and orthogonal mirror synthesis filtering processing, the problem of insufficient code rate in audio encoding and decoding is solved, and bandwidth reduction and audio playback effect are improved.
Patent Information
- Application Number
- CN202111371005.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-18
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2041-11-18
AI Technical Summary
In the prior art, the code rate reduction is limited during the audio encoding and decoding process, resulting in poor transmission bandwidth reduction and poor audio playback effect.
By receiving the spectrum characteristic parameters of the high-frequency signal and the low-frequency coded data of the low-frequency signal, decoding and prediction network matching processing are performed to generate the predicted high-frequency signal, and combined with the orthogonal mirror synthesis filtering process, an audio output signal is generated.
It effectively reduces the transmission bandwidth of audio data, while ensuring the audio playback effect, realizing the controllable generation of audio signals and accurate restoration of high-frequency signals.
Smart Images

Figure CN114333861B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and in particular, to an audio processing method, apparatus, storage medium, device, and product. Background Art
[0002] Audio processing usually mainly performs audio encoding and decoding processing. In the audio encoding and decoding processing process, the sound signal is mainly collected by the acquisition end, and the acquisition end will encode and compress the audio signal of the collected sound signal and then send it to the receiving end, and the receiving end decodes and plays the sound.
[0003] Currently, in the related art, the acquisition end will adopt a certain method to reduce the bit rate of the audio signal and then transmit it to the receiving end to reduce the transmission bandwidth. However, in the current method, the reduction of the bit rate is limited, resulting in a poor effect of reducing the transmission bandwidth, and often in order to reduce the bit rate, the generation of the audio output signal at the receiving end is uncontrollable, resulting in a poor audio playback effect. Summary of the Invention
[0004] The embodiments of the present application provide an audio processing solution, which can effectively reduce the transmission bandwidth of audio data and ensure the audio playback effect.
[0005] To solve the above technical problems, the embodiments of the present application provide the following technical solutions:
[0006] According to an embodiment of the present application, an audio processing method includes: receiving the spectral feature parameters of a high-frequency signal and the low-frequency encoded data of a low-frequency signal, where the high-frequency signal and the low-frequency signal belong to a target audio signal; performing decoding processing on the low-frequency encoded data to generate a decoded low-frequency signal; performing prediction network matching processing based on the spectral feature parameters to obtain an audio prediction network that matches the spectral feature parameters; performing audio prediction processing based on the audio prediction network and the decoded low-frequency signal to generate a predicted high-frequency signal corresponding to the high-frequency signal; and generating an audio output signal corresponding to the target audio signal according to the predicted high-frequency signal and the decoded low-frequency signal.
[0007] According to an embodiment of the present application, an audio processing device includes: a receiving module for receiving the spectral feature parameters of a high-frequency signal and the low-frequency encoded data of a low-frequency signal, where the high-frequency signal and the low-frequency signal belong to a target audio signal; a decoding module for performing decoding processing on the low-frequency encoded data to generate a decoded low-frequency signal; a matching module for performing prediction network matching processing based on the spectral feature parameters to obtain an audio prediction network that matches the spectral feature parameters; a prediction module for performing audio prediction processing based on the audio prediction network and the decoded low-frequency signal to generate a predicted high-frequency signal corresponding to the high-frequency signal; and an output module for generating an audio output signal corresponding to the target audio signal according to the predicted high-frequency signal and the decoded low-frequency signal.
[0008] In some embodiments of the present application, the spectral feature parameters include a spectral envelope type; the matching module includes: an information acquisition unit for acquiring the network information of at least one preset audio prediction network, and each piece of network information corresponds to a preset spectral envelope type; a network matching unit for determining the network information corresponding to the preset spectral envelope type that matches the spectral envelope type to obtain target network information; and a network determination unit for determining the preset audio prediction network corresponding to the target network information as the audio prediction network that matches the spectral feature parameters.
[0009] In some embodiments of the present application, the prediction module includes: an extraction processing unit for performing spectral feature extraction processing on the decoded low-frequency signal to obtain low-frequency spectral information; an information prediction unit for using the audio prediction network to perform audio prediction processing based on the low-frequency spectral information to obtain predicted spectral information; and a signal generation unit for generating a predicted high-frequency signal corresponding to the high-frequency signal based on the predicted spectral information.
[0010] In some embodiments of the present application, the extraction processing unit is used to: perform modified discrete cosine transform processing on the decoded low-frequency signal to obtain the low-frequency spectral information; and the signal generation unit is used to: perform inverse modified discrete cosine transform processing on the predicted spectral information to generate a predicted high-frequency signal corresponding to the high-frequency signal.
[0011] In some embodiments of the present application, the output module is used to: perform quadrature mirror synthesis filtering processing on the predicted high-frequency signal and the decoded low-frequency signal to generate the audio output signal.
[0012] According to an embodiment of the present application, an audio processing method includes: decomposing a target audio signal to generate a high-frequency signal and a low-frequency signal; performing feature extraction processing on the high-frequency signal to obtain spectral feature parameters corresponding to the high-frequency signal; performing audio encoding processing on the low-frequency signal to generate low-frequency encoded data corresponding to the low-frequency signal; and sending the spectral feature parameters and the low-frequency encoded data to a receiving end, so that the receiving end determines an audio prediction network matching the spectral feature parameters, and generates an audio output signal based on the audio prediction network and a decoded low-frequency signal obtained by decoding the low-frequency encoded data.
[0013] According to an embodiment of the present application, an audio processing device includes: a decomposition module configured to decompose a target audio signal to generate a high-frequency signal and a low-frequency signal; an extraction module configured to perform feature extraction processing on the high-frequency signal to obtain spectral feature parameters corresponding to the high-frequency signal; an encoding module configured to perform audio encoding processing on the low-frequency signal to generate low-frequency encoded data corresponding to the low-frequency signal; and a transmission module configured to send the spectral feature parameters and the low-frequency encoded data to a receiving end, so that the receiving end determines an audio prediction network matching the spectral feature parameters, and generates an audio output signal based on the audio prediction network and a decoded low-frequency signal obtained by decoding the low-frequency encoded data.
[0014] In some embodiments of the present application, the extraction module includes: a frequency-domain conversion unit configured to perform frequency-domain conversion processing on the high-frequency signal to obtain a frequency-domain signal; a power spectral value calculation unit configured to calculate power spectral values of each frequency point in the frequency-domain signal; and a spectral feature parameter acquisition unit configured to perform feature extraction processing based on the power spectral values of each frequency point to obtain spectral feature parameters describing the spectral distribution characteristics of the high-frequency signal.
[0015] In some embodiments of the present application, the spectral feature parameter acquisition unit includes: an element calculation sub-unit configured to calculate an average value of the power spectral values of each frequency point and determine a maximum power spectral value among the power spectral values of each frequency point; a difference calculation sub-unit configured to perform a difference calculation between the maximum power spectral value and the average value to obtain a first difference value; and a spectral feature parameter determination sub-unit configured to determine the spectral feature parameters corresponding to the high-frequency signal according to the first difference value.
[0016] In some embodiments of the present application, the spectral feature parameter includes a spectral envelope type; the spectral feature parameter determination subunit is configured to: if the first difference is less than a first predetermined threshold value and the maximum power spectral value is less than a second predetermined threshold value, determine that the spectral envelope type corresponding to the high-frequency signal is a first type; if the first difference is less than the first predetermined threshold value and the maximum power spectral value is greater than the second predetermined threshold value, determine that the spectral envelope type corresponding to the high-frequency signal is a second type.
[0017] In some embodiments of the present application, the spectral feature parameter includes a spectral envelope type; the spectral feature parameter determination subunit is configured to: if the first difference is greater than a first predetermined threshold value, perform a normalization process on the power spectral values of each of the frequency points to obtain a normalized value corresponding to each of the frequency points; obtain at least one preset target value, where each target value corresponds to a preset spectral envelope type; calculate the mean square error value between the normalized value corresponding to each of the frequency points and each target value; determine the preset spectral envelope type corresponding to the target value corresponding to the smallest mean square error value as the spectral envelope type of the high-frequency signal.
[0018] In some embodiments of the present application, the spectral feature parameter determination subunit is configured to: perform a difference operation on the power spectral values of each of the frequency points and the average value respectively to obtain a second difference corresponding to each of the frequency points; calculate the square value of the second difference corresponding to each of the frequency points, and calculate the average value of the square values to obtain a normalized score value; divide the second difference corresponding to each of the frequency points by the normalized score value respectively to obtain a normalized value corresponding to each of the frequency points.
[0019] In some embodiments of the present application, the decomposition module is configured to: perform an orthogonal mirror decomposition filtering process on the target audio signal to generate the high-frequency signal and the low-frequency signal.
[0020] According to another embodiment of the present application, a computer-readable storage medium stores a computer program thereon, and when the computer program is executed by a processor of a computer, the computer is caused to execute the method described in the embodiments of the present application.
[0021] According to another embodiment of the present application, an electronic device includes: a memory storing a computer program; a processor reading the computer program stored in the memory to execute the method described in the embodiments of the present application.
[0022] According to another embodiment of the present application, a computer program product or a computer program includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the methods provided in the various alternative implementations described in the embodiments of the present application.
[0023] In the embodiments of the present application, spectral feature parameters of a high-frequency signal and low-frequency encoded data of a low-frequency signal are received, where the high-frequency signal and the low-frequency signal are generated by decomposing a target audio signal; the low-frequency encoded data is decoded to generate a decoded low-frequency signal; a prediction network matching process is performed based on the spectral feature parameters to obtain an audio prediction network matching the spectral feature parameters; an audio prediction process is performed based on the audio prediction network and the decoded low-frequency signal to generate a predicted high-frequency signal corresponding to the high-frequency signal; and an audio output signal corresponding to the target audio signal is generated according to the predicted high-frequency signal and the decoded low-frequency signal.
[0024] In this way, for a target audio signal, the spectral distribution characteristics of the high-frequency signal therein can be described by spectral feature parameters with a very small data size. When receiving data, only the spectral feature parameters and the low-frequency encoded data of the low-frequency signal need to be transmitted, effectively reducing the transmission bandwidth. At the same time, a matching audio prediction network is selected based on the spectral feature parameters to restore the high-frequency signal and generate a predicted high-frequency signal. Since the general spectral distribution characteristics can be described by a very small data size, the error between the predicted high-frequency signal and the original high-frequency signal is controllable, making the generation of the audio output signal controllable. Furthermore, the overall coding rate in the audio processing process is effectively reduced and the ability to restore the predicted high-frequency signal is strong, effectively reducing the transmission bandwidth of audio data and ensuring the audio playback effect. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those skilled in the art, other drawings can be obtained based on these drawings without creative efforts.
[0026] Figure 1 FIG. shows a schematic diagram of a system to which the embodiments of the present application can be applied.
[0027] Figure 2 FIG. shows a flowchart of an audio processing method according to an embodiment of the present application.
[0028] Figure 3Shows a schematic diagram of a spectrum envelope type according to an embodiment of the present application.
[0029] Figure 4 Shows a flowchart of an audio processing method according to another embodiment of the present application.
[0030] Figure 5 Shows a flowchart of an audio processing process in a scenario.
[0031] Figure 6 Shows another flowchart of an audio processing process in a scenario.
[0032] Figure 7 Shows a block diagram of an audio processing apparatus according to an embodiment of the present application.
[0033] Figure 8 Shows a block diagram of an audio processing apparatus according to another embodiment of the present application.
[0034] Figure 9 Shows a block diagram of an electronic device according to an embodiment of the present application. Detailed Embodiments
[0035] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative efforts shall fall within the protection scope of the present application.
[0036] In the following description, specific embodiments of the present application will be described with reference to steps and symbols executed by one or more computers, unless otherwise stated. Therefore, these steps and operations will be mentioned several times as being executed by a computer. The computer execution referred to herein includes operations of a computer processing unit that represents electronic signals in a structured form of data. This operation transforms the data or maintains it at a location in the computer's memory system, which can be reconfigured or otherwise changed in a manner well known to those skilled in the art. The data structure maintained by the data is the physical location of the memory, which has specific characteristics defined by the data format. However, the principles of the present application are described in the above text, which does not represent a limitation. Those skilled in the art will understand that the various steps and operations described below can also be implemented in hardware.
[0037] Figure 1 Shows a schematic diagram of a system 100 to which embodiments of the present application can be applied. As Figure 1 shown, the system 100 may include a terminal 101, a terminal 102, and a server 103.
[0038] The terminals 101 and 102 can be any devices. The terminal 102 includes, but is not limited to, mobile phones, computers, intelligent voice interaction devices, smart home appliances, vehicle-mounted terminals, VR / AR devices, smart watches, and computers, etc.
[0039] The server 103 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, as well as big data and artificial intelligence platforms. In one implementation manner of this example, the server 103 is a cloud server.
[0040] In some examples, the terminals 101, 102, and the server 103 can be nodes in a blockchain network, which can improve the security of audio processing.
[0041] In one implementation manner of this example, the terminal 101 can: receive the spectral feature parameters of the high-frequency signal and the low-frequency coding data of the low-frequency signal, where the high-frequency signal and the low-frequency signal belong to the target audio signal; perform decoding processing on the low-frequency coding data to generate a decoded low-frequency signal; perform prediction network matching processing based on the spectral feature parameters to obtain an audio prediction network that matches the spectral feature parameters; perform audio prediction processing based on the audio prediction network and the decoded low-frequency signal to generate a predicted high-frequency signal corresponding to the high-frequency signal; generate an audio output signal corresponding to the target audio signal according to the predicted high-frequency signal and the decoded low-frequency signal.
[0042] Among them, in one example, the spectral feature parameters of the high-frequency signal and the low-frequency coding data of the low-frequency signal can be directly sent by the terminal 102 as a collection end to the terminal 101 as a receiving end; in one example, the spectral feature parameters of the high-frequency signal and the low-frequency coding data of the low-frequency signal can be sent by the terminal 102 to the terminal 101 through the server 103; in one example, the spectral feature parameters of the high-frequency signal and the low-frequency coding data of the low-frequency signal can be sent from the A processing unit as the collection end in the terminal 101 to the B processing unit as the receiving end in the terminal 101.
[0043] In one implementation of this example, the terminal 102 may: decompose the target audio signal to generate a high-frequency signal and a low-frequency signal; perform feature extraction processing on the high-frequency signal to obtain the spectral feature parameters corresponding to the high-frequency signal; perform audio encoding processing on the low-frequency signal to generate low-frequency encoded data corresponding to the low-frequency signal; and send the spectral feature parameters and the low-frequency encoded data to the receiving end, so that the receiving end determines an audio prediction network that matches the spectral feature parameters, and generates an audio output signal based on the audio prediction network and the decoded low-frequency signal decoded from the low-frequency encoded data.
[0044] Among them, in one example, the receiving end may be the terminal 101, and the terminal 102 may directly send the spectral feature parameters and the low-frequency encoded data to the terminal 101, or the terminal 102 may send the spectral feature parameters and the low-frequency encoded data to the terminal 101 through the server 103; in one example, the receiving end may be a C processing unit in the terminal 102, and another D processing unit in the terminal 102 as the acquisition end may send the spectral feature parameters and the low-frequency encoded data to the C processing unit.
[0045] Figure 2 Schematically shows a flowchart of an audio processing method according to an embodiment of the present application. The execution subject of this audio processing method may be any receiving end, and the receiving end may decode the received audio-related data to generate an audio output signal for playing sound. The receiving end is, for example Figure 1 the shown terminal 101 or terminal 102, etc.
[0046] As Figure 2 shown, this audio processing method may include steps S210 to S250.
[0047] Step S210: Receive the spectral feature parameters of the high-frequency signal and the low-frequency encoded data of the low-frequency signal. The high-frequency signal and the low-frequency signal belong to the target audio signal; Step S220: Decode the low-frequency encoded data to generate a decoded low-frequency signal; Step S230: Perform a prediction network matching process based on the spectral feature parameters to obtain an audio prediction network that matches the spectral feature parameters; Step S240: Perform audio prediction processing based on the audio prediction network and the decoded low-frequency signal to generate a predicted high-frequency signal corresponding to the high-frequency signal; Step S250: Generate an audio output signal corresponding to the target audio signal according to the predicted high-frequency signal and the decoded low-frequency signal.
[0048] In this way, for the target audio signal, the spectral distribution characteristics of the high-frequency signal therein can be described by spectral feature parameters with a very small data size. When receiving data, only the spectral feature parameters and the low-frequency encoded data of the low-frequency signal need to be transmitted, effectively reducing the transmission bandwidth. At the same time, based on the spectral feature parameters, a matching audio prediction network is selected to restore the high-frequency signal and generate a predicted high-frequency signal. Since the general spectral distribution characteristics can be described by a very small data size, the error between the predicted high-frequency signal and the original high-frequency signal is controllable, making the generation of the audio output signal controllable. Furthermore, during the audio processing process, the overall coding rate is effectively reduced and the ability to restore and predict the high-frequency signal is strong, effectively reducing the transmission bandwidth of the audio data and ensuring the audio playback effect.
[0049] The following description Figure 2 In the embodiments shown, when performing audio processing, the specific processes of each step are as follows.
[0050] In step S210, the spectral feature parameters of the high-frequency signal and the low-frequency encoded data of the low-frequency signal are received, and the high-frequency signal and the low-frequency signal are generated by decomposing the target audio signal.
[0051] The target audio signal can be a digital sound signal generated by the acquisition end through analog-to-digital conversion of the collected sound signal. The target audio signal can be decomposed into a high-frequency signal and a low-frequency signal. The high-frequency signal can be a part of the audio signal in the target audio signal that is higher than a predetermined frequency, and the low-frequency signal can be a part of the audio signal in the target audio signal that is lower than a predetermined frequency.
[0052] The acquisition end can extract the spectral distribution characteristics of the high-frequency signal and generate spectral feature parameters describing the spectral distribution characteristics according to the extracted spectral distribution characteristics. The spectral feature parameters can be numbers or identifiers, etc. The data size of the spectral feature parameters can be controlled to be very small. For example, in one example, a spectral distribution characteristic can be described by "1", and only 3 bits (bits) of "1" are needed to describe a spectral distribution characteristic. And the low-frequency signal can be encoded using a traditional speech encoder (which can be an encoder such as CELP, SILK, AAC, etc.) to generate low-frequency encoded data.
[0053] After the acquisition end generates the spectral feature parameters of the high-frequency signal and the low-frequency encoded data of the low-frequency signal, it can send the spectral feature parameters of the high-frequency signal and the low-frequency encoded data of the low-frequency signal to the receiving end. Among them, when sending, the acquisition end can form an encoded bitstream with the spectral feature parameters and the low-frequency encoded data and send it to the receiving end, and the bitstream of this encoded bitstream can be extremely small.
[0054] In step S220, the low-frequency encoded data is decoded to generate a decoded low-frequency signal.
[0055] The receiving end can decode the low-frequency encoded data through a traditional speech decoder to generate a decoded low-frequency signal, which is also the decoded low-frequency signal.
[0056] In step S230, a prediction network matching process is performed based on the spectral feature parameters to obtain an audio prediction network that matches the spectral feature parameters.
[0057] The audio prediction network is a prediction network for predicting high-frequency signals, and the audio prediction network can be a deep learning network. Performing a prediction network matching process based on spectral feature parameters means determining an audio prediction network that matches the spectral feature parameters. It can be understood that each spectral feature parameter can correspond to an audio prediction network.
[0058] The audio prediction network is used to predict the predicted high-frequency signal corresponding to the high-frequency signal based on the low-frequency encoded data corresponding to the low-frequency signal to achieve the mapping of the audio from low frequency to high frequency. There is diversity in the mapping of the audio from low frequency to high frequency. Adding the spectral feature parameters of the high-frequency signal to match the corresponding audio prediction network can limit the diversity of the samples, that is, training multiple preset audio prediction networks, each network corresponding to its own training samples (the training samples can include input samples: the signal feature information of the low-frequency signal (such as spectral information), and the expected output: the signal feature information of the high-frequency signal (such as spectral information)) and training independently, and an audio prediction network that matches the spectral feature parameters can be determined from multiple preset audio prediction networks.
[0059] Among them, each network corresponds to its own training samples and is trained independently. For example, the preset audio prediction network A can be trained based on the training samples corresponding to the spectral feature parameter 1, and the preset audio prediction network B can be trained based on the training samples corresponding to the spectral feature parameter 2. At this time, if the received spectral feature parameter is 1, the audio prediction network that matches the received spectral feature parameter is the preset audio prediction network A.
[0060] In one embodiment, the spectral feature parameters include the spectral envelope type; step S230, performing a prediction network matching process based on the spectral feature parameters to obtain an audio prediction network that matches the spectral feature parameters, includes: obtaining the network information of at least one preset audio prediction network, each network information corresponding to a preset spectral envelope type; determining the network information corresponding to the preset spectral envelope type that matches the spectral envelope type to obtain the target network information; and determining the preset audio prediction network corresponding to the target network information as the audio prediction network that matches the spectral feature parameters.
[0061] The spectral feature parameter is information describing the spectral distribution feature. In this embodiment, the spectral feature parameter is the spectral envelope type, the spectral distribution feature is the spectral envelope feature, the spectral envelope type is the information describing the spectral envelope feature, and the spectral envelope feature can represent the change trend of the spectral values in the signal spectrum, that is, each change trend of the spectral values corresponds to a spectral envelope type.
[0062] The preset spectral envelope type is the preset spectral envelope type, the network information can be the identifier of the preset audio prediction network, and the preset audio prediction network is the pre-trained audio prediction network. Each network information corresponds to a preset spectral envelope type, that is, each preset audio prediction network corresponds to a preset spectral envelope type.
[0063] In this embodiment, each network corresponds to its own training samples and is trained independently. For example, the A preset audio prediction network can be trained based on the training samples corresponding to the spectral envelope type 1, and the B preset audio prediction network can be trained based on the training samples corresponding to the spectral envelope type 2. At this time, if the received spectral envelope type is 1, the audio prediction network that matches the received spectral envelope type is the A preset audio prediction network.
[0064] Determine the preset spectral envelope type that matches the received spectral envelope type. For example, if the received spectral envelope type is 1, the preset spectral envelope type that matches the received spectral envelope type is 1. If the network information corresponding to the matching preset spectral envelope type 1 is X, then the target network information is X. Furthermore, the audio prediction network that matches the spectral feature parameter is the preset audio prediction network indicated by X.
[0065] The method of matching the prediction network with the spectral envelope type can effectively improve the accuracy of the audio prediction network in predicting high-frequency signals.
[0066] In one implementation, refer to Figure 3 , the preset spectral envelope types include 8 types: type "0" is the low-energy flat type; type "1" is the high-energy flat type; type "2" is the energy convex type; type "3" is the energy concave type; type "4" is the energy increasing type; type "5" is the energy decreasing type; type "6" is the step type with high energy in the front and low energy in the back; type "7" is the step type with low energy in the front and high energy in the back. The applicant finds that with this type of classification method, the method of matching the prediction network with the spectral envelope type can further improve the prediction accuracy of the audio prediction network.
[0067] It can be understood that in some other embodiments, the spectral feature parameters may include other feature information, such as information describing features such as the proportion of power spectral values exceeding a predetermined threshold in the spectrum; and further, step S230, performing a prediction network matching process based on the spectral feature parameters to obtain an audio prediction network matching the spectral feature parameters, may include: obtaining network information of at least one preset audio prediction network, each piece of the network information corresponding to a preset spectral envelope type; determining the network information corresponding to the preset spectral envelope type that matches the other feature information to obtain target network information; and determining the preset audio prediction network corresponding to the target network information as the audio prediction network matching the spectral feature parameters.
[0068] In step S240, an audio prediction process is performed based on the audio prediction network and the decoded low-frequency signal to generate a predicted high-frequency signal corresponding to the high-frequency signal.
[0069] When making a prediction based on an audio prediction network matched with spectral feature parameters, since the general spectral distribution characteristics can be described by a very small data size, the error between the predicted high-frequency signal and the original high-frequency signal is controllable. The matched audio prediction network can accurately predict the high-frequency signal. Compared with the method of using a unified prediction network, the mediocrity of the predicted high-frequency signal can be avoided, and the accuracy of the predicted high-frequency signal can be improved.
[0070] Among them, the signal feature information (such as spectral information) of the decoded low-frequency signal can be input into the audio prediction network, and the audio prediction network can predict and output the signal feature information (such as spectral information) of the high-frequency signal. The output signal features can be used to restore the high-frequency signal to generate a predicted audio signal.
[0071] In one embodiment, step S240, performing an audio prediction process based on the audio prediction network and the decoded low-frequency signal to generate a predicted high-frequency signal corresponding to the high-frequency signal, includes: performing a spectral feature extraction process on the decoded low-frequency signal to obtain low-frequency spectral information; using the audio prediction network to perform an audio prediction process based on the low-frequency spectral information to obtain predicted spectral information; and generating a predicted high-frequency signal corresponding to the high-frequency signal based on the predicted spectral information.
[0072] The spectral feature extraction process is performed on the decoded low-frequency signal through a neural network or a modified discrete cosine transform (MDCT) to obtain low-frequency spectral information. Then, the low-frequency spectral information is input into the audio prediction network, and the audio prediction network performs an audio prediction process, thereby outputting predicted spectral information. The predicted high-frequency signal corresponding to the high-frequency signal can be obtained by performing audio restoration based on the predicted spectral information.
[0073] In one embodiment, the process of extracting spectral features from the decoded low-frequency signal to obtain low-frequency spectral information includes: performing an improved discrete cosine transform on the decoded low-frequency signal to obtain the low-frequency spectral information; and generating a predicted high-frequency signal corresponding to the high-frequency signal based on the predicted spectral information includes: performing an inverse improved discrete cosine transform on the predicted spectral information to generate a predicted high-frequency signal corresponding to the high-frequency signal.
[0074] In this embodiment, an improved discrete cosine transform can be performed on the decoded low-frequency signal by an improved discrete cosine transformer (MDCT, Modified Discrete Cosine Transform) to obtain low-frequency spectral information. Then, for the predicted spectral information predicted by the audio prediction network, an inverse improved discrete cosine transform can be performed by an inverse modified discrete cosine transformer (IMDCT, Inverse Modified Discrete Cosine Transform) to generate a predicted high-frequency signal. In this prediction method, the applicant has found that the accuracy of the predicted audio signal can be further improved.
[0075] In step S250, an audio output signal corresponding to the target audio signal is generated according to the predicted high-frequency signal and the decoded low-frequency signal.
[0076] The predicted high-frequency signal is the predicted high-frequency signal, and the decoded low-frequency signal is the original low-frequency signal. By synthesizing the predicted high-frequency signal and the decoded low-frequency signal, an audio output signal corresponding to the original target audio signal can be generated.
[0077] In one embodiment, step S250 of generating an audio output signal corresponding to the target audio signal according to the predicted high-frequency signal and the decoded low-frequency signal includes: performing quadrature mirror synthesis filtering on the predicted high-frequency signal and the decoded low-frequency signal to generate the audio output signal. Among them, a quadrature mirror filter (QMF) can be used to perform quadrature mirror synthesis filtering on the predicted high-frequency signal and the decoded low-frequency signal to generate a full-band audio output signal corresponding to the target audio signal.
[0078] It can be understood that in other embodiments, other existing synthesis filters can be used to perform synthesis filtering on the predicted high-frequency signal and the decoded low-frequency signal to generate an audio output signal.
[0079] Figure 4 A flowchart of an audio processing method according to an embodiment of the present application is schematically shown. The execution subject of this audio processing method can be any terminal, such as Figure 1 the shown terminal 101 or terminal 102, etc.
[0080] As shown Figure 4 in the figure, the audio processing method may include steps S310 to S340.
[0081] Step S310: Decompose the target audio signal to generate a high-frequency signal and a low-frequency signal; Step S320: Extract features from the high-frequency signal to obtain spectral feature parameters corresponding to the high-frequency signal; Step S330: Perform audio encoding on the low-frequency signal to generate low-frequency encoded data corresponding to the low-frequency signal; Step S340: Send the spectral feature parameters and the low-frequency encoded data to the receiving end so that the receiving end determines an audio prediction network that matches the spectral feature parameters, and generates an audio output signal based on the audio prediction network and the decoded low-frequency signal obtained by decoding the low-frequency encoded data.
[0082] In this way, for the target audio signal, the spectral distribution characteristics of the high-frequency signal therein can be described by spectral feature parameters with a very small data size. When transmitting, only the spectral feature parameters and the low-frequency encoded data of the low-frequency signal need to be transmitted, effectively reducing the transmission bandwidth. At the same time, based on the spectral feature parameters, a matching audio prediction network is selected to restore the high-frequency signal and generate a predicted high-frequency signal. Since the general spectral distribution characteristics can be described by a very small data size, the error between the predicted high-frequency signal and the original high-frequency signal is controllable, making the generation of the audio output signal controllable. Furthermore, during the audio processing process, the overall coding rate is effectively reduced and the ability to restore and predict the high-frequency signal is strong, effectively reducing the transmission bandwidth of the audio data and ensuring the audio playback effect.
[0083] The following describes Figure 3 the specific processes of the respective steps when performing audio processing in the embodiments shown.
[0084] In step S310, the target audio signal is decomposed to generate a high-frequency signal and a low-frequency signal.
[0085] The target audio signal may be a digital sound signal generated by the acquisition end through analog-to-digital conversion of the acquired sound signal. The target audio signal can be decomposed into a high-frequency signal and a low-frequency signal. The high-frequency signal may be a part of the audio signal in the target audio signal that is higher than a predetermined frequency, and the low-frequency signal may be a part of the audio signal in the target audio signal that is lower than a predetermined frequency.
[0086] Among them, a band-pass filter (BPF) group or a quadrature mirror filter (QMF) group, etc., may be used to decompose the target audio signal in a corresponding manner to generate a high-frequency signal and a low-frequency signal.
[0087] In one embodiment, in step S310, the target audio signal is decomposed to generate a high-frequency signal and a low-frequency signal, including: performing quadrature mirror decomposition filtering on the target audio signal to generate the high-frequency signal and the low-frequency signal. Among them, the target audio signal can be subjected to quadrature mirror decomposition filtering through a quadrature mirror filter (QMF) bank to generate a high-frequency signal and a low-frequency signal.
[0088] In step S320, feature extraction is performed on the high-frequency signal to obtain spectral feature parameters corresponding to the high-frequency signal.
[0089] The acquisition end can extract the spectral distribution features of the high-frequency signal and generate spectral feature parameters describing the spectral distribution features according to the extracted spectral distribution features. Among them, the spectral distribution features can include spectral envelope features and the proportion of power spectral values exceeding a predetermined threshold in the spectrum, etc. In one implementation manner of this example, the spectral distribution feature is the spectral envelope feature, and the spectral feature parameter is the spectral envelope type.
[0090] The spectral feature parameter can be a number or an identifier, etc. The data size of the spectral feature parameter can be controlled to be extremely small. For example, in one example, a spectral distribution feature is described by "1", and only 3 bits (bits) of "1" are needed to describe a spectral distribution feature.
[0091] In one embodiment, in step S320, feature extraction is performed on the high-frequency signal to obtain spectral feature parameters corresponding to the high-frequency signal, including: performing frequency-domain conversion on the high-frequency signal to obtain a frequency-domain signal; calculating the power spectral values of each frequency point in the frequency-domain signal; and performing feature extraction based on the power spectral values of each frequency point to obtain spectral feature parameters describing the spectral distribution features of the high-frequency signal.
[0092] The high-frequency signal is a time-domain signal. Through frequency-domain conversion processing (such as Fourier transform processing), a frequency-domain signal in the frequency domain can be obtained. Based on the spectrum (i.e., the spectrogram) of the frequency-domain signal, the power spectral values of each frequency point can be extracted. Performing feature extraction based on the power spectral values of each frequency point can accurately analyze the spectral distribution features of the high-frequency signal, and then generate spectral feature parameters describing the spectral distribution features. Among them, in one example, the power spectral values of each frequency point can be the logarithms of the power spectra of each frequency point.
[0093] In one embodiment, feature extraction processing is performed based on the power spectral values of each of the frequency points to obtain spectral feature parameters describing the spectral distribution characteristics of the high-frequency signal, including: calculating the average value of the power spectral values of each of the frequency points and determining the maximum power spectral value among the power spectral values of each of the frequency points; performing a difference operation on the maximum power spectral value and the average value to obtain a first difference; and determining the spectral feature parameters corresponding to the high-frequency signal according to the first difference.
[0094] For example, the power spectral values of each frequency point are x(i), where i ∈ [N1, N2] and i is the serial number of the frequency point. The average value xavg of the power spectral values of each frequency point is the mean value of all x(i). The maximum power spectral value xmax among the power spectral values of each frequency point is the largest one among all x(i). The maximum power spectral value xmax minus the average value xavg is the first difference. In this way, the spectral distribution characteristics can be accurately reflected based on the magnitude of the first difference, especially for the spectral envelope characteristics.
[0095] In some other ways, when performing feature extraction processing based on the power spectral values of each of the frequency points to obtain spectral feature parameters describing the spectral distribution characteristics of the high-frequency signal, the proportion of the power spectral values exceeding a predetermined threshold in the spectrum can be calculated, and the corresponding spectral feature parameters can be determined according to the proportion size.
[0096] In one embodiment, the spectral feature parameters include the spectral envelope type; the determining the spectral feature parameters corresponding to the high-frequency signal according to the first difference includes: if the first difference is less than a first predetermined threshold value and the maximum power spectral value is less than a second predetermined threshold value, determining that the spectral envelope type corresponding to the high-frequency signal is the first type; if the first difference is less than the first predetermined threshold value and the maximum power spectral value is greater than the second predetermined threshold value, determining that the spectral envelope type corresponding to the high-frequency signal is the second type. The first type and the second type respectively describe the corresponding spectral envelope characteristics.
[0097] In one implementation manner under this embodiment, referring to Figure 3 , the preset spectral envelope types include: type "0" is a low-energy flat type; type "1" is a high-energy flat type. If the first difference is less than the first predetermined threshold value C1 and the maximum power spectral value is less than the second predetermined threshold value C2, at this time, it can be determined that the spectral envelope type corresponding to the high-frequency signal is the first type: the low-energy flat type of type "0". If the first difference is less than the first predetermined threshold value C1 and the maximum power spectral value is greater than the second predetermined threshold value C2, it is determined that the spectral envelope type corresponding to the high-frequency signal is the second type: the high-energy flat type of type "1".
[0098] In one embodiment, the spectral feature parameter includes a spectral envelope type; determining the spectral feature parameter corresponding to the high-frequency signal according to the first difference includes: if the first difference is greater than a first predetermined threshold value, normalizing the power spectral values of the respective frequency points to obtain normalized values corresponding to the respective frequency points; obtaining at least one preset target value, each target value corresponding to a preset spectral envelope type; calculating the mean square error values between the normalized values corresponding to the respective frequency points and each target value; and determining the preset spectral envelope type corresponding to the target value with the smallest mean square error value as the spectral envelope type of the high-frequency signal.
[0099] In an implementation manner under this embodiment, referring to Figure 3 , the preset spectral envelope types include: type "2" is energy convex; type "3" is energy concave; type "4" is energy increasing; type "5" is energy decreasing; type "6" is a step type with high energy at the front and low energy at the back; type "7" is a step type with low energy at the front and high energy at the back.
[0100] Each target value corresponds to a preset spectral envelope type, and the target value is a preset value. In one example, the target value can be z(i), i ∈ [N1, N2], i is the serial number of the frequency point, z(i) is less than 1. Taking type "2" as an example, if N2 - N1 + 1 equals 9, then the target value z(i) corresponding to type "2" can be set to 000111000.
[0101] Calculating the mean square error values between the normalized values corresponding to the respective frequency points and each target value, based on the smallest mean square error value, the preset spectral envelope type closest to the spectral distribution characteristics of the high-frequency signal can be accurately determined. For example, if the mean square error value between the target value corresponding to type "2" and the normalized values corresponding to the respective frequency points is smaller than those of types "3" to "7", then it can be determined that type "2" energy convex is the spectral envelope type of the high-frequency signal.
[0102] In one embodiment, the normalizing the power spectral values of the respective frequency points to obtain normalized values corresponding to the respective frequency points includes: subtracting the power spectral values of the respective frequency points from the average value respectively to obtain second differences corresponding to the respective frequency points; calculating the square values of the second differences corresponding to the respective frequency points, and calculating the average value of the square values to obtain a normalized score; and dividing the second differences corresponding to the respective frequency points by the normalized score respectively to obtain normalized values corresponding to the respective frequency points.
[0103] Under this embodiment, specifically, normalization can be performed according to the formula and to obtain the normalized value y corresponding to each frequency point i ( i ), where N2 - N1 + 1 is the total number of frequency points, N2 to N1 is the range of frequency point numbers, xavg is the average value, x(i) is the power spectrum value of each frequency point i, x(i) - xavg is the second difference corresponding to each frequency point, and std is the normalized score (average value of squared values).
[0104] In step S330, perform audio encoding processing on the low-frequency signal to generate low-frequency encoded data corresponding to the low-frequency signal.
[0105] The low-frequency signal can be encoded using a traditional speech encoder (such as CELP, SILK, AAC, etc.) to generate low-frequency encoded data.
[0106] In step S340, send the spectral feature parameters and the low-frequency encoded data to the receiving end so that the receiving end can determine the audio prediction network that matches the spectral feature parameters and generate an audio output signal based on the decoded low-frequency signal obtained by decoding the audio prediction network and the low-frequency encoded data.
[0107] When sending the spectral feature parameters and the low-frequency encoded data to the receiving end, the spectral feature parameters and the low-frequency encoded data can be combined into an encoded bitstream and sent to the receiving end, and the bitstream of this encoded bitstream can be extremely small. The receiving end can be based on Figure 2 the steps in the illustrated embodiment to determine the audio prediction network that matches the spectral feature parameters and generate an audio output signal based on the decoded low-frequency signal obtained by decoding the audio prediction network and the low-frequency encoded data.
[0108] According to the method described in the above embodiments, the following will be further described in detail by way of examples in application scenarios. The meanings of the relevant terms in this scenario are the same as those in the foregoing embodiments, and specific references can be made to the descriptions in the foregoing embodiments. The process of audio processing in this application scenario can refer to Figure 5 and Figure 6 shown, and the foregoing embodiments of the present application are applied to audio processing in this scenario.
[0109] First, refer to Figure 5 , and perform encoding processing during the audio processing at the acquisition end, and this process may include steps S410 to S450.
[0110] In step S410, input the target audio signal: Specifically, the acquisition end can collect the target audio signal generated by analog-to-digital conversion of the sound signal.
[0111] In step S420, QMF decomposition: Specifically, the target audio signal is decomposed to generate a high-frequency signal and a low-frequency signal. Decomposing the target audio signal to generate a high-frequency signal and a low-frequency signal includes: performing quadrature mirror decomposition filtering on the target audio signal to generate the high-frequency signal and the low-frequency signal. Among them, the target audio signal can be subjected to quadrature mirror decomposition filtering through a quadrature mirror filter (QMF) bank to generate a high-frequency signal and a low-frequency signal.
[0112] In step S430, spectral feature frame extraction: Specifically, feature extraction is performed on the high-frequency signal to obtain spectral feature parameters corresponding to the high-frequency signal.
[0113] Performing feature extraction on the high-frequency signal to obtain spectral feature parameters corresponding to the high-frequency signal includes: performing frequency-domain conversion on the high-frequency signal to obtain a frequency-domain signal; calculating the power spectral values of each frequency point in the frequency-domain signal; performing feature extraction based on the power spectral values of each frequency point to obtain spectral feature parameters describing the spectral distribution characteristics of the high-frequency signal.
[0114] Performing feature extraction based on the power spectral values of each frequency point to obtain spectral feature parameters describing the spectral distribution characteristics of the high-frequency signal includes: calculating the average value of the power spectral values of each frequency point and determining the maximum power spectral value among the power spectral values of each frequency point; performing a difference operation on the maximum power spectral value and the average value to obtain a first difference; determining the spectral feature parameters corresponding to the high-frequency signal according to the first difference.
[0115] Among them, the spectral feature parameters include spectral envelope types; specifically, refer to Figure 3 , the preset spectral envelope types include: type "0" is a low-energy flat type; type "1" is a high-energy flat type; type "2" is an energy convex type; type "3" is an energy concave type; type "4" is an energy increasing type; type "5" is an energy decreasing type; type "6" is a step type with high energy in the front and low energy in the back; type "7" is a step type with low energy in the front and high energy in the back.
[0116] For the "0" type which is a low-energy tiled type and the "1" type which is a high-energy tiled type. Determining the spectral characteristic parameters corresponding to the high-frequency signal according to the first difference includes: if the first difference is less than a first predetermined threshold value and the maximum power spectral value is less than a second predetermined threshold value, determining that the spectral envelope type corresponding to the high-frequency signal is the first type; if the first difference is less than the first predetermined threshold value and the maximum power spectral value is greater than the second predetermined threshold value, determining that the spectral envelope type corresponding to the high-frequency signal is the second type. Among them, if the first difference is less than the first predetermined threshold value C1 and the maximum power spectral value is less than the second predetermined threshold value C2, at this time, it can be determined that the spectral envelope type corresponding to the high-frequency signal is the first type: the low-energy tiled type of "0" type. If the first difference is less than the first predetermined threshold value C1 and the maximum power spectral value is greater than the second predetermined threshold value C2, then determine that the spectral envelope type corresponding to the high-frequency signal is the second type: the high-energy tiled type of "1" type.
[0117] For the "2" type to the "7" type. Determining the spectral characteristic parameters corresponding to the high-frequency signal according to the first difference includes: if the first difference is greater than the first predetermined threshold value, performing normalization processing on the power spectral values of each frequency point to obtain the normalized value corresponding to each frequency point; obtaining at least one preset target value, and each target value corresponds to a preset spectral envelope type; calculating the mean square error value between the normalized value corresponding to each frequency point and each target value; determining the preset spectral envelope type corresponding to the target value corresponding to the smallest mean square error value as the spectral envelope type of the high-frequency signal.
[0118] Each target value corresponds to a preset spectral envelope type, and the target value is a preset value. The target value can be z(i), i ∈ [N1, N2], i is the serial number of the frequency point, and z(i) is less than 1. Taking the type "2" as an example, if N2 - N1 + 1 is equal to 9, then the target value z(i) corresponding to the type "2" can be set to 000111000. Calculating the mean square error value between the normalized value corresponding to each frequency point and each target value, and based on the smallest mean square error value, the preset spectral envelope type closest to the spectral distribution characteristic of the high-frequency signal can be accurately determined. For example, if the mean square error value between the target value corresponding to the type "2" and the normalized value corresponding to each frequency point is the smallest compared with the types "3" to "7", then it can be determined that the energy convex type of "2" is the spectral envelope type of the high-frequency signal.
[0119] Normalizing the power spectral values of the respective frequency points to obtain the normalization values corresponding to the respective frequency points includes: performing a difference operation on the power spectral values of the respective frequency points and the average value to obtain the second differences corresponding to the respective frequency points; calculating the square values of the second differences corresponding to the respective frequency points, and calculating the average value of the square values to obtain the normalization score; dividing the second differences corresponding to the respective frequency points by the normalization score to obtain the normalization values corresponding to the respective frequency points. In this embodiment, specifically, it can be based on the formula and to perform normalization processing to obtain the normalization value y corresponding to each frequency point i ( i ) , where N2 - N1 + 1 is the total number of frequency points, N2 to N1 is the frequency point serial number range, xavg is the average value, x(i) is the power spectral value of each frequency point i, x(i) - xavg is the second difference corresponding to each frequency point, and std is the normalization score (average value of the square values).
[0120] In step S440, low-frequency speech coding: Specifically, perform audio coding processing on the low-frequency signal to generate low-frequency coding data corresponding to the low-frequency signal. The low-frequency signal can be encoded using a traditional speech encoder (which can be an encoder such as CELP, SILK, AAC, etc.) to generate low-frequency coding data.
[0121] In step S450, data output: Send the spectral feature parameters and the low-frequency coding data to the receiving end. The spectral feature parameters and the low-frequency coding data can be encapsulated together to form a coding bitstream and sent to the receiving end.
[0122] Further, referring to Figure 6 , perform decoding processing during the audio processing at the acquisition end, and this process can include steps S510 to S550.
[0123] In step S510, input bitstream: Specifically, receive the bitstream sent by the acquisition end, and the bitstream includes the spectral feature parameters of the high-frequency signal and the low-frequency coding data of the low-frequency signal. That is, receive the spectral feature parameters of the high-frequency signal and the low-frequency coding data of the low-frequency signal, and the high-frequency signal and the low-frequency signal are generated by decomposing the target audio signal.
[0124] In step S520, bitstream parsing: Specifically, parse the received bitstream to parse out the spectral feature parameters of the high-frequency signal and the low-frequency coding data of the low-frequency signal in the bitstream.
[0125] In step S530, low-frequency speech decoding: Specifically, perform decoding processing on the low-frequency coding data to generate a decoded low-frequency signal. The receiving end can decode the low-frequency coding data through a traditional speech decoder to generate a decoded low-frequency signal, and the decoded low-frequency signal is the decoded low-frequency signal.
[0126] In step S540, network matching: Specifically, based on the spectral feature parameters, a prediction network matching process is performed to obtain an audio prediction network that matches the spectral feature parameters.
[0127] The spectral feature parameters include a spectral envelope type; performing a prediction network matching process based on the spectral feature parameters to obtain an audio prediction network that matches the spectral feature parameters includes: obtaining network information of at least one preset audio prediction network, where each network information corresponds to a preset spectral envelope type; determining the network information corresponding to the preset spectral envelope type that matches the spectral envelope type to obtain target network information; and determining the preset audio prediction network corresponding to the target network information as the audio prediction network that matches the spectral feature parameters.
[0128] In step S550, prediction processing: Specifically, based on the audio prediction network and the decoded low-frequency signal, audio prediction processing is performed to generate a predicted high-frequency signal corresponding to the high-frequency signal.
[0129] Performing audio prediction processing based on the audio prediction network and the decoded low-frequency signal in step S550 to generate a predicted high-frequency signal corresponding to the high-frequency signal includes: step S551, performing spectral feature extraction processing on the decoded low-frequency signal to obtain low-frequency spectral information; S552, using the audio prediction network to perform audio prediction processing based on the low-frequency spectral information to obtain predicted spectral information; and S553, generating a predicted high-frequency signal corresponding to the high-frequency signal based on the predicted spectral information.
[0130] Performing spectral feature extraction processing on the decoded low-frequency signal to obtain low-frequency spectral information includes: performing modified discrete cosine transform processing on the decoded low-frequency signal to obtain the low-frequency spectral information; generating a predicted high-frequency signal corresponding to the high-frequency signal based on the predicted spectral information includes: performing inverse modified discrete cosine transform processing on the predicted spectral information to generate a predicted high-frequency signal corresponding to the high-frequency signal.
[0131] The modified discrete cosine transform processing of the decoded low-frequency signal can be performed by a modified discrete cosine transformer (MDCT, Modified Discrete Cosine Transform) to obtain low-frequency spectral information. Then, for the predicted spectral information predicted by the audio prediction network, inverse modified discrete cosine transform processing can be performed by an inverse modified discrete cosine transformer (IMDCT, Inverse Modified Discrete Cosine Transform) to generate a predicted high-frequency signal.
[0132] In step S560, QMF synthesis: Specifically, based on the predicted high-frequency signal and the decoded low-frequency signal, an audio output signal corresponding to the target audio signal is generated. Generating an audio output signal corresponding to the target audio signal based on the predicted high-frequency signal and the decoded low-frequency signal includes: performing quadrature mirror synthesis filtering on the predicted high-frequency signal and the decoded low-frequency signal to generate the audio output signal. Among them, a quadrature mirror filter (QMF) can be used to perform quadrature mirror synthesis filtering on the predicted high-frequency signal and the decoded low-frequency signal to generate a full-band audio output signal corresponding to the target audio signal.
[0133] In this way, at least it can be achieved that for the target audio signal, the acquisition end can describe the spectral distribution characteristics of the high-frequency signal in it through spectral feature parameters with a very small data size. During transmission, only the spectral feature parameters and the low-frequency encoded data of the low-frequency signal need to be transmitted, effectively reducing the transmission bandwidth. At the same time, based on the spectral feature parameters, a matching audio prediction network is selected to restore the high-frequency signal and generate a predicted high-frequency signal. Since the general spectral distribution characteristics can be described through a very small data size, the error between the predicted high-frequency signal and the original high-frequency signal is controllable, making the generation of the audio output signal controllable. Furthermore, during the audio processing process, the overall coding rate is effectively reduced and the ability to restore the predicted high-frequency signal is strong, effectively reducing the transmission bandwidth of the audio data and ensuring the audio playback effect.
[0134] To facilitate better implementation of the audio processing method provided by the embodiments of the present application, the embodiments of the present application also provide an audio processing device based on the above audio processing method. The meanings of the terms are the same as those in the above audio processing method, and the specific implementation details can refer to the descriptions in the method embodiments. Figure 7 The block diagram of an audio processing device according to an embodiment of the present application is shown. Figure 8 The block diagram of an audio processing device according to another embodiment of the present application is shown.
[0135] As Figure 7 shown, the audio processing device 600 may include a receiving module 610, a decoding module 620, a matching module 630, a prediction module 640, and an output module 650. The audio processing device 600 can be applied to the device corresponding to the receiving end of the audio.
[0136] The receiving module 610 can be used to receive the spectral feature parameters of the high-frequency signal and the low-frequency encoded data of the low-frequency signal, where the high-frequency signal and the low-frequency signal belong to the target audio signal; the decoding module 620 can be used to perform decoding processing on the low-frequency encoded data to generate a decoded low-frequency signal; the matching module 630 can be used to perform prediction network matching processing based on the spectral feature parameters to obtain an audio prediction network that matches the spectral feature parameters; the prediction module 640 can be used to perform audio prediction processing based on the audio prediction network and the decoded low-frequency signal to generate a predicted high-frequency signal corresponding to the high-frequency signal; the output module 650 can be used to generate an audio output signal corresponding to the target audio signal according to the predicted high-frequency signal and the decoded low-frequency signal.
[0137] In some embodiments of the present application, the spectral feature parameters include a spectral envelope type; the matching module 630 includes: an information acquisition unit, configured to acquire network information of at least one preset audio prediction network, where each piece of network information corresponds to a preset spectral envelope type; a network matching unit, configured to determine the network information corresponding to the preset spectral envelope type that matches the spectral envelope type to obtain target network information; and a network determination unit, configured to determine the preset audio prediction network corresponding to the target network information as the audio prediction network that matches the spectral feature parameters.
[0138] In some embodiments of the present application, the prediction module 640 includes: an extraction processing unit, configured to perform spectral feature extraction processing on the decoded low-frequency signal to obtain low-frequency spectral information; an information prediction unit, configured to use the audio prediction network to perform audio prediction processing based on the low-frequency spectral information to obtain predicted spectral information; and a signal generation unit, configured to generate a predicted high-frequency signal corresponding to the high-frequency signal based on the predicted spectral information.
[0139] In some embodiments of the present application, the extraction processing unit is configured to: perform improved discrete cosine transform processing on the decoded low-frequency signal to obtain the low-frequency spectral information; the signal generation unit is configured to: perform inverse improved discrete cosine transform processing on the predicted spectral information to generate a predicted high-frequency signal corresponding to the high-frequency signal.
[0140] In some embodiments of the present application, the output module 650 is configured to: perform quadrature mirror synthesis filtering processing on the predicted high-frequency signal and the decoded low-frequency signal to generate the audio output signal.
[0141] In this way, based on the audio processing device 600, for the target audio signal, the spectral distribution characteristics of the high-frequency signal therein can be described by spectral feature parameters with a very small data size. When receiving data, only the spectral feature parameters and the low-frequency encoded data of the low-frequency signal need to be transmitted, effectively reducing the transmission bandwidth. At the same time, based on the spectral feature parameters, a matching audio prediction network is selected to restore the high-frequency signal and generate a predicted high-frequency signal. Since the general spectral distribution characteristics can be described by a very small data size, the error between the predicted high-frequency signal and the original high-frequency signal is controllable, making the generation of the audio output signal controllable. Furthermore, during the audio processing, the overall coding rate is effectively reduced and the ability to restore and predict the high-frequency signal is strong, effectively reducing the transmission bandwidth of the audio data and ensuring the audio playback effect.
[0142] As Figure 8 shown, the audio processing device 700 may include a decomposition module 710, an extraction module 720, an encoding module 730, and a delivery module 740. The audio processing device 700 can be applied to the device corresponding to the audio acquisition end.
[0143] The decomposition module 710 can be used to decompose the target audio signal to generate a high-frequency signal and a low-frequency signal; the extraction module 720 can be used to perform feature extraction processing on the high-frequency signal to obtain the spectral feature parameters corresponding to the high-frequency signal; the encoding module 730 can be used to perform audio encoding processing on the low-frequency signal to generate the low-frequency encoded data corresponding to the low-frequency signal; the delivery module 740 can be used to send the spectral feature parameters and the low-frequency encoded data to the receiving end, so that the receiving end determines the audio prediction network matching the spectral feature parameters, and generates an audio output signal based on the audio prediction network and the decoded low-frequency signal obtained by decoding the low-frequency encoded data.
[0144] In some embodiments of the present application, the extraction module 720 includes: a frequency domain conversion unit for performing frequency domain conversion processing on the high-frequency signal to obtain a frequency domain signal; a power spectral value calculation unit for calculating the power spectral values of each frequency point in the frequency domain signal; and a spectral feature parameter acquisition unit for performing feature extraction processing based on the power spectral values of each frequency point to obtain spectral feature parameters describing the spectral distribution characteristics of the high-frequency signal.
[0145] In some embodiments of the present application, the spectral feature parameter acquisition unit includes: an element calculation sub-unit for calculating the average value of the power spectral values of each frequency point and determining the maximum power spectral value among the power spectral values of each frequency point; a difference calculation sub-unit for performing a difference calculation on the maximum power spectral value and the average value to obtain a first difference; and a spectral feature parameter determination sub-unit for determining the spectral feature parameters corresponding to the high-frequency signal according to the first difference.
[0146] In some embodiments of the present application, the spectral feature parameter includes a spectral envelope type; the spectral feature parameter determination subunit is configured to: if the first difference is less than a first predetermined threshold value and the maximum power spectral value is less than a second predetermined threshold value, determine that the spectral envelope type corresponding to the high-frequency signal is a first type; if the first difference is less than the first predetermined threshold value and the maximum power spectral value is greater than the second predetermined threshold value, determine that the spectral envelope type corresponding to the high-frequency signal is a second type.
[0147] In some embodiments of the present application, the spectral feature parameter includes a spectral envelope type; the spectral feature parameter determination subunit is configured to: if the first difference is greater than a first predetermined threshold value, perform normalization processing on the power spectral values of each frequency point to obtain a normalized value corresponding to each frequency point; obtain at least one preset target value, where each target value corresponds to a preset spectral envelope type; calculate the mean square error value between the normalized value corresponding to each frequency point and each target value; and determine the preset spectral envelope type corresponding to the target value corresponding to the smallest mean square error value as the spectral envelope type of the high-frequency signal.
[0148] In some embodiments of the present application, the spectral feature parameter determination subunit is configured to: perform a difference operation on the power spectral values of each frequency point and the average value respectively to obtain a second difference corresponding to each frequency point; calculate the square value of the second difference corresponding to each frequency point, and calculate the average value of the square values to obtain a normalized score; and divide the second difference corresponding to each frequency point by the normalized score to obtain a normalized value corresponding to each frequency point.
[0149] In some embodiments of the present application, the decomposition module 710 is configured to: perform orthogonal mirror decomposition filtering on the target audio signal to generate the high-frequency signal and the low-frequency signal.
[0150] In this way, based on the audio processing device 700, for the target audio signal, the spectral distribution characteristics of the high-frequency signal therein can be described by spectral feature parameters with a very small data size. When transmitting data, only the spectral feature parameters and the low-frequency coding data of the low-frequency signal need to be transmitted, effectively reducing the transmission bandwidth. At the same time, based on the spectral feature parameters, a matching audio prediction network is selected to restore the high-frequency signal to generate a predicted high-frequency signal. Since the general spectral distribution characteristics can be described by a very small data size, the error between the predicted high-frequency signal and the original high-frequency signal is controllable, making the generation of the audio output signal controllable. Furthermore, the overall coding rate in the audio processing process is effectively reduced and the ability to restore the predicted high-frequency signal is strong, effectively reducing the transmission bandwidth of the audio data and ensuring the audio playback effect.
[0151] It should be noted that although several modules or units of the device for action execution are mentioned in the above detailed description, this division is not mandatory. In fact, according to the embodiments of the present application, the features and functions of the two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.
[0152] In addition, an embodiment of the present application further provides an electronic device, which can be a terminal or a server. For example, Figure 9 as shown, it shows a schematic structural diagram of the electronic device involved in the embodiments of the present application. Specifically:
[0153] The electronic device may include a processor 801 with one or more processing cores, a memory 802 with one or more computer-readable storage media, a power supply 803, an input unit 804, and other components. Those skilled in the art can understand that Figure 9 the structure of the electronic device shown in does not constitute a limitation on the electronic device, and it may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements. Among them:
[0154] The processor 801 is the control center of the electronic device, connecting various parts of the entire computer device through various interfaces and lines. By running or executing software programs and / or modules stored in the memory 802, and calling data stored in the memory 802, it executes various functions of the computer device and processes data, thereby controlling the electronic device as a whole. Optionally, the processor 801 may include one or more processing cores; preferably, the processor 801 may integrate an application processor and a modem processor. Among them, the application processor mainly processes the operating system, user interfaces, and application programs, etc., and the modem processor mainly processes wireless communications. It can be understood that the above-mentioned modem processor may not be integrated into the processor 801.
[0155] The memory 802 can be used to store software programs and modules. The processor 801 executes various functional applications and data processing by running the software programs and modules stored in the memory 802. The memory 802 mainly includes a program storage area and a data storage area. Among them, the program storage area can store the operating system, application programs required for at least one function (such as the sound playback function, the image playback function, etc.); the data storage area can store data created according to the use of the computer device. In addition, the memory 802 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other volatile solid-state storage devices. Correspondingly, the memory 802 may also include a memory controller to provide the processor 801 with access to the memory 802.
[0156] The electronic device further includes a power supply 803 for supplying power to each component. Preferably, the power supply 803 can be logically connected to the processor 801 through a power management system, so as to realize functions such as management of charging, discharging, and power consumption management through the power management system. The power supply 803 may also include any components such as one or more DC or AC power supplies, a recharge system, a power failure detection circuit, a power converter or inverter, and a power status indicator.
[0157] The electronic device may further include an input unit 804, which can be used to receive input digital or character information, and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function controls.
[0158] Although not shown, the electronic device may further include a display unit, etc., which will not be elaborated here. Specifically, in this embodiment, the processor 801 in the electronic device will load the executable files corresponding to the processes of one or more computer programs into the memory 802 according to the following instructions, and the processor 801 will run the computer programs stored in the memory 802 to implement various functions in the foregoing embodiments of the present application.
[0159] For example, the processor 801 can execute: receiving the spectral feature parameters of a high-frequency signal and the low-frequency coding data of a low-frequency signal, where the high-frequency signal and the low-frequency signal belong to a target audio signal; performing decoding processing on the low-frequency coding data to generate a decoded low-frequency signal; performing prediction network matching processing based on the spectral feature parameters to obtain an audio prediction network matching the spectral feature parameters; performing audio prediction processing based on the audio prediction network and the decoded low-frequency signal to generate a predicted high-frequency signal corresponding to the high-frequency signal; generating an audio output signal corresponding to the target audio signal according to the predicted high-frequency signal and the decoded low-frequency signal.
[0160] For example, the processor 801 may execute: decomposing the target audio signal to generate a high-frequency signal and a low-frequency signal; extracting features from the high-frequency signal to obtain spectral feature parameters corresponding to the high-frequency signal; performing audio encoding on the low-frequency signal to generate low-frequency encoded data corresponding to the low-frequency signal; and sending the spectral feature parameters and the low-frequency encoded data to a receiving end, so that the receiving end determines an audio prediction network that matches the spectral feature parameters, and generates an audio output signal based on the audio prediction network and a decoded low-frequency signal obtained by decoding the low-frequency encoded data.
[0161] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by a computer program or by controlling related hardware through a computer program. The computer program can be stored in a computer-readable storage medium and loaded and executed by a processor.
[0162] Therefore, an embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored, and the computer program can be loaded by a processor to execute the steps in any one of the methods provided by the embodiments of the present application.
[0163] Wherein, the computer-readable storage medium may include: a read-only memory (ROM, Read Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk, an optical disk, or the like.
[0164] Since the computer program stored in the computer-readable storage medium can execute the steps in any one of the methods provided by the embodiments of the present application, the beneficial effects that can be achieved by the methods provided by the embodiments of the present application can be realized. For details, see the previous embodiments and will not be elaborated here.
[0165] According to one aspect of the present application, there is provided a computer program product or a computer program, the computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the methods provided in various optional implementation manners in the above embodiments of the present application.
[0166] Those skilled in the art will readily think of other implementation manners of the present application after considering the specification and practicing the disclosed embodiments herein. The present application is intended to cover any variations, uses, or adaptations of the present application, which follow the general principles of the present application and include known common general knowledge or conventional technical means in the technical field not disclosed in the present application.
[0167] It should be understood that the present application is not limited to the embodiments described above and illustrated in the accompanying drawings, and various modifications and changes can be made without departing from its scope.
Claims
1. An audio processing method, characterized in that, Comprising: Receiving spectral characteristic parameters of a high-frequency signal and low-frequency encoded data of a low-frequency signal, where the high-frequency signal and the low-frequency signal belong to a target audio signal, the spectral characteristic parameters are information describing the spectral distribution characteristics of the high-frequency signal, and include a spectral envelope type; Performing decoding processing on the low-frequency encoded data to generate a decoded low-frequency signal; Performing prediction network matching processing based on the spectral characteristic parameters to obtain an audio prediction network matching the spectral characteristic parameters, where the audio prediction network matching the spectral characteristic parameters is a deep learning network trained based on training samples corresponding to the spectral characteristic parameters; Performing audio prediction processing based on the audio prediction network and the decoded low-frequency signal to generate a predicted high-frequency signal corresponding to the high-frequency signal; Generating an audio output signal corresponding to the target audio signal according to the predicted high-frequency signal and the decoded low-frequency signal; Wherein, the performing prediction network matching processing based on the spectral characteristic parameters to obtain an audio prediction network matching the spectral characteristic parameters includes: Obtaining network information of at least one preset audio prediction network, and each piece of network information corresponds to a preset spectral envelope type; Determining the network information corresponding to the preset spectral envelope type that matches the spectral envelope type to obtain target network information; Determining the preset audio prediction network corresponding to the target network information as the audio prediction network matching the spectral characteristic parameters.
2. The method according to claim 1, characterized in that, The performing audio prediction processing based on the audio prediction network and the decoded low-frequency signal to generate a predicted high-frequency signal corresponding to the high-frequency signal includes: Performing spectral characteristic extraction processing on the decoded low-frequency signal to obtain low-frequency spectral information; Using the audio prediction network to perform audio prediction processing based on the low-frequency spectral information to obtain predicted spectral information; Generating a predicted high-frequency signal corresponding to the high-frequency signal based on the predicted spectral information.
3. The method according to claim 2, wherein The performing spectral characteristic extraction processing on the decoded low-frequency signal to obtain low-frequency spectral information includes: Performing improved discrete cosine transform processing on the decoded low-frequency signal to obtain the low-frequency spectral information; The generating a predicted high-frequency signal corresponding to the high-frequency signal based on the predicted spectral information includes: Performing inverse improved discrete cosine transform processing on the predicted spectral information to generate a predicted high-frequency signal corresponding to the high-frequency signal.
4. The method according to any one of claims 1 to 3, characterized in that The generating an audio output signal corresponding to the target audio signal according to the predicted high-frequency signal and the decoded low-frequency signal includes: Performing quadrature mirror synthesis filtering processing on the predicted high-frequency signal and the decoded low-frequency signal to generate the audio output signal.
5. An audio processing method, characterized in that, Comprising: Decomposing a target audio signal to generate a high-frequency signal and a low-frequency signal; Performing feature extraction processing on the high-frequency signal to obtain spectral characteristic parameters corresponding to the high-frequency signal, where the spectral characteristic parameters are information describing the spectral distribution characteristics of the high-frequency signal, and include a spectral envelope type; Performing audio encoding processing on the low-frequency signal to generate low-frequency encoded data corresponding to the low-frequency signal; Send the spectral feature parameters and the low-frequency encoded data to a receiving end, so that the receiving end determines an audio prediction network that matches the spectral feature parameters, and generates an audio output signal based on the audio prediction network and a decoded low-frequency signal obtained by decoding the low-frequency encoded data, where the audio prediction network that matches the spectral feature parameters is a deep learning network trained based on training samples corresponding to the spectral feature parameters; Among them, determining the audio prediction network that matches the spectral feature parameters includes: Obtain network information of at least one preset audio prediction network, and each piece of network information corresponds to a preset spectral envelope type; Determine network information corresponding to the preset spectral envelope type that matches the spectral envelope type to obtain target network information; Determine the preset audio prediction network corresponding to the target network information as the audio prediction network that matches the spectral feature parameters.
6. The method according to claim 5, wherein The performing feature extraction processing on the high-frequency signal to obtain spectral feature parameters corresponding to the high-frequency signal includes: Perform frequency-domain conversion processing on the high-frequency signal to obtain a frequency-domain signal; Calculate power spectral values of each frequency point in the frequency-domain signal; Perform feature extraction processing based on the power spectral values of each frequency point to obtain spectral feature parameters describing the spectral distribution characteristics of the high-frequency signal.
7. The method according to claim 6, wherein The performing feature extraction processing based on the power spectral values of each frequency point to obtain spectral feature parameters describing the spectral distribution characteristics of the high-frequency signal includes: Calculate an average value of the power spectral values of each frequency point, and determine a maximum power spectral value among the power spectral values of each frequency point; Perform a difference operation on the maximum power spectral value and the average value to obtain a first difference; Determine spectral feature parameters corresponding to the high-frequency signal according to the first difference.
8. The method according to claim 7, characterized in that The determining spectral feature parameters corresponding to the high-frequency signal according to the first difference includes: If the first difference is less than a first predetermined threshold and the maximum power spectral value is less than a second predetermined threshold, determine that the spectral envelope type corresponding to the high-frequency signal is a first type; If the first difference is less than a first predetermined threshold and the maximum power spectral value is greater than a second predetermined threshold, determine that the spectral envelope type corresponding to the high-frequency signal is a second type.
9. The method according to claim 7, wherein The determining spectral feature parameters corresponding to the high-frequency signal according to the first difference includes: If the first difference is greater than a first predetermined threshold, perform normalization processing on the power spectral values of each frequency point to obtain a normalized value corresponding to each frequency point; Obtain at least one preset target value, and each target value corresponds to a preset spectral envelope type; Calculate a mean square error value between the normalized value corresponding to each frequency point and each target value; Determine the preset spectral envelope type corresponding to the target value corresponding to the smallest mean square error value as the spectral envelope type of the high-frequency signal.
10. The method according to claim 9, wherein The performing normalization processing on the power spectral values of each frequency point to obtain a normalized value corresponding to each frequency point includes: Perform a difference operation on the power spectral values of each frequency point and the average value respectively to obtain a second difference corresponding to each frequency point; Calculate the square value of the second difference corresponding to each of the frequency points, and calculate the average value of the square values to obtain a normalized score; Divide the second difference corresponding to each of the frequency points by the normalized score to obtain a normalized value corresponding to each of the frequency points.
11. The method according to any one of claims 5 to 10, characterized in that, The decomposing the target audio signal to generate a high-frequency signal and a low-frequency signal includes: Performing an orthogonal mirror decomposition filtering process on the target audio signal to generate the high-frequency signal and the low-frequency signal.
12. An audio processing device, characterized in that, Includes: A receiving module, configured to receive spectral feature parameters of a high-frequency signal and low-frequency coding data of a low-frequency signal, where the high-frequency signal and the low-frequency signal belong to a target audio signal, and the spectral feature parameters are information describing the spectral distribution characteristics of the high-frequency signal and include a spectral envelope type; A decoding module, configured to perform a decoding process on the low-frequency coding data to generate a decoded low-frequency signal; A matching module, configured to perform a prediction network matching process based on the spectral feature parameters to obtain an audio prediction network matching the spectral feature parameters, where the audio prediction network matching the spectral feature parameters is a deep learning network trained based on training samples corresponding to the spectral feature parameters; A prediction module, configured to perform an audio prediction process based on the audio prediction network and the decoded low-frequency signal to generate a predicted high-frequency signal corresponding to the high-frequency signal; An output module, configured to generate an audio output signal corresponding to the target audio signal according to the predicted high-frequency signal and the decoded low-frequency signal; Wherein, the performing a prediction network matching process based on the spectral feature parameters to obtain an audio prediction network matching the spectral feature parameters includes: Obtaining network information of at least one preset audio prediction network, where each network information corresponds to a preset spectral envelope type; Determining the network information corresponding to the preset spectral envelope type that matches the spectral envelope type to obtain target network information; Determining the preset audio prediction network corresponding to the target network information as the audio prediction network matching the spectral feature parameters.
13. An audio processing device, characterized in that, Includes: A decomposing module, configured to decompose a target audio signal to generate a high-frequency signal and a low-frequency signal; An extracting module, configured to perform a feature extraction process on the high-frequency signal to obtain spectral feature parameters corresponding to the high-frequency signal, where the spectral feature parameters are information describing the spectral distribution characteristics of the high-frequency signal and include a spectral envelope type; An encoding module, configured to perform an audio encoding process on the low-frequency signal to generate low-frequency coding data corresponding to the low-frequency signal; A transmitting module, configured to send the spectral feature parameters and the low-frequency coding data to a receiving end, so that the receiving end determines an audio prediction network matching the spectral feature parameters, and generates an audio output signal based on the audio prediction network and the decoded low-frequency signal obtained by decoding the low-frequency coding data, where the audio prediction network matching the spectral feature parameters is a deep learning network trained based on training samples corresponding to the spectral feature parameters; Wherein, the determining the audio prediction network matching the spectral feature parameters includes: Obtain network information of at least one preset audio prediction network, where each piece of the network information corresponds to a preset spectral envelope type; Determine the network information corresponding to the preset spectral envelope type that matches the spectral envelope type to obtain target network information; Determine the preset audio prediction network corresponding to the target network information as the audio prediction network that matches the spectral feature parameters.
14. A computer-readable storage medium, characterized in that, A computer program is stored thereon, and when the computer program is executed by a processor of a computer, the computer is caused to execute the method according to any one of claims 1 to 4 and 5 to 11.
15. An electronic device, characterized in that, Comprising: A memory storing a computer program; A processor that reads the computer program stored in the memory to execute the method according to any one of claims 1 to 4 and 5 to 11.
16. A computer program product, characterized in that, The computer program product includes a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 4 and 5 to 11 is implemented.
Citation Information
Patent Citations
Audio encoding and decoding method and device, medium and electronic equipment
CN112767954A
Blind Bandwidth Extension using K-Means and a Support Vector Machine
US20180040336A1