Audio signal enhancement method and device, computer device and storage medium
By decoding the speech packets and extracting their feature parameters, converting them into filter speech excitation signals, and then enhancing them, the problems of large delay and limited effectiveness of traditional audio signal enhancement methods are solved, achieving efficient audio signal enhancement.
Patent Information
- Application Number
- CN202110484196.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-04-30
- Publication Date
- 2026-02-27
- Estimated Expiration
- 2041-10-06
AI Technical Summary
Traditional audio signal enhancement methods suffer from large time delays and limited improvement in speech quality, resulting in poor timeliness of audio signal enhancement.
The received voice packets are decoded to obtain residual signals, long-term filtering parameters, and linear filtering parameters. Feature parameters are extracted and converted into filter voice excitation signals. Then, voice enhancement processing is performed using linear filtering parameters and feature parameters, and finally, voice synthesis is performed to enhance the audio signal.
It can effectively enhance audio signals in a shorter time, improving the timeliness and quality of audio signal enhancement.
Smart Images

Figure CN113763973B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of computer, in particular to an audio signal enhancement method and device, computer equipment and storage medium. BACKGROUND
[0002] In the process of audio signal coding and decoding, quantization noise is usually introduced, which causes the synthesized speech to be distorted. In the traditional scheme, pitch filter or neural network-based post-processing technology is usually used to enhance the audio signal to reduce the influence of quantization noise on the quality of the speech.
[0003] However, the speed of signal processing in the traditional scheme is low, there is a large time delay, and the improvement effect of the speech quality that can be achieved is limited, resulting in poor timeliness of audio signal enhancement. SUMMARY
[0004] Therefore, it is necessary to provide an audio signal enhancement method and device, computer equipment and storage medium capable of improving the timeliness of audio signal enhancement.
[0005] An audio signal enhancement method, the method comprising:
[0006] sequentially decoding a received speech packet to obtain a residual signal, a long-time filtering parameter and a linear filtering parameter; filtering the residual signal to obtain an audio signal;
[0007] extracting a feature parameter from the audio signal when the audio signal is a forward error correction frame signal;
[0008] converting the audio signal into a filter speech excitation signal based on the linear filtering parameter;
[0009] performing speech enhancement processing on the filter speech excitation signal according to the feature parameter, the long-time filtering parameter and the linear filtering parameter to obtain an enhanced speech excitation signal;
[0010] performing speech synthesis based on the enhanced speech excitation signal and the linear filtering parameter to obtain a speech enhancement signal.
[0011] In one embodiment, the linear filtering parameter comprises a linear filtering coefficient and an energy gain value; the parameter configuration of the linear prediction filter based on the linear filtering parameter comprises:
[0012] configuring the linear prediction filter based on the linear filtering coefficient;
[0013] obtaining an energy gain value corresponding to a historical speech packet decoded before decoding the speech packet;
[0014] determining an energy adjustment parameter based on the energy gain value corresponding to the historical speech packet and the energy gain value corresponding to the speech packet;
[0015] performing energy adjustment on a historical long-term filtered excitation signal corresponding to the historical speech packet by the energy adjustment parameter to obtain an adjusted historical long-term filtered excitation signal;
[0016] inputting the adjusted historical long-term filtered excitation signal and the enhanced speech excitation signal into a linear prediction filter configured with parameters, so that the linear prediction filter performs linear synthesis filtering on the enhanced speech excitation signal based on the adjusted historical long-term filtered excitation signal.
[0017] An audio signal enhancement device, the device comprising:
[0018] a speech packet processing module, configured to sequentially decode a received speech packet to obtain a residual signal, a long-term filtering parameter and a linear filtering parameter, and filter the residual signal to obtain an audio signal;
[0019] a feature parameter extraction module, configured to extract a feature parameter from the audio signal when the audio signal is a forward error correction frame signal;
[0020] a signal conversion module, configured to convert the audio signal into a filter speech excitation signal based on the linear filtering parameter;
[0021] a speech enhancement module, configured to perform speech enhancement processing on the filter speech excitation signal according to the feature parameter, the long-term filtering parameter and the linear filtering parameter to obtain an enhanced speech excitation signal;
[0022] a speech synthesis module, configured to perform speech synthesis based on the enhanced speech excitation signal and the linear filtering parameter to obtain a speech enhancement signal.
[0023] A computer device, comprising a memory and a processor, the memory storing a computer program, and the processor implementing the following steps when executing the computer program:
[0024] sequentially decoding a received speech packet to obtain a residual signal, a long-term filtering parameter and a linear filtering parameter, and filtering the residual signal to obtain an audio signal;
[0025] extracting a feature parameter from the audio signal when the audio signal is a forward error correction frame signal;
[0026] convert the audio signal into a filter speech excitation signal based on the linear filter parameter;
[0027] perform speech enhancement processing on the filter speech excitation signal according to the feature parameter, the long-term filter parameter and the linear filter parameter, to obtain an enhanced speech excitation signal;
[0028] perform speech synthesis based on the enhanced speech excitation signal and the linear filter parameter, to obtain a speech enhanced signal.
[0029] A computer readable storage medium having stored thereon a computer program, the computer program being executed by a processor to implement the following steps:
[0030] decode the received speech packet in sequence to obtain a residual signal, a long-term filter parameter and a linear filter parameter; filter the residual signal to obtain an audio signal;
[0031] extract a feature parameter from the audio signal when the audio signal is a forward error correction frame signal;
[0032] convert the audio signal into a filter speech excitation signal based on the linear filter parameter;
[0033] perform speech enhancement processing on the filter speech excitation signal according to the feature parameter, the long-term filter parameter and the linear filter parameter, to obtain an enhanced speech excitation signal;
[0034] perform speech synthesis based on the enhanced speech excitation signal and the linear filter parameter, to obtain a speech enhanced signal.
[0035] A computer program, the computer program comprising computer instructions stored in a computer readable storage medium, a processor of a computer device reading the computer instructions from the computer readable storage medium, the processor executing the computer instructions to cause the computer device to perform the following steps:
[0036] decode the received speech packet in sequence to obtain a residual signal, a long-term filter parameter and a linear filter parameter; filter the residual signal to obtain an audio signal;
[0037] extract a feature parameter from the audio signal when the audio signal is a forward error correction frame signal;
[0038] convert the audio signal into a filter speech excitation signal based on the linear filter parameter;
[0039] According to the feature parameter, the long-time filtering parameter and the linear filtering parameter, a speech enhancement processing is performed on the filter speech excitation signal to obtain an enhanced speech excitation signal.
[0040] Based on the enhanced speech excitation signal and the linear filtering parameter, a speech synthesis is performed to obtain a speech enhancement signal.
[0041] The audio signal enhancement method, device, computer device and storage medium, by sequentially decoding the received speech packet to obtain a residual signal, a long-time filtering parameter and a linear filtering parameter, filtering the residual signal to obtain an audio signal, and when the audio signal is a forward error correction frame signal, extracting a feature parameter from the audio signal, converting the audio signal into a filter speech excitation signal based on the linear filtering coefficient obtained from the decoded speech packet, and performing a speech enhancement processing on the filter speech excitation signal according to the feature parameter and the long-time filtering parameter and the linear filtering parameter obtained from the decoded speech packet to obtain an enhanced speech excitation signal, and performing a speech synthesis based on the enhanced speech excitation signal and the linear filtering parameter to obtain a speech enhancement signal, thereby completing the enhancement processing of the audio signal in a short time and achieving a good signal enhancement effect, and improving the timeliness of the audio signal enhancement. BRIEF DESCRIPTION OF DRAWINGS
[0042] Figure 1 A speech generation model based on an excitation signal in an embodiment;
[0043] Figure 2 An application environment diagram of the audio signal enhancement method in an embodiment;
[0044] Figure 3 A flowchart of the audio signal enhancement method in an embodiment;
[0045] Figure 4 An audio signal transmission flowchart in an embodiment;
[0046] Figure 5 An amplitude-frequency response diagram of a long-time prediction filter in an embodiment;
[0047] Figure 6 A flowchart of a speech packet decoding and filtering step in an embodiment;
[0048] Figure 7 An amplitude-frequency response diagram of a long-time inverse filter in an embodiment;
[0049] Figure 8 A signal enhancement model diagram in an embodiment;
[0050] Figure 9 A flowchart of the audio signal enhancement method in another embodiment;
[0051] Figure 10 Flowchart of the audio signal enhancement method in another embodiment;
[0052] Figure 11 Block diagram of the audio signal enhancement device in an embodiment;
[0053] Figure 12 Block diagram of the audio signal enhancement device in another embodiment;
[0054] Figure 13 Internal structure diagram of the computer device in an embodiment;
[0055] Figure 14 Internal structure diagram of the computer device in another embodiment. DETAILED DESCRIPTION
[0056] In order to make the purposes, technical solutions and advantages of the present application clearer, the present application is further described in detail below with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and do not limit the present application.
[0057] Before the audio signal enhancement method provided by the present application is described, the speech generation model is described first, referring to the speech generation model based on excitation signal shown in Figure 1 The physical theoretical basis of the speech generation model based on excitation signal is the occurrence process of human voice, which includes:
[0058] (1) At the trachea, an impact signal with certain energy and noise-like is generated, which corresponds to the excitation signal in the speech generation model based on excitation signal.
[0059] (2) The impact signal impacts the vocal cords of the human, generating a periodic opening and closing, and after amplification through the oral cavity, the sound is emitted, which corresponds to the filter in the speech generation model based on excitation signal.
[0060] In the actual process, considering the characteristics of the sound, the filter in the speech generation model based on excitation signal is subdivided into a long-term prediction (LTP) filter and a linear predictive coding (LPC) filter, wherein the LTP filter is used to strengthen the audio signal by using the long-term correlation of the speech, and the LPC filter is used to strengthen the audio signal by using the short-term correlation of the speech. Specifically, for voiced signals which are periodic signals, in the speech generation model based on excitation signal, the excitation signal will impact the LTP filter and the LPC filter respectively; for unvoiced signals which are non-periodic signals, the excitation signal will only impact the LPC filter.
[0061] The audio signal enhancement method provided in the application can be implemented based on cloud technology. The cloud technology refers to a hosting technology of unifying a series of resources such as hardware, software and network in a wide area network or a local area network to realize data calculation, storage, processing and sharing. The cloud technology is a general term of network technology, information technology, integration technology, management platform technology and application technology applied based on a cloud computing business model, can form a resource pool, and is used on demand, flexibly and conveniently. The cloud computing technology will become an important support. The background service of a technical network system needs a large amount of calculation and storage resources, such as a video website, a picture website and more portals. With the high development and application of the Internet industry, in the future, each item may have its own identification mark, and needs to be transmitted to the background system for logical processing. Different levels of data will be processed separately, and the data of various industries all need strong system support, which can only be realized through cloud computing.
[0062] The audio signal processed by the audio signal enhancement method provided in the application can be an audio signal generated in a cloud conference process. The cloud conference is an efficient, convenient and low-cost conference form based on cloud computing technology. Users only need to perform simple and easy-to-use operations through an Internet interface, and can quickly and efficiently share voice, data files and video with teams and customers all over the world, and the complex technologies such as data transmission and processing in the conference are operated by cloud conference service providers to help users. At present, domestic cloud conferences mainly focus on the service content of the SaaS (Software as a Service) mode, including telephone, network, video and other service forms, and the video conference based on cloud computing is called cloud conference. In the cloud conference era, the transmission, processing and storage of data are all processed by the computer resources of the video conference manufacturer, and users no longer need to purchase expensive hardware and install complicated software. They only need to open a browser and log in to the corresponding interface to perform efficient remote conferences.
[0063] Artificial intelligence (AI) is to use a digital computer or a machine controlled by a digital computer to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use the knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology of computer science, which tries to understand the essence of intelligence and produce a new intelligent machine that can react in a similar way to human intelligence. Artificial intelligence is to study the design principles and implementation methods of various intelligent machines, so that the machine has the functions of perception, reasoning and decision-making.
[0064] Machine Learning (ML) is a multi-disciplinary subject, involving probability theory, statistics, approximation theory, convex analysis, algorithm complexity theory, etc. It is a specialized study of how computers simulate or implement human learning behavior to acquire new knowledge or skills, reorganize existing knowledge structure to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental approach to making computers intelligent, and its applications are widespread in various fields of artificial intelligence. Machine learning and deep learning usually include artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and rule-based learning.
[0065] The scheme provided by the embodiments of the present application relates to machine learning and other technologies of artificial intelligence. Specifically, the audio signal enhancement method provided by the present application can be applied to the application environment as shown in Figure 2 The terminal 202 communicates with the server 204 through the network, and the terminal 202 can receive the voice packet sent by the server 204 or the voice packet forwarded by other devices through the server 204, and the server 204 can receive the voice packet sent by the terminal or the voice packet sent by other devices. The above-mentioned audio signal enhancement method can be applied to the terminal 202 or the server 204, and will be described taking the terminal 202 as an example. The terminal 202 sequentially decodes the received voice packet to obtain a residual signal, a long-time filtering parameter and a linear filtering parameter, filters the residual signal to obtain an audio signal, extracts a feature parameter from the audio signal when the audio signal is a forward error correction frame signal, converts the audio signal into a filter speech excitation signal based on the linear filtering parameter, performs speech enhancement processing on the filter speech excitation signal according to the feature parameter, the long-time filtering parameter and the linear filtering parameter to obtain an enhanced speech excitation signal, and performs speech synthesis based on the enhanced speech excitation signal and the linear filtering parameter to obtain a speech enhancement signal.
[0066] The terminal 202 can be, but is not limited to, various personal computers, notebook computers, smart phones, tablet computers and portable wearable devices, and the server 204 can be a standalone physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDNs, and basic cloud computing services such as big data and artificial intelligence platforms.
[0067] In one embodiment, as shown in Figure 3 An audio signal enhancement method is provided, which will be described taking the computer device (terminal or server) in Figure 2 as an example, including the following steps:
[0068] S302 decodes the received voice packets sequentially to obtain the residual signal, long-time filtering parameters, and linear filtering parameters; then filters the residual signal to obtain the audio signal.
[0069] The received voice packets can be voice packets in a packet loss-resistant scenario based on Feedforward Error Correction (FEC) technology.
[0070] Forward error correction is an error control method that involves encoding a signal according to a certain algorithm before it is sent into the transmission channel, adding redundant codes with the characteristics of the signal itself, and then decoding the received signal at the receiving end according to the corresponding algorithm to find and correct the error codes generated during transmission.
[0071] Redundant codes, also known as redundant information, are referenced in this embodiment of the application. Figure 4 When encoding the audio signal of the current voice frame (hereinafter referred to as the current frame), the signal transmitting end can encode the audio signal information of the previous voice frame (hereinafter referred to as the previous frame) as redundant information into the voice packet corresponding to the current frame audio signal. After encoding, the voice packet corresponding to the current frame audio signal is sent to the receiving end. The receiving end receives this voice packet. In this way, even if a fault occurs during signal transmission, causing the receiving end to miss a certain voice packet or have a bit error in a certain voice packet, the audio signal corresponding to the lost or erroneous voice packet can be obtained by decoding the voice packet corresponding to the audio signal of the next voice frame (hereinafter referred to as the next frame), thereby improving the reliability of signal transmission. The receiving end can be... Figure 2 Terminal 202 in the middle.
[0072] Specifically, when the terminal receives a voice packet, it stores the received voice packet in a buffer, then retrieves the voice packet corresponding to the voice frame to be played from the buffer, and decodes and filters the voice packet to obtain an audio signal. If the voice packet is an adjacent packet of the previously decoded historical voice packet and the previously decoded historical voice packet is normal, the obtained audio signal is directly output, or the audio signal is enhanced to obtain a voice enhancement signal, which is then output. If the voice packet is not an adjacent packet of the previously decoded historical voice packet, or if the voice packet is an adjacent packet of the previously decoded historical voice packet but the previously decoded historical voice packet is abnormal, the audio signal is enhanced to obtain a voice enhancement signal, which is then output. The voice enhancement signal carries the audio signal corresponding to the adjacent packet of the previously decoded historical voice packet.
[0073] The decoding can be specifically entropy decoding, which is a decoding scheme corresponding to entropy encoding. Specifically, the sender can encode the audio signal by using an entropy encoding scheme to obtain a speech packet, so that the receiver can decode the received speech packet by using an entropy decoding scheme.
[0074] In one embodiment, the terminal decodes the received speech packet to obtain a residual signal and filter parameters, performs signal synthesis filtering on the residual signal based on the filter parameters, and obtains the audio signal. The filter parameters include long-time filter parameters and linear filter parameters.
[0075] Specifically, when encoding the current frame of audio signal, the sender obtains filter parameters by analyzing the previous frame of audio signal, configures the filter based on the obtained filter parameters, then analyzes and filters the current frame of audio signal by using the configured filter to obtain a residual signal of the current frame of audio signal, and encodes the audio signal by using the residual signal and the analyzed filter parameters to obtain a speech packet, and sends the speech packet to the receiver. Thus, after receiving the speech packet, the receiver decodes the received speech packet to obtain a residual signal and filter parameters, performs signal synthesis filtering on the residual signal based on the filter parameters, and obtains the audio signal.
[0076] In one embodiment, the filter parameters include linear filter parameters and long-time filter parameters. When encoding the current frame of audio signal, the sender obtains linear filter parameters and long-time filter parameters by analyzing the previous frame of audio signal, then performs linear analysis filtering on the current frame of audio signal based on the linear filter parameters to obtain a linear filter excitation signal, performs long-time analysis filtering on the linear filter excitation signal based on the long-time filter parameters to obtain a residual signal corresponding to the current frame of audio signal, and encodes the current frame of audio signal by using the residual signal, the analyzed linear filter parameters and the long-time filter parameters to obtain a speech packet, and sends the speech packet to the receiver.
[0077] Specifically, performing linear analysis filtering on the current frame of audio signal based on the linear filter parameters specifically includes: configuring a linear prediction filter based on the linear filter parameters, and performing linear analysis filtering on the audio signal by using the configured linear prediction filter to obtain a linear filter excitation signal. The linear filter parameters include linear filter coefficients and energy gain values. The linear filter coefficients can be denoted as LPC AR, and the energy gain values can be denoted as LPC gain. The formula of the linear prediction filter is as follows:
[0078]
[0079] wherein e(n) is a linear filter excitation signal corresponding to the current frame audio signal, s(n) is the current frame audio signal, p is the number of sampling points contained in each frame of audio signal, a i is a linear filter coefficient obtained by analyzing the previous frame audio signal, s adj (n-i) is the energy-adjusted state of the previous frame audio signal s(n-i) of the current frame audio signal s(n), s adj (n-i) can be obtained by the following formula:
[0080] s adj (n-i) = gain adj ·s(n-i) (2)
[0081] wherein s(n-i) is the previous frame audio signal of the current frame audio signal s(n), gain adj is an energy adjustment parameter of the previous frame audio signal s(n-i), gain adj can be obtained by the following formula:
[0082]
[0083] wherein gain(n) is an energy gain value corresponding to the current frame audio signal, and gain(n-i) is an energy gain value corresponding to the previous frame audio signal.
[0084] The long-time analysis filtering of the linear filter excitation signal based on the long-time filter parameter specifically includes: parameter configuration of a long-time prediction filter based on the long-time filter parameter, long-time analysis filtering of a residual signal by the long-time prediction filter after the parameter configuration, to obtain a corresponding residual signal of the current frame audio signal, wherein the long-time filter parameter includes a pitch and a corresponding amplitude gain value, the pitch can be denoted as LTP pitch, and the corresponding amplitude gain value can be denoted as LTP gain, and the frequency domain representation of the long-time prediction filter is as follows, and the frequency domain can be denoted as Z domain:
[0085] p(z) = 1 - γz -T (4)
[0086] In the above formula, p(z) is the amplitude-frequency response of the long-time prediction filter, z is a rotation factor of frequency domain transformation, γ is the amplitude gain value LTP gain, and T is the pitch LTP pitch, Figure 5 The amplitude-frequency response diagram of the long-time prediction filter corresponding to γ = 1 and T = 80 in an embodiment is shown.
[0087] The time domain representation of the long-time prediction filter is as follows:
[0088] δ(n) = e(n) - γe(n-T) (5)
[0089] wherein, δ(n) is a residual signal corresponding to the current frame of audio signal, e(n) is a linear filtered excitation signal corresponding to the current frame of audio signal, γ is an LTP gain, T is an LTP pitch, e(n-T) is a linear filtered excitation signal corresponding to the audio signal of the previous LTP pitch of the current frame of audio signal.
[0090] In one embodiment, the filter parameters decoded by the terminal include long-term filter parameters and linear filter parameters, and the signal synthesis filtering includes long-term synthesis filtering based on the long-term filter parameters and linear synthesis filtering based on the linear filter parameters. After the terminal decodes the residual signal, the long-term filter parameters and the linear filter parameters for the speech packet, the terminal performs long-term synthesis filtering on the residual signal based on the long-term filter parameters to obtain a long-term filtered excitation signal, and then performs linear synthesis filtering on the long-term filtered excitation signal based on the linear filter parameters to obtain the audio signal.
[0091] In one embodiment, after the terminal obtains the residual signal, the terminal divides the obtained residual signal into a plurality of sub-frames to obtain a plurality of sub-residual signals, and for each sub-residual signal, performs long-term synthesis filtering thereon based on the corresponding long-term filter parameters to obtain a long-term filtered excitation signal corresponding to each sub-frame, and then combines the long-term filtered excitation signals corresponding to each sub-frame in the time sequence of the sub-frames to obtain the corresponding long-term filtered excitation signal.
[0092] For example, a speech packet corresponds to an audio signal of 20 ms, i.e., the obtained residual signal is 20 ms, and the residual signal can be divided into 4 sub-frames to obtain 4 sub-residual signals of 5 ms. For each sub-residual signal of 5 ms, long-term synthesis filtering is performed thereon based on the corresponding long-term filter parameters to obtain a long-term filtered excitation signal of 5 ms. Then, the 4 long-term filtered excitation signals of 5 ms are combined in the time sequence of the sub-frames to obtain a long-term filtered excitation signal of 20 ms.
[0093] In one embodiment, after the terminal obtains the long-term filtered excitation signal, the terminal divides the obtained long-term filtered excitation signal into a plurality of sub-frames to obtain a plurality of sub-long-term filtered excitation signals, and for each sub-long-term filtered excitation signal, performs linear synthesis filtering thereon based on the corresponding linear filter parameters to obtain a sub-linear filtered excitation signal corresponding to each sub-frame, and then combines the linear filtered excitation signals corresponding to each sub-frame in the time sequence of the sub-frames to obtain the corresponding linear filtered excitation signal.
[0094] For example, one speech packet corresponds to 20 ms of audio signal, i.e. the resulting long-time filtering excitation signal is 20 ms, the long-time filtering excitation signal can be divided into 2 sub-frames to obtain 2 sub-long-time filtering excitation signals of 10 ms, and linear synthesis filtering is performed on each 10 ms sub-long-time filtering excitation signal based on the corresponding linear filtering parameters to obtain 2 sub-audio signals of 10 ms, and then the 2 sub-audio signals of 10 ms are combined according to the time sequence of each sub-frame to obtain an audio signal of 20 ms.
[0095] S304, when the audio signal is a forward error correction frame signal, extracting a feature parameter from the audio signal.
[0096] The audio signal is a forward error correction frame signal, which means that the audio signal of the historical adjacent frame of the audio signal is abnormal. The abnormality of the audio signal of the historical adjacent frame specifically includes: the speech packet corresponding to the audio signal of the historical adjacent frame is not received, or the speech packet corresponding to the audio signal of the historical adjacent frame cannot be normally decoded. The feature parameter includes a cepstrum feature parameter.
[0097] In one embodiment, after the terminal decodes and filters the received speech packet to obtain an audio signal, it is determined whether the historical speech packet decoded before decoding the speech packet has data abnormality. If the decoded historical speech packet has data abnormality, it is determined that the audio signal obtained by decoding and filtering is a forward error correction frame signal.
[0098] Specifically, the terminal determines whether the historical audio signal corresponding to the historical speech packet decoded at the previous time of decoding the speech packet is the previous frame audio signal of the audio signal obtained by decoding the speech packet. If yes, it is determined that the historical speech packet has no data abnormality. If no, it is determined that the historical speech packet has data abnormality.
[0099] In this embodiment, the terminal determines whether the audio signal obtained by decoding and filtering is a forward error correction frame signal by determining whether the historical speech packet decoded before decoding the current speech packet has data abnormality, and then can perform audio signal enhancement processing on the audio signal when the audio signal is a forward error correction frame signal, to further improve the quality of the audio signal.
[0100] In one embodiment, when the decoded audio signal is a forward error correction frame signal, a feature parameter is extracted from the decoded audio signal. The extracted feature parameter can be a cepstrum feature parameter, which specifically includes the following steps: performing Fourier transform on the audio signal to obtain a Fourier transformed audio signal; performing logarithmic processing on the Fourier transformed audio signal to obtain a logarithmic result; and performing inverse Fourier transform on the obtained logarithmic result to obtain a cepstrum feature parameter. The cepstrum feature parameter can be extracted from the audio signal by the following formula:
[0101]
[0102] wherein C(n) is the cepstrum feature parameter of the audio signal S(n) obtained after decoding and filtering, and S(F) is the Fourier transformed audio signal obtained by performing Fourier transform on the audio signal S(n).
[0103] In the above embodiment, the terminal extracts the cepstrum feature parameter from the audio signal, so that the audio signal can be enhanced based on the extracted cepstrum feature parameter, and the quality of the audio signal is improved.
[0104] In one embodiment, when the audio signal is not a forward error correction frame signal, i.e., the previous frame of the audio signal obtained after decoding and filtering is not abnormal, the feature parameter can also be extracted from the current audio signal obtained after decoding and filtering, so as to perform audio signal enhancement processing on the current audio signal obtained after decoding and filtering.
[0105] S306, converting the audio signal into a filter speech excitation signal based on the linear filter parameter.
[0106] Specifically, after obtaining the audio signal by decoding and filtering the speech packet, the terminal can also obtain the linear filter parameter obtained when decoding the speech packet, and perform linear analysis filtering on the obtained audio signal based on the linear filter parameter, so as to convert the audio signal into a filter speech excitation signal.
[0107] In one embodiment, S306 specifically includes the following steps: performing parameter configuration on a linear prediction filter based on the linear filter parameter, and performing linear decomposition filtering on the audio signal by the parameter configured linear prediction filter to obtain the filter speech excitation signal.
[0108] wherein the linear decomposition filtering is also called linear analysis filtering, and in the embodiment of the application, the linear analysis filtering is directly performed on the whole frame of audio signal without performing sub-frame processing on the whole frame of audio signal.
[0109] Specifically, the terminal can perform linear decomposition filtering on the audio signal by using the following formula to obtain the filter speech excitation signal:
[0110]
[0111] wherein D(n) is the filter speech excitation signal corresponding to the audio signal S(n) obtained after decoding and filtering the speech packet, S(n) is the audio signal obtained after decoding and filtering the speech packet, S adj(n-i) is an energy-adjusted state of a previous frame of the obtained audio signal S(n), p is a number of sampling points contained in each frame of the audio signal, A i is a linear filter coefficient obtained by decoding a speech packet.
[0112] In the above embodiment, the terminal converts the audio signal into the filter speech excitation signal based on the linear filter parameter, so that the enhancement of the audio signal can be realized by enhancing the filter speech excitation signal, and the quality of the audio signal is improved.
[0113] S308, performing speech enhancement processing on the filter speech excitation signal according to the feature parameter, the long-term filter parameter and the linear filter parameter, to obtain an enhanced speech excitation signal.
[0114] The long-term filter parameter includes a pitch period and an amplitude gain value.
[0115] In one embodiment, S308 includes the following steps: performing speech enhancement processing on the filter speech excitation signal according to the pitch period, the amplitude gain value, the linear filter parameter and the cepstral feature parameter, to obtain an enhanced speech excitation signal.
[0116] Specifically, the speech enhancement processing on the audio signal can be realized by a pre-trained signal enhancement model, and the signal enhancement model is a neural network (NN) model, which can specifically adopt an LSTM and CNN level structure.
[0117] In the above embodiment, the terminal performs speech enhancement processing on the filter speech excitation signal according to the pitch period, the amplitude gain value, the linear filter parameter and the cepstral feature parameter, to obtain an enhanced speech excitation signal, and then the enhancement of the audio signal can be realized based on the enhanced speech excitation signal, and the quality of the audio signal is improved.
[0118] In one embodiment, the terminal inputs the obtained feature parameter, long-term filter parameter, linear filter parameter and filter speech excitation signal into a pre-trained signal enhancement model, so that the signal enhancement model performs speech enhancement processing on the filter speech excitation signal based on the feature parameter, to obtain an enhanced speech excitation signal.
[0119] In the above embodiment, the terminal realizes the enhanced speech excitation signal through the pre-trained signal enhancement model, and then the enhancement of the audio signal can be realized based on the enhanced speech excitation signal, and the quality of the audio signal and the efficiency of the enhancement processing on the audio signal are improved.
[0120] It should be noted that in the embodiments of the present application, the speech enhancement processing of the filter speech excitation signal by the pre-trained signal enhancement model is performed on the whole frame of the filter speech excitation signal, and the whole frame of the filter speech excitation signal does not need to be processed in sub-frames.
[0121] In S310, speech synthesis is performed based on the enhanced speech excitation signal and the linear filter parameter to obtain a speech enhancement signal.
[0122] The speech synthesis can be linear synthesis filtering based on the linear filter parameter.
[0123] In one embodiment, after obtaining the enhanced speech excitation signal, the terminal performs parameter configuration on the linear prediction filter based on the linear filter parameter, and performs linear synthesis filtering on the enhanced speech excitation signal by the parameter-configured linear prediction filter to obtain the speech enhancement signal.
[0124] The linear filter parameter includes a linear filter coefficient and an energy gain value, the linear filter coefficient can be denoted as LPCAR, and the energy gain value can be denoted as LPC gain. The linear synthesis filtering is the inverse process of the linear analysis filtering performed when the sending terminal encodes the audio signal, and thus the linear prediction filter performing the linear synthesis filtering is also called a linear inverse filter. The time-domain representation of the linear prediction filter is as follows:
[0125]
[0126] S enh (n) is the speech enhancement signal, D enh (n) is the enhanced speech excitation signal obtained by performing speech enhancement processing on the filter speech excitation signal D(n), S adj (n-i) is the energy-adjusted state of the previous frame of the obtained audio signal S(n), p is the number of sampling points contained in each frame of the audio signal, A i is the linear filter coefficient obtained by decoding the speech packet.
[0127] The energy-adjusted state of the previous frame of the audio signal S(n), S adj (n-i) can be obtained by the following formula:
[0128] S adj (n-i)=gain adj ·S(n-i) (9)
[0129] In the above formula, S adj (n-i) is the energy-adjusted state of the previous frame of the audio signal S(n-i), gain adjThe energy adjustment parameter is used for adjusting the energy of the previous frame of audio signal S(n-i).
[0130] In this embodiment, the terminal can obtain the speech enhancement signal by performing linear synthesis filtering on the enhanced speech excitation signal, that is, the enhancement processing of the audio signal is realized, and the quality of the audio signal is improved.
[0131] It should be noted that the speech synthesis process in the embodiment of the present application is speech synthesis on the whole frame of the enhanced speech excitation signal, and does not need to perform sub-frame processing on the whole frame of the enhanced speech excitation signal.
[0132] The above-mentioned audio signal enhancement method, when the terminal receives the speech packet, sequentially decodes and filters the speech packet to obtain the audio signal, and when the audio signal is a forward error correction frame signal, extracts the feature parameter from the audio signal, converts the audio signal into a filter speech excitation signal based on the linear filter coefficient obtained by decoding the speech packet, performs speech enhancement processing on the filter speech excitation signal according to the feature parameter and the long-time filter parameter obtained by decoding the speech packet to obtain an enhanced speech excitation signal, and performs speech synthesis based on the enhanced speech excitation signal and the linear filter parameter to obtain a speech enhancement signal. Thus, the enhancement processing of the audio signal is completed in a shorter time, and a better signal enhancement effect is achieved, improving the timeliness of the audio signal enhancement.
[0133] In one embodiment, as shown in FIG. 3, Figure 6 S302 specifically includes the following steps:
[0134] S602, parameter configuration of the long-time prediction filter based on the long-time filter parameter, long-time synthesis filtering of the residual signal by the parameter-configured long-time prediction filter to obtain a long-time filter excitation signal.
[0135] The long-time filter parameter includes a pitch period and a corresponding amplitude gain value, the pitch period can be denoted as LTP pitch, the LTP pitch can also be referred to as the pitch period, and the corresponding amplitude gain value can be denoted as LTP gain. The long-time prediction filter performs long-time synthesis filtering on the residual signal after parameter configuration, and the long-time synthesis filtering is the inverse process of the long-time analysis filtering performed by the sending end when encoding the audio signal. Therefore, the long-time prediction filter performing the long-time synthesis filtering is also referred to as a long-time inverse filter, that is, the long-time inverse filter is used to process the residual signal. The frequency domain expression of the long-time inverse filter corresponding to formula (1) is as follows:
[0136]
[0137] Wherein, p -1(z) is the amplitude frequency response of the long-term inverse filter, z is the rotation factor of the frequency domain transform, γ is the amplitude gain value LTP gain, T is the pitch LTP pitch, Figure 7 The amplitude frequency response diagram of the long-term inverse prediction filter corresponding to γ = 1, T = 80 in an embodiment is shown.
[0138] The time domain representation of the long-term inverse filter corresponding to formula (10) is as follows:
[0139] E(n) = γE(n-T) + δ(n) (11)
[0140] In the above formula, E(n) is the long-term filtered excitation signal corresponding to the speech packet, δ(n) is the residual signal corresponding to the speech packet, γ is the amplitude gain value LTP gain, T is the pitch LTP pitch, and E(n-T) is the long-term filtered excitation signal corresponding to the audio signal of the previous pitch period of the speech packet. It can be understood that in the embodiment, the long-term filtered excitation signal E(n) obtained by the receiving end through long-term synthesis filtering of the residual signal by the long-term inverse filter is the same as the linear filtered excitation signal e(n) obtained by the transmitting end through linear analysis filtering of the audio signal by the linear filter when encoding.
[0141] S604, parameter configuration is carried out on the linear prediction filter based on the linear filtering parameter, and the linear synthesis filtering of the long-term filtered excitation signal is carried out through the linear prediction filter after the parameter configuration, so that the audio signal is obtained.
[0142] Wherein, the linear filtering parameter includes the linear filtering coefficient and the energy gain value, the linear filtering coefficient can be denoted as LPCAR, the energy gain value can be denoted as LPC gain, the linear synthesis filtering is the inverse process of the linear analysis filtering carried out by the transmitting end when encoding the audio signal, therefore, the linear prediction filter carrying out the linear synthesis filtering is also called linear inverse filter, and the time domain representation of the linear prediction filter is as follows:
[0143]
[0144] In the above formula, S(n) is the audio signal corresponding to the speech packet, E(n) is the long-term filtered excitation signal corresponding to the speech packet, S adj (n-i) is the energy adjusted state of the previous frame audio signal S(n-i) of the audio signal S(n), p is the number of sampling points contained in each frame of audio signal, A i is the linear filtering coefficient obtained by decoding the speech packet.
[0145] The energy adjusted state of the previous frame audio signal S(n-i) of the audio signal S(n), S adj (n-i) can be obtained by the following formula:
[0146]
[0147] wherein, gain adj is the energy adjustment parameter of the previous frame of audio signal S(n-i), gain(n) is the energy gain value obtained by decoding the speech packet, and gain(n-i) is the energy gain value corresponding to the previous frame of audio signal.
[0148] In the above embodiment, the terminal performs long-time synthesis filtering on the residual signal based on the long-time filtering parameter to obtain a long-time filtering excitation signal, and performs linear synthesis filtering on the long-time filtering excitation signal based on the linear filtering parameter obtained by decoding to obtain the audio signal. Therefore, when the audio signal is not a forward error correction frame signal, the audio signal can be directly output, and when the audio signal is a forward error correction frame signal, the audio signal is output after being enhanced, thereby improving the timeliness of the audio signal output.
[0149] In one embodiment, S604 specifically includes the following steps: dividing the long-time filtering excitation signal into at least two subframes to obtain a sub long-time filtering excitation signal; grouping the linear filtering parameter obtained by decoding to obtain at least two linear filtering parameter sets; performing parameter configuration on the at least two linear prediction filters based on the linear filtering parameter sets; and inputting the obtained sub long-time filtering excitation signal into the linear prediction filters after parameter configuration, so that the linear prediction filters perform linear synthesis filtering on the sub long-time filtering excitation signal based on the linear filtering parameter sets to obtain sub audio signals corresponding to the subframes; and combining the sub audio signals according to the time sequence of the subframes to obtain the audio signal.
[0150] The linear filtering parameter set has two types of linear filtering coefficient set and energy gain value set.
[0151] Specifically, for the sub long-time filtering excitation signal corresponding to each subframe, when linear synthesis filtering is performed by using the linear inverse filter corresponding to formula (12), S(n) in formula (12) is the sub audio signal corresponding to any subframe, E(n) is the long-time filtering excitation signal corresponding to the subframe, S adj (n-i) is the S(n-i) energy adjusted state of the sub audio signal of the previous subframe of the sub audio signal S(n) obtained, p is the number of sampling points contained in each subframe of audio signal, A i is the linear filtering coefficient set corresponding to the subframe; gain adj in formula (13) is the energy adjustment parameter of the sub audio signal of the previous subframe of the sub audio signal, gain(n) is the energy gain value of the sub audio signal, and gain(n-i) is the energy gain value of the sub audio signal of the previous subframe of the sub audio signal.
[0152] In the above embodiment, the terminal divides the long-time filtering excitation signal into at least two subframes to obtain sub long-time filtering excitation signals; groups the linear filtering parameters obtained after decoding to obtain at least two sets of linear filtering parameters; configures parameters of at least two linear prediction filters based on the sets of linear filtering parameters respectively; inputs the obtained sub long-time filtering excitation signals into the linear prediction filters configured with parameters respectively, so that the linear prediction filters perform linear synthesis filtering on the sub long-time filtering excitation signals based on the sets of linear filtering parameters to obtain sub audio signals corresponding to the subframes; and combines the sub audio signals according to the time sequence of the subframes to obtain an audio signal, thereby ensuring that the obtained audio signal can restore the audio signal sent by the sending terminal well and improving the quality of the restored audio signal.
[0153] In one embodiment, the linear filtering parameters include linear filtering coefficients and energy gain values; and S604 further includes the following steps: obtaining, for the sub long-time filtering excitation signal corresponding to the first subframe of the long-time filtering excitation signal, an energy gain value of a historical sub long-time filtering excitation signal of a subframe adjacent to the sub long-time filtering excitation signal in a historical long-time filtering excitation signal; determining an energy adjustment parameter corresponding to the sub long-time filtering excitation signal based on the energy gain value corresponding to the historical sub long-time filtering excitation signal and the energy gain value of the sub long-time filtering excitation signal corresponding to the first subframe; and performing energy adjustment on the historical sub long-time filtering excitation signal by using the energy adjustment parameter to obtain an energy-adjusted historical sub long-time filtering excitation signal.
[0154] In one embodiment, the linear filtering parameters include linear filtering coefficients and energy gain values; and S604 further includes the following steps: obtaining, for the sub long-time filtering excitation signal corresponding to the first subframe of the long-time filtering excitation signal, an energy gain value of a historical sub long-time filtering excitation signal of a subframe adjacent to the sub long-time filtering excitation signal in a historical long-time filtering excitation signal; determining an energy adjustment parameter corresponding to the sub long-time filtering excitation signal based on the energy gain value corresponding to the historical sub long-time filtering excitation signal and the energy gain value of the sub long-time filtering excitation signal corresponding to the first subframe; and performing energy adjustment on the historical sub long-time filtering excitation signal by using the energy adjustment parameter to obtain an energy-adjusted historical sub long-time filtering excitation signal.
[0155] For example, the long-time filtering excitation signal of the current frame is divided into two subframes to obtain a sub long-time filtering excitation signal corresponding to a first subframe and a sub long-time filtering excitation signal corresponding to a second subframe, and the sub long-time filtering excitation signal corresponding to the second subframe of the previous frame long-time filtering excitation signal is adjacent to the sub long-time filtering excitation signal corresponding to the first subframe of the current frame.
[0156] In one embodiment, after obtaining the energy-adjusted historical sub long-time filtering excitation signal, the terminal inputs the obtained sub long-time filtering excitation signal and the energy-adjusted historical sub long-time filtering excitation signal into the linear prediction filter configured with parameters, so that the linear prediction filter performs linear synthesis filtering on the sub long-time filtering excitation signal corresponding to the first subframe based on the linear filtering coefficients and the energy-adjusted historical sub long-time filtering excitation signal to obtain a sub audio signal corresponding to the first subframe.
[0157] For example, one speech packet corresponds to 20ms audio signal, i.e. the obtained long-time filter excitation signal is 20ms, the AR coefficients obtained by decoding the speech packet are {A1, A2, …, A p-1 ,A p ,A p+1 ,…A 2p-1 ,A 2p}, the energy gain values obtained by decoding the speech packet are {gain1(n), gain2(n)}, the long-time filter excitation signal can be divided into two sub-frames to obtain the first sub-filter excitation signal E1(n) corresponding to the first 10ms and the second sub-filter excitation signal E2(n) corresponding to the last 10ms, the AR coefficients are grouped to obtain AR coefficient set 1 {A1, A2, …, A p-1 ,A p} and AR coefficient set 2 {A p+1 ,…A 2p-1 ,A 2p}, the energy gain values are grouped to obtain energy gain value set 1 {gain1(n)} and energy gain value set 2 {gain2(n)}, the sub-filter excitation signal of the previous sub-frame of the first sub-filter excitation signal E1(n) is E2(n-i), the energy gain value set of the previous sub-frame of the first sub-filter excitation signal E1(n) is {gain2(n-i)}, the sub-filter excitation signal of the previous sub-frame of the second sub-filter excitation signal E2(n) is E1(n), the energy gain value set of the previous sub-frame of the second sub-filter excitation signal E2(n) is {gain1(n)}, then the sub-audio signal corresponding to the first sub-filter excitation signal E1(n) can be obtained by substituting the corresponding parameters into formula (12) and formula (13), and the sub-audio signal corresponding to the second sub-filter excitation signal E2(n) can be obtained by substituting the corresponding parameters into formula (12) and formula (13).
[0158] In the above embodiment, the terminal obtains an energy gain value of a historical sub-long-term filtered excitation signal corresponding to a subframe adjacent to a first subframe corresponding to the sub-long-term filtered excitation signal in the long-term filtered excitation signal, determines an energy adjustment parameter corresponding to the sub-long-term filtered excitation signal based on the energy gain value corresponding to the historical sub-long-term filtered excitation signal and the energy gain value of the sub-long-term filtered excitation signal corresponding to the first subframe, and performs energy adjustment on the historical sub-long-term filtered excitation signal through the energy adjustment parameter. The obtained sub-long-term filtered excitation signal and the historical sub-long-term filtered excitation signal obtained after energy adjustment are input into the linear prediction filter after parameter configuration, so that the linear prediction filter performs linear synthesis filtering on the sub-long-term filtered excitation signal corresponding to the first subframe based on the linear filter coefficient and the historical sub-long-term filtered excitation signal obtained after energy adjustment, to obtain a sub-audio signal corresponding to the first subframe. Thus, the obtained each subframe audio signal can restore each subframe audio signal sent by the sending end, and the quality of the restored audio signal is improved.
[0159] In one embodiment, the feature parameter includes a cepstrum feature parameter, and S308 includes the following steps: vectorizing the cepstrum feature parameter, the long-term filtering parameter, and the linear filtering parameter, and splicing the results obtained after vectorization to obtain a feature vector; inputting the feature vector and the filter voice excitation signal into a pre-trained signal enhancement model; extracting features of the feature vector through the signal enhancement model to obtain a target feature vector; and performing enhancement processing on the filter voice excitation signal based on the target feature vector to obtain an enhanced voice excitation signal.
[0160] The signal enhancement model is a multi-level network structure, specifically including a first feature splicing layer, a second feature splicing layer, a first neural network layer, and a second neural network layer. The target feature vector is an enhanced feature vector.
[0161] Specifically, the terminal vectorizes the cepstrum feature parameter, the long-term filtering parameter, and the linear filtering parameter through the first feature splicing layer of the signal enhancement model, splices the results obtained after vectorization to obtain a feature vector, then inputs the obtained feature vector into the first neural network layer of the signal enhancement model, extracts features of the feature vector through the first neural network layer to obtain a primary feature vector, inputs the primary feature vector and envelope information obtained by performing Fourier transform on the linear filter coefficient in the linear filtering parameter into the second feature splicing layer of the signal enhancement model, splices the primary feature vector, inputs the spliced primary feature vector into the second neural network layer of the signal enhancement model, extracts features of the spliced primary feature vector through the second neural network layer to obtain a target feature vector, and then performs enhancement processing on the filter voice excitation signal based on the target feature vector to obtain an enhanced voice excitation signal.
[0162] In the above embodiment, the terminal obtains a feature vector by vectorizing the cepstrum feature parameter, the long-time filtering parameter and the linear filtering parameter, and splicing the results obtained by the vectorization processing; inputs the feature vector and the filter voice excitation signal into the pre-trained signal enhancement model; extracts features of the feature vector by the signal enhancement model to obtain a target feature vector; and performs enhancement processing on the filter voice excitation signal based on the target feature vector to obtain an enhanced voice excitation signal. Thus, the signal enhancement model can be used to perform enhancement processing on the audio signal, and the quality of the audio signal and the efficiency of the enhancement processing on the audio signal are improved.
[0163] In one embodiment, the terminal performs enhancement processing on the filter voice excitation signal based on the target feature vector to obtain an enhanced voice excitation signal, including: performing Fourier transform on the filter voice excitation signal to obtain a frequency domain voice excitation signal; enhancing the amplitude feature of the frequency domain voice excitation signal based on the target feature vector; and performing inverse Fourier transform on the frequency domain voice excitation signal with the enhanced amplitude feature to obtain the enhanced voice excitation signal.
[0164] Specifically, after performing Fourier transform on the filter voice excitation signal to obtain a frequency domain voice excitation signal, and after enhancing the amplitude feature of the frequency domain voice excitation signal based on the target feature vector, the terminal performs inverse Fourier transform on the frequency domain voice excitation signal with the enhanced amplitude feature in combination with the phase feature of the frequency domain voice excitation signal that is not enhanced to obtain the enhanced voice excitation signal.
[0165] As Figure 8As shown, two feature concatenation layers are concat1 and concat2 respectively, and two neural network layers are NN part1 and NN part2 respectively, the cepstrum feature parameter Cepstrum with a dimension of 40, the pitch cycle LTP pitch with a dimension of 1 and the amplitude gain value LTP Gain with a dimension of 1 are concatenated together through concat1 to form a feature vector with a dimension of 42, and the feature vector with a dimension of 42 is input into NN part1, NN part1 is composed of a two-layer convolutional neural network and two fully connected layers, the dimension of the first layer of convolution kernel is (1, 128, 3, 1), the dimension of the second layer of convolution kernel is (128, 128, 3, 1), the number of nodes of the fully connected layers is 128 and 8, and the activation function at the end of each layer is a Tanh function, high-level features are extracted from the feature vector through NN part1 to obtain a primary feature vector with a dimension of 1024, then the primary feature vector with a dimension of 1024 is concatenated with envelope information Envelope with a dimension of 161 obtained by Fourier transform on the linear filter coefficient LTP AR in the linear filter parameter to obtain a concatenated primary feature vector with a dimension of 1185, and the concatenated primary feature vector with a dimension of 1185 is input into NN part2, NN part2 is a two-layer fully connected network with node numbers of 256 and 161 respectively, the activation function at the end of each layer is a Tanh function, and the target feature vector is obtained through NN part2, then based on the target feature vector, the amplitude feature Excitation of the frequency domain speech excitation signal obtained by Fourier transform on the filter speech excitation signal is enhanced, and the filter speech excitation signal with the enhanced amplitude feature Excitation is inverse Fourier transformed to obtain an enhanced speech excitation signal D enh (n).
[0166] In the above embodiment, the terminal performs Fourier transform on the filter speech excitation signal to obtain a frequency domain speech excitation signal, enhances the amplitude feature of the frequency domain speech excitation signal based on the target feature vector, and inverse Fourier transforms the frequency domain speech excitation signal with the enhanced amplitude feature to obtain an enhanced speech excitation signal, so that the enhancement processing of the audio signal can be realized while ensuring that the phase information of the audio signal is unchanged, and the quality of the audio signal is improved.
[0167] In one embodiment, the linear filtering parameter includes a linear filtering coefficient and an energy gain value; the terminal performs parameter configuration on the linear prediction filter based on the linear filtering parameter, and the step of performing linear synthesis filtering on the enhanced speech excitation signal by the parameter-configured linear prediction filter includes: performing parameter configuration on the linear prediction filter based on the linear filtering coefficient; obtaining an energy gain value corresponding to a historical speech packet decoded before the current speech packet; determining an energy adjustment parameter based on the energy gain value corresponding to the historical speech packet and the energy gain value corresponding to the current speech packet; performing energy adjustment on a historical long-term filtering excitation signal corresponding to the historical speech packet by the energy adjustment parameter to obtain an adjusted historical long-term filtering excitation signal; and inputting the adjusted historical long-term filtering excitation signal and the enhanced speech excitation signal into the parameter-configured linear prediction filter, so that the linear prediction filter performs linear synthesis filtering on the enhanced speech excitation signal based on the adjusted historical long-term filtering excitation signal.
[0168] wherein the historical audio signal corresponding to the historical speech packet is a previous frame audio signal of a current frame audio signal corresponding to the current speech packet. The energy gain value corresponding to the historical speech packet can be an energy gain value corresponding to an integral frame audio signal of the historical speech packet, or an energy gain value corresponding to a partial sub-frame audio signal of the historical speech packet.
[0169] Specifically, when the audio signal is a non-forward error correction frame signal, i.e., the previous frame audio signal of the current frame audio signal is obtained by normal decoding of the historical speech packet by the terminal, the energy gain value of the historical speech packet obtained by the terminal when decoding the historical speech packet can be obtained, and the energy adjustment parameter is determined based on the energy gain value of the historical speech packet; when the audio signal is a forward error correction frame, i.e., the previous frame audio signal of the current frame audio signal is not obtained by normal decoding of the historical speech packet by the terminal, a compensation energy gain value corresponding to the previous frame audio signal is determined based on a preset energy gain compensation mechanism, and the compensation energy gain value is determined as the energy gain value of the historical speech packet, so as to determine the energy adjustment parameter based on the energy gain value of the historical speech packet.
[0170] In one embodiment, when the audio signal is a non-forward error correction frame signal, the energy adjustment parameter gain adj The gain can be calculated by the following formula:
[0171]
[0172] wherein the gain adjgain(n-i) is an energy adjustment parameter of the previous frame audio signal S(n-i), gain(n-i) is an energy gain value of the previous frame audio signal S(n-i), and gain(n) is an energy gain value of the current frame audio signal. Formula (14) is to calculate the energy adjustment parameter based on the energy gain value corresponding to the historical speech of the whole frame audio signal.
[0173] In one embodiment, when the audio signal is a non-forward error correction frame signal, the energy adjustment parameter gain adj may be obtained by the following formula:
[0174]
[0175] wherein, gain adj is an energy adjustment parameter of the previous frame audio signal S(n-i), gain m (n-i) is an energy gain value of the mth subframe of the previous frame audio signal S(n-i), gain m (n) is an energy gain value of the mth subframe of the current frame audio signal, m is the number of subframes corresponding to each audio signal, and {gain1(n)+…+gain(n)} / m is an energy gain value of the current frame audio signal. Formula (15) is to calculate the energy adjustment parameter based on the energy gain value corresponding to the historical speech of the partial subframe audio signal.
[0176] In the above embodiment, the terminal performs parameter configuration on the linear prediction filter based on the linear filter coefficient; obtains an energy gain value corresponding to the historical speech packet decoded before the speech packet is decoded; determines an energy adjustment parameter based on the energy gain value corresponding to the historical speech packet and the energy gain value corresponding to the speech packet; performs energy adjustment on the historical long-term filter excitation signal corresponding to the historical speech packet based on the energy adjustment parameter to obtain an adjusted historical long-term filter excitation signal; inputs the adjusted historical long-term filter excitation signal and the enhanced speech excitation signal into the linear prediction filter configured with the parameters, so that the linear prediction filter performs linear synthesis filtering on the enhanced speech excitation signal based on the adjusted historical long-term filter excitation signal, thereby smoothing the audio signals between different frames and improving the quality of the speech composed of the audio signals of different frames.
[0177] In one embodiment, as shown in Figure 9 , an audio signal enhancement method is provided, which is applied to a computer device (terminal or server) in Figure 2 for example, and includes the following steps:
[0178] S902, decoding the speech packet to obtain a residual signal, a long-term filter parameter, and a linear filter parameter.
[0179] S904, parameter configuration is performed on the long-time prediction filter based on the long-time filtering parameter, and long-time synthesis filtering is performed on the residual signal by the parameter-configured long-time prediction filter to obtain a long-time filtering excitation signal.
[0180] S906, the long-time filtering excitation signal is divided into at least two subframes to obtain a sub long-time filtering excitation signal.
[0181] S908, the de-linear filtering parameters are grouped to obtain at least two sets of linear filtering parameters.
[0182] S910, parameter configuration is performed on the at least two linear prediction filters based on the sets of linear filtering parameters, respectively.
[0183] S912, the obtained sub long-time filtering excitation signal is input into the parameter-configured linear prediction filter, respectively, so that the linear prediction filter performs linear synthesis filtering on the sub long-time filtering excitation signal based on the sets of linear filtering parameters to obtain a sub audio signal corresponding to each subframe.
[0184] S914, the sub audio signals are combined according to the time sequence of the subframes to obtain an audio signal.
[0185] S916, it is determined whether a historical speech packet decoded before the current speech packet has data anomaly.
[0186] S918, if the historical speech packet has data anomaly, the audio signal obtained through decoding and filtering is determined as a forward error correction frame signal.
[0187] S920, when the audio signal is the forward error correction frame signal, Fourier transform is performed on the audio signal to obtain a Fourier-transformed audio signal, logarithmic processing is performed on the Fourier-transformed audio signal to obtain a logarithmic result, and inverse Fourier transform is performed on the logarithmic result to obtain a cepstrum feature parameter.
[0188] S922, parameter configuration is performed on the linear prediction filter based on the linear filtering parameter, and linear decomposition filtering is performed on the audio signal by the parameter-configured linear prediction filter to obtain a filter speech excitation signal.
[0189] S924, the feature parameter, the long-time filtering parameter, the linear filtering parameter, the linear filtering parameter, and the filter speech excitation signal are input into a pre-trained signal enhancement model, so that the signal enhancement model performs speech enhancement processing on the filter speech excitation signal based on the feature parameter to obtain an enhanced speech excitation signal.
[0190] S926, parameter configuration is performed on the linear prediction filter based on the linear filtering parameter, and linear synthesis filtering is performed on the enhanced speech excitation signal by the parameter-configured linear prediction filter to obtain a speech enhancement signal.
[0191] The application also provides an application scenario of the audio signal enhancement method.
[0192] Specifically, the audio signal enhancement method is applied in the application scenario as follows:
[0193] Taking a wideband signal with Fs of 16000Hz as an example, it can be understood that the application is also applicable to scenarios with other sampling rates, such as Fs of 8000Hz, 32000Hz or 48000Hz. The frame length of the audio signal is set to 20ms; for Fs = 16000Hz, it is equivalent to containing 320 sample points per frame. Referring to Figure 10 , after the terminal receives a speech packet corresponding to a frame of audio signal, the speech packet is entropy decoded to obtain δ(n), LTP pitch, LTP gain, LPC AR and LPC gain, LTP synthesis filtering is performed on δ(n) based on LTP pitch and LTP gain to obtain E(n), LPC synthesis filtering is performed on each subframe of E(n) based on LPC AR and LPC gain, and the LPC synthesis filtering results are combined to obtain a frame of S(n), then cepstrum analysis is performed on S(n) to obtain C(n), and LPC decomposition filtering is performed on the whole frame of S(n) based on LPC AR and LPC gain to obtain the whole frame of D(n), the envelope information after Fourier transform of LTP pitch, LTP gain, LPC AR, C(n) and D(n) are input into a pre-trained signal enhancement model NN postfilter, the whole frame of D enh (n) is enhanced by the NN postfilter to obtain the whole frame of D enh (n), and LPC synthesis filtering is performed on the whole frame of D enh (n) based on LPC AR and LPC gain to obtain S
[0194] It should be understood that although Figure 3 、 Figure 4 、 Figure 6 、 Figure 9 and Figure 10 the steps in the flowcharts are displayed in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified herein, the execution of these steps is not strictly limited in sequence, and these steps can be executed in other orders. Moreover, Figure 3 、 Figure 4 、 Figure 6 、 Figure 9 and Figure 10At least one of the steps in the above method can comprise a plurality of steps or stages which are not necessarily performed at the same time but can be performed at different times and in which the order of the steps or stages is not necessarily sequential but can be performed in rotation or alternation with other steps or steps or stages in other steps.
[0195] In one embodiment, as shown in FIG. 11, there is provided an audio signal enhancement device which can be a software module or a hardware module or a combination of both as part of a computer device, the device comprising: Figure 11 a speech packet processing module 1102, a feature parameter extraction module 1104, a signal conversion module 1106, a speech enhancement module 1108 and a speech synthesis module 1110, wherein:
[0196] The speech packet processing module 1102 is configured to sequentially decode and filter the received speech packet to obtain a residual signal, a long-term filter parameter and a linear filter parameter, and filter the residual signal to obtain an audio signal.
[0197] The feature parameter extraction module 1104 is configured to extract a feature parameter from the audio signal when the audio signal is a forward error correction frame signal.
[0198] The signal conversion module 1106 is configured to convert the audio signal into a filter speech excitation signal based on the linear filter parameter.
[0199] The speech enhancement module 1108 is configured to perform speech enhancement processing on the filter speech excitation signal according to the feature parameter, the long-term filter parameter and the linear filter parameter to obtain an enhanced speech excitation signal.
[0200] The speech synthesis module 1110 is configured to perform speech synthesis based on the enhanced speech excitation signal and the linear filter parameter to obtain a speech enhancement signal.
[0201] In the above embodiment, the computer device sequentially decodes the received speech packet to obtain a residual signal, a long-term filter parameter and a linear filter parameter, filters the residual signal to obtain an audio signal, extracts a feature parameter from the audio signal when the audio signal is a forward error correction frame signal, converts the audio signal into a filter speech excitation signal based on the linear filter parameter obtained from the decoded speech packet, and performs speech enhancement processing on the filter speech excitation signal according to the feature parameter and the long-term filter parameter obtained from the decoded speech packet to obtain an enhanced speech excitation signal, and performs speech synthesis based on the enhanced speech excitation signal and the linear filter parameter to obtain a speech enhancement signal, thereby completing the enhancement processing of the audio signal in a shorter time and achieving a better signal enhancement effect, improving the timeliness of the audio signal enhancement.
[0202] In an embodiment, the voice packet processing module 1102 is further configured to: configure a long-time prediction filter based on the long-time filter parameter, perform long-time synthesis filtering on the residual signal by the long-time prediction filter to obtain a long-time filter excitation signal; and configure a linear prediction filter based on the linear filter parameter, perform linear synthesis filtering on the long-time filter excitation signal by the linear prediction filter to obtain the audio signal.
[0203] In the above embodiment, the terminal performs long-time synthesis filtering on the residual signal based on the long-time filter parameter to obtain a long-time filter excitation signal, and performs linear synthesis filtering on the long-time filter excitation signal based on the decoded linear filter parameter to obtain the audio signal, so that the audio signal can be directly output when the audio signal is not a forward error correction frame signal, and the audio signal can be output after being enhanced when the audio signal is a forward error correction frame signal, thereby improving the timeliness of the audio signal output.
[0204] In an embodiment, the voice packet processing module 1102 is further configured to: divide the long-time filter excitation signal into at least two subframes to obtain a sub-long-time filter excitation signal; group the linear filter parameter to obtain at least two linear filter parameter sets; configure at least two linear prediction filters based on the linear filter parameter sets, respectively; input the obtained sub-long-time filter excitation signal into the linear prediction filters configured based on the linear filter parameter sets, respectively, so that the linear prediction filters perform linear synthesis filtering on the sub-long-time filter excitation signal based on the linear filter parameter sets to obtain sub-audio signals corresponding to the subframes, respectively; and combine the sub-audio signals according to the time sequence of the subframes to obtain the audio signal.
[0205] In the above embodiment, the terminal divides the long-time filter excitation signal into at least two subframes to obtain a sub-long-time filter excitation signal, groups the linear filter parameter to obtain at least two linear filter parameter sets, configures at least two linear prediction filters based on the linear filter parameter sets, respectively, inputs the obtained sub-long-time filter excitation signal into the linear prediction filters configured based on the linear filter parameter sets, respectively, so that the linear prediction filters perform linear synthesis filtering on the sub-long-time filter excitation signal based on the linear filter parameter sets to obtain sub-audio signals corresponding to the subframes, respectively, and combines the sub-audio signals according to the time sequence of the subframes to obtain the audio signal, thereby ensuring that the obtained audio signal can restore the audio signal sent by the sending terminal well, and improving the quality of the restored audio signal.
[0206] In one embodiment, the linear filtering parameters include linear filtering coefficients and energy gain values. The voice packet processing module 1102 is further configured to: for the sub-long-term filtering excitation signal corresponding to the first sub-frame in the long-term filtering excitation signal, obtain the energy gain values corresponding to the historical sub-long-term filtering excitation signals of the sub-frames adjacent to the sub-long-term filtering excitation signal corresponding to the first sub-frame in the historical long-term filtering excitation signal; determine the energy adjustment parameters corresponding to the sub-long-term filtering excitation signal based on the energy gain values corresponding to the historical sub-long-term filtering excitation signals and the energy gain values of the sub-long-term filtering excitation signals corresponding to the first sub-frame; adjust the energy of the historical sub-long-term filtering excitation signals using the energy adjustment parameters; and input the obtained sub-long-term filtering excitation signals and the historical sub-long-term filtering excitation signals obtained after energy adjustment to the parameter-configured linear prediction filter, so that the linear prediction filter performs linear synthesis filtering on the sub-long-term filtering excitation signals corresponding to the first sub-frame based on the linear filtering coefficients and the historical sub-long-term filtering excitation signals obtained after energy adjustment, to obtain the sub-audio signal corresponding to the first sub-frame.
[0207] In the above embodiments, the terminal obtains the energy gain value of the historical sub-long-term filtering excitation signal of the sub-long-term filtering excitation signal of the sub-frame adjacent to the sub-long-term filtering excitation signal of the first sub-frame in the historical long-term filtering excitation signal; based on the energy gain value of the historical sub-long-term filtering excitation signal and the energy gain value of the sub-long-term filtering excitation signal of the first sub-frame, it determines the energy adjustment parameter corresponding to the sub-long-term filtering excitation signal; the energy adjustment parameter is used to adjust the energy of the historical sub-long-term filtering excitation signal, and the obtained sub-long-term filtering excitation signal and the energy-adjusted historical sub-long-term filtering excitation signal are input to the parameter-configured linear prediction filter, so that the linear prediction filter performs linear synthesis filtering on the sub-long-term filtering excitation signal of the first sub-frame based on the linear filtering coefficients and the energy-adjusted historical sub-long-term filtering excitation signal, to obtain the sub-audio signal corresponding to the first sub-frame. This ensures that the obtained sub-frame audio signal can better reproduce the sub-frame audio signal sent by the transmitter, thus improving the quality of the reproduced audio signal.
[0208] In one embodiment, such as Figure 12 As shown, the device further includes: a data anomaly determination module 1112 and a forward error correction frame signal determination module 1114, wherein: the data anomaly determination module 1112 is used to determine whether a data anomaly has occurred in a historical voice packet decoded before the decoded voice packet; the forward error correction frame signal determination module 1114 is used to determine that the audio signal obtained after decoding and filtering is a forward error correction frame signal if a data anomaly has occurred in a historical voice packet.
[0209] In the above embodiment, the terminal determines whether the audio signal after decoding and filtering is a forward error correction frame signal by determining whether the historical speech packets decoded before decoding the current speech packet have data exceptions, and then can perform audio signal enhancement processing on the audio signal when the audio signal is a forward error correction frame signal, thereby further improving the quality of the audio signal.
[0210] In one embodiment, the feature parameter includes a cepstrum feature parameter; the feature parameter extraction module 1104 is further configured to perform Fourier transform on the audio signal to obtain a Fourier transformed audio signal; perform logarithmic processing on the Fourier transformed audio signal to obtain a logarithmic result; and perform inverse Fourier transform on the logarithmic result to obtain the cepstrum feature parameter.
[0211] In the above embodiment, the terminal extracts the cepstrum feature parameter from the audio signal, so that the audio signal can be enhanced based on the extracted cepstrum feature parameter, thereby improving the quality of the audio signal.
[0212] In one embodiment, the long-time filter parameter includes a pitch period and an amplitude gain value; and the speech enhancement module 1108 is further configured to perform speech enhancement processing on the filter speech excitation signal based on the pitch period, the amplitude gain value, the linear filter parameter, and the cepstrum feature parameter to obtain an enhanced speech excitation signal.
[0213] In the above embodiment, the terminal performs speech enhancement processing on the filter speech excitation signal based on the pitch period, the amplitude gain value, the linear filter parameter, and the cepstrum feature parameter to obtain an enhanced speech excitation signal, and then can enhance the audio signal based on the enhanced speech excitation signal, thereby improving the quality of the audio signal.
[0214] In one embodiment, the signal conversion module 1106 is further configured to perform parameter configuration on the linear prediction filter based on the linear filter parameter, and perform linear decomposition filtering on the audio signal through the parameter configured linear prediction filter to obtain the filter speech excitation signal.
[0215] In the above embodiment, the terminal converts the audio signal into the filter speech excitation signal based on the linear filter parameter, so that the audio signal can be enhanced by enhancing the filter speech excitation signal, thereby improving the quality of the audio signal.
[0216] In one embodiment, the speech enhancement module 1108 is further configured to input the feature parameter, the long-time filter parameter, the linear filter parameter, and the filter speech excitation signal into a pre-trained signal enhancement model, so that the signal enhancement model performs speech enhancement processing on the filter speech excitation signal based on the feature parameter to obtain an enhanced speech excitation signal.
[0217] In the above embodiment, the terminal implements enhancement on the speech excitation signal after enhancement through the pre-trained signal enhancement model, and then can implement enhancement on the audio signal based on the speech excitation signal after enhancement, thereby improving the quality of the audio signal and the efficiency of the enhancement processing on the audio signal.
[0218] In one embodiment, the feature parameter includes a cepstrum feature parameter; the speech enhancement module 1108 is further configured to: perform vectorization processing on the cepstrum feature parameter, the long-time filtering parameter and the linear filtering parameter, and splice the results obtained through the vectorization processing to obtain a feature vector; input the feature vector and the filter speech excitation signal into the pre-trained signal enhancement model; perform feature extraction on the feature vector through the signal enhancement model to obtain a target feature vector; and perform enhancement processing on the filter speech excitation signal based on the target feature vector to obtain a speech excitation signal after enhancement.
[0219] In the above embodiment, the terminal performs vectorization processing on the cepstrum feature parameter, the long-time filtering parameter and the linear filtering parameter, and splices the results obtained through the vectorization processing to obtain a feature vector; inputs the feature vector and the filter speech excitation signal into the pre-trained signal enhancement model; performs feature extraction on the feature vector through the signal enhancement model to obtain a target feature vector; and performs enhancement processing on the filter speech excitation signal based on the target feature vector to obtain a speech excitation signal after enhancement, thereby implementing enhancement on the audio signal through the signal enhancement model, improving the quality of the audio signal and the efficiency of the enhancement processing on the audio signal.
[0220] In one embodiment, the speech enhancement module 1108 is further configured to: perform Fourier transform on the filter speech excitation signal to obtain a frequency domain speech excitation signal; perform enhancement on the amplitude feature of the frequency domain speech excitation signal based on the target feature vector; and perform inverse Fourier transform on the frequency domain speech excitation signal with the enhanced amplitude feature to obtain a speech excitation signal after enhancement.
[0221] In the above embodiment, the terminal performs Fourier transform on the filter speech excitation signal to obtain a frequency domain speech excitation signal; performs enhancement on the amplitude feature of the frequency domain speech excitation signal based on the target feature vector; and performs inverse Fourier transform on the frequency domain speech excitation signal with the enhanced amplitude feature to obtain a speech excitation signal after enhancement, thereby implementing enhancement on the audio signal while ensuring that the phase information of the audio signal is unchanged, and improving the quality of the audio signal.
[0222] In one embodiment, the speech synthesis module 1110 is further configured to: perform parameter configuration on the linear prediction filter based on the linear filtering parameter, and perform linear synthesis filtering on the speech excitation signal after enhancement through the linear prediction filter after the parameter configuration to obtain a speech enhancement signal.
[0223] In the embodiment, the terminal can obtain the speech enhancement signal by performing linear synthesis filtering on the enhanced speech excitation signal, that is, the enhancement processing of the audio signal is realized, and the quality of the audio signal is improved.
[0224] In one embodiment, the linear filtering parameter includes a linear filtering coefficient and an energy gain value; the speech synthesis module 1110 is further configured to: perform parameter configuration on the linear prediction filter based on the linear filtering coefficient; obtain an energy gain value corresponding to a historical speech packet decoded before the speech packet; determine an energy adjustment parameter based on the energy gain value corresponding to the historical speech packet and the energy gain value corresponding to the speech packet; perform energy adjustment on a historical long-term filter excitation signal corresponding to the historical speech packet by using the energy adjustment parameter to obtain an adjusted historical long-term filter excitation signal; and input the adjusted historical long-term filter excitation signal and the enhanced speech excitation signal into the linear prediction filter configured with the parameters, so that the linear prediction filter performs linear synthesis filtering on the enhanced speech excitation signal based on the adjusted historical long-term filter excitation signal.
[0225] In the above embodiment, the terminal performs parameter configuration on the linear prediction filter based on the linear filtering coefficient; obtains an energy gain value corresponding to a historical speech packet decoded before the speech packet; determines an energy adjustment parameter based on the energy gain value corresponding to the historical speech packet and the energy gain value corresponding to the speech packet; performs energy adjustment on a historical long-term filter excitation signal corresponding to the historical speech packet by using the energy adjustment parameter to obtain an adjusted historical long-term filter excitation signal; and inputs the adjusted historical long-term filter excitation signal and the enhanced speech excitation signal into the linear prediction filter configured with the parameters, so that the linear prediction filter performs linear synthesis filtering on the enhanced speech excitation signal based on the adjusted historical long-term filter excitation signal, thereby smoothing the audio signals between different frames and improving the quality of the speech composed of the audio signals of different frames.
[0226] The specific limitations of the audio signal enhancement device can be referred to the limitations of the audio signal enhancement method in the above, which will not be repeated here. Each module in the above audio signal enhancement device can be realized by software, hardware and combinations thereof, in whole or in part. The above modules can be embedded in or independent of the processor in the computer device in hardware form, or can be stored in the memory in the computer device in software form, so as to call and execute the operations corresponding to each module by the processor.
[0227] In one embodiment, a computer device is provided, which can be a server, and the internal structure diagram thereof can be as shown in Figure 13As shown in the figure. The computer device includes a processor, a memory and a network interface connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operating system and the computer program in the non-volatile storage medium to run. The database of the computer device is used to store voice packet data. The network interface of the computer device is used to communicate with external terminals through network connection. The computer program is executed by the processor to implement an audio signal enhancement method.
[0228] In one embodiment, a computer device is provided, which can be a terminal, and its internal structure diagram can be as shown in the figure. Figure 14 As shown in the figure. The computer device includes a processor, a memory, a communication interface, a display screen and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operating system and the computer program in the non-volatile storage medium to run. The communication interface of the computer device is used to communicate with external terminals in a wired or wireless manner. Wireless mode can be achieved through WIFI, operator network, NFC (near field communication) or other technologies. The computer program is executed by the processor to implement an audio signal enhancement method. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer overlaid on the display screen, or a key, trackball or touchpad arranged on the shell of the computer device, or an external keyboard, touchpad or mouse, etc.
[0229] Those skilled in the art can understand that, Figure 13 Or Figure 14 The structure shown in the figure is only a block diagram of part of the structure related to the scheme of the present application, and does not constitute a limitation on the computer device to which the scheme of the present application is applied. The specific computer device can include more or fewer components than those shown in the figure, or combine certain components, or have a different component arrangement.
[0230] In one embodiment, a computer device is also provided, including a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the steps in each of the above method embodiments.
[0231] In one embodiment, a computer readable storage medium is provided, storing a computer program, which is executed by a processor to implement the steps in each of the above method embodiments.
[0232] In one embodiment, a computer program product or computer program is provided, which includes computer instructions stored in a computer readable storage medium. A processor of a computer device reads the computer instructions from the computer readable storage medium, and the processor executes the computer instructions, so that the computer device performs the steps in each of the above method embodiments.
[0233] A person of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiments can be completed by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer readable storage medium, and when executed, can include the processes of the above-mentioned embodiments. Any reference to memory, storage, database or other medium used in each embodiment provided by the present application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory or optical memory. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration but not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM).
[0234] Each of the technical features of the above embodiments can be combined arbitrarily. In order to make the description simple, not all possible combinations of each technical feature in the above embodiments are described, however, as long as the combination of the technical features does not exist, it should be considered as the scope of the present application.
[0235] The above embodiments only express several implementation manners of the present application, and the description is more specific and detailed, but it should not be understood as a limitation on the scope of the patent. It should be pointed out that for those of ordinary skill in the art, without departing from the concept of the present application, some modifications and improvements can be made, which are all within the scope of the present application. Therefore, the scope of the patent of the present application should be subject to the appended claims.
Claims
1. An audio signal enhancement method, characterized in that, Executed at the receiving end of a voice packet, the method includes: The received voice packets are decoded sequentially to obtain residual signals, long-term filtering parameters, and linear filtering parameters; based on the long-term filtering parameters and the linear filtering parameters, the residual signals are synthesized and filtered to obtain audio signals. Extract feature parameters from the audio signal; Based on the linear filtering parameters, the audio signal is converted into a filter speech excitation signal; Based on the feature parameters, the long-term filtering parameters, and the linear filtering parameters, the filtered speech excitation signal is subjected to speech enhancement processing to obtain the enhanced speech excitation signal. Speech synthesis is performed based on the enhanced speech excitation signal and the linear filtering parameters to obtain the enhanced speech signal.
2. The method according to claim 1, characterized in that, The step of performing signal synthesis filtering on the residual signal based on the long-term filtering parameters and the linear filtering parameters to obtain an audio signal includes: Based on the long-time filtering parameters, the parameters of the long-time prediction filter are configured, and the residual signal is subjected to long-time synthesis filtering through the long-time prediction filter with the configured parameters to obtain the long-time filtering excitation signal. Based on the linear filtering parameters, the linear prediction filter is configured with parameters, and the long-time filtered excitation signal is linearly synthesized and filtered using the parameter-configured linear prediction filter to obtain the audio signal.
3. The method according to claim 2, characterized in that, The process of configuring the linear prediction filter based on the linear filtering parameters, and then using the parameter-configured linear prediction filter to perform linear synthesis filtering on the long-time filtered excitation signal to obtain an audio signal includes: The long-time filtered excitation signal is divided into at least two sub-frames to obtain a sub-long-time filtered excitation signal; The linear filter parameters are grouped to obtain at least two sets of linear filter parameters. Based on the aforementioned set of linear filtering parameters, at least two linear prediction filters are configured with parameters respectively; The obtained sub-long-time filtering excitation signals are respectively input into the parameter-configured linear prediction filter, so that the linear prediction filter performs linear synthesis filtering on the sub-long-time filtering excitation signals based on the linear filtering parameter set to obtain the sub-audio signals corresponding to each sub-frame; The sub-audio signals are combined according to the timing of each sub-frame to obtain an audio signal.
4. The method according to claim 3, characterized in that, The linear filtering parameters include linear filtering coefficients and energy gain values; the method further includes: For the sub-long-term filtered excitation signal corresponding to the first sub-frame in the long-term filtered excitation signal, obtain the energy gain value of the historical sub-long-term filtered excitation signal of the sub-frame adjacent to the sub-long-term filtered excitation signal corresponding to the first sub-frame in the historical long-term filtered excitation signal; Based on the energy gain value corresponding to the historical sub-time-filtered excitation signal and the energy gain value of the sub-time-filtered excitation signal corresponding to the first subframe, the energy adjustment parameter corresponding to the sub-time-filtered excitation signal is determined. The energy of the historical sub-long-term filtered excitation signal is adjusted using the energy adjustment parameters. The step of inputting the obtained sub-long-time filtered excitation signals into the parameter-configured linear prediction filter, so that the linear prediction filter performs linear synthesis filtering on the sub-long-time filtered excitation signals based on the linear filtering parameter set, to obtain sub-audio signals corresponding to each sub-frame, includes: The obtained sub-long-time filtered excitation signal and the historical sub-long-time filtered excitation signal obtained after energy adjustment are input to the linear prediction filter with configured parameters, so that the linear prediction filter performs linear synthesis filtering on the sub-long-time filtered excitation signal corresponding to the first sub-frame based on the linear filtering coefficients and the historical sub-long-time filtered excitation signal obtained after energy adjustment, to obtain the sub-audio signal corresponding to the first sub-frame.
5. The method according to claim 1, characterized in that, The method further includes: Determine whether any data anomalies occurred in historical voice packets decoded prior to the decoded voice packet; If the historical voice packet shows data anomalies, the audio signal obtained after decoding and filtering is determined to be a forward error correction frame signal; Extracting feature parameters from the audio signal includes: When the audio signal is a forward error correction frame signal, feature parameters are extracted from the audio signal.
6. The method according to claim 1, characterized in that, The feature parameters include cepstral feature parameters; Extracting feature parameters from the audio signal includes: Perform a Fourier transform on the audio signal to obtain the Fourier transformed audio signal; The Fourier-transformed audio signal is logarithmically processed to obtain the logarithmic result; Perform an inverse Fourier transform on the logarithmic result to obtain the cepstral characteristic parameters.
7. The method according to claim 6, characterized in that, The long-term filtering parameters include the pitch period and amplitude gain value; The step of performing speech enhancement processing on the filtered speech excitation signal based on the feature parameters, the long-term filtering parameters, and the linear filtering parameters to obtain the enhanced speech excitation signal includes: Based on the pitch period, amplitude gain value, linear filtering parameters, and cepstral feature parameters, the filtered speech excitation signal is subjected to speech enhancement processing to obtain the enhanced speech excitation signal.
8. The method according to claim 1, characterized in that, The process of converting the audio signal into a filter speech excitation signal based on the linear filtering parameters includes: Based on the linear filtering parameters, the linear prediction filter is configured with parameters, and the audio signal is linearly decomposed and filtered using the configured linear prediction filter to obtain the filter speech excitation signal.
9. The method according to claim 1, characterized in that, The step of performing speech enhancement processing on the filtered speech excitation signal based on the feature parameters, the long-term filtering parameters, and the linear filtering parameters to obtain the enhanced speech excitation signal includes: The feature parameters, the long-term filtering parameters, the linear filtering parameters, and the filter speech excitation signal are input into a pre-trained signal enhancement model, so that the signal enhancement model performs speech enhancement processing on the filter speech excitation signal based on the feature parameters to obtain an enhanced speech excitation signal.
10. The method according to claim 9, characterized in that, The feature parameters include cepstral feature parameters; the step of inputting the feature parameters, the long-term filtering parameters, the linear filtering parameters, and the filtered speech excitation signal into a pre-trained signal enhancement model, so that the signal enhancement model performs speech enhancement processing on the filtered speech excitation signal based on the feature parameters to obtain an enhanced speech excitation signal, includes: The cepstral feature parameters, the long-term filtering parameters, and the linear filtering parameters are vectorized, and the results of the vectorization are concatenated to obtain the feature vector. The feature vector and the filter speech excitation signal are input into the pre-trained signal enhancement model; The target feature vector is obtained by extracting features from the feature vector using the signal enhancement model. The enhanced speech excitation signal is obtained by performing enhancement processing on the filter speech excitation signal based on the target feature vector.
11. The method according to claim 10, characterized in that, The enhancement process based on the target feature vector to the filter speech excitation signal to obtain the enhanced speech excitation signal includes: Perform a Fourier transform on the filtered speech excitation signal to obtain a frequency domain speech excitation signal; The amplitude features of the frequency domain speech excitation signal are enhanced based on the target feature vector; The enhanced speech excitation signal is obtained by performing an inverse Fourier transform on the frequency domain speech excitation signal with enhanced amplitude characteristics.
12. The method according to claim 1, characterized in that, The process of synthesizing speech based on the enhanced speech excitation signal and the linear filtering parameters to obtain the enhanced speech signal includes: Based on the linear filtering parameters, the parameters of the linear prediction filter are configured, and the enhanced speech excitation signal is linearly synthesized and filtered by the parameter-configured linear prediction filter to obtain the speech enhancement signal.
13. An audio signal enhancement device, characterized in that, Executed at the receiving end of a voice packet, the device includes: The voice packet processing module is used to decode the received voice packets sequentially to obtain residual signals, long-term filtering parameters, and linear filtering parameters; and to perform signal synthesis filtering on the residual signals based on the long-term filtering parameters and the linear filtering parameters to obtain audio signals. A feature parameter extraction module is used to extract feature parameters from the audio signal; A signal conversion module is used to convert the audio signal into a filter speech excitation signal based on the linear filtering parameters; The speech enhancement module is used to perform speech enhancement processing on the filter speech excitation signal according to the feature parameters, the long-term filtering parameters and the linear filtering parameters to obtain the enhanced speech excitation signal; The speech synthesis module is used to perform speech synthesis based on the enhanced speech excitation signal and the linear filtering parameters to obtain the speech enhancement signal.
14. The apparatus according to claim 13, characterized in that, The voice packet processing module is also used for: Based on the long-time filtering parameters, the parameters of the long-time prediction filter are configured, and the residual signal is subjected to long-time synthesis filtering through the long-time prediction filter with the configured parameters to obtain the long-time filtering excitation signal. Based on the linear filtering parameters, the linear prediction filter is configured with parameters, and the long-time filtered excitation signal is linearly synthesized and filtered using the parameter-configured linear prediction filter to obtain the audio signal.
15. The apparatus according to claim 14, characterized in that, The voice packet processing module is also used for: The long-time filtered excitation signal is divided into at least two sub-frames to obtain a sub-long-time filtered excitation signal; The linear filter parameters are grouped to obtain at least two sets of linear filter parameters. Based on the aforementioned set of linear filtering parameters, at least two linear prediction filters are configured with parameters respectively; The obtained sub-long-time filtering excitation signals are respectively input into the parameter-configured linear prediction filter, so that the linear prediction filter performs linear synthesis filtering on the sub-long-time filtering excitation signals based on the linear filtering parameter set to obtain the sub-audio signals corresponding to each sub-frame; The sub-audio signals are combined according to the timing of each sub-frame to obtain an audio signal.
16. The apparatus according to claim 15, characterized in that, The linear filtering parameters include linear filtering coefficients and energy gain values; the speech packet processing module is also used for: For the sub-long-term filtered excitation signal corresponding to the first sub-frame in the long-term filtered excitation signal, obtain the energy gain value of the historical sub-long-term filtered excitation signal of the sub-frame adjacent to the sub-long-term filtered excitation signal corresponding to the first sub-frame in the historical long-term filtered excitation signal; Based on the energy gain value corresponding to the historical sub-time-filtered excitation signal and the energy gain value of the sub-time-filtered excitation signal corresponding to the first subframe, the energy adjustment parameter corresponding to the sub-time-filtered excitation signal is determined. The energy of the historical sub-long-term filtered excitation signal is adjusted using the energy adjustment parameters. The obtained sub-long-time filtered excitation signal and the historical sub-long-time filtered excitation signal obtained after energy adjustment are input to the linear prediction filter with configured parameters, so that the linear prediction filter performs linear synthesis filtering on the sub-long-time filtered excitation signal corresponding to the first sub-frame based on the linear filtering coefficients and the historical sub-long-time filtered excitation signal obtained after energy adjustment, to obtain the sub-audio signal corresponding to the first sub-frame.
17. The apparatus according to claim 13, characterized in that, The device further includes: The data anomaly determination module is used to determine whether there are data anomalies in historical voice packets decoded before the voice packet is decoded; A forward error correction frame signal determination module is used to determine that the audio signal obtained after decoding and filtering is a forward error correction frame signal if the historical voice packet has data anomalies. The feature parameter extraction module is also used for: When the audio signal is a forward error correction frame signal, feature parameters are extracted from the audio signal.
18. The apparatus according to claim 13, characterized in that, The feature parameters include cepstral feature parameters; The feature parameter extraction module is also used for: Perform a Fourier transform on the audio signal to obtain the Fourier transformed audio signal; The Fourier-transformed audio signal is logarithmically processed to obtain the logarithmic result; Perform an inverse Fourier transform on the logarithmic result to obtain the cepstral characteristic parameters.
19. The apparatus according to claim 18, characterized in that, The long-term filtering parameters include the pitch period and amplitude gain value; The speech enhancement module is also used for: Based on the pitch period, amplitude gain value, linear filtering parameters, and cepstral feature parameters, the filtered speech excitation signal is subjected to speech enhancement processing to obtain the enhanced speech excitation signal.
20. The apparatus according to claim 13, characterized in that, The signal conversion module is also used for: Based on the linear filtering parameters, the linear prediction filter is configured with parameters, and the audio signal is linearly decomposed and filtered using the configured linear prediction filter to obtain the filter speech excitation signal.
21. The apparatus according to claim 13, characterized in that, The voice enhancement module is also used for: The feature parameters, the long-term filtering parameters, the linear filtering parameters, and the filter speech excitation signal are input into a pre-trained signal enhancement model, so that the signal enhancement model performs speech enhancement processing on the filter speech excitation signal based on the feature parameters to obtain an enhanced speech excitation signal.
22. The apparatus according to claim 21, characterized in that, The feature parameters include cepstral feature parameters; The speech enhancement module is also used for: The cepstral feature parameters, the long-term filtering parameters, and the linear filtering parameters are vectorized, and the results of the vectorization are concatenated to obtain the feature vector. The feature vector and the filter speech excitation signal are input into the pre-trained signal enhancement model; The target feature vector is obtained by extracting features from the feature vector using the signal enhancement model. The enhanced speech excitation signal is obtained by performing enhancement processing on the filter speech excitation signal based on the target feature vector.
23. The apparatus according to claim 22, characterized in that, The speech enhancement module is also used for: Perform a Fourier transform on the filtered speech excitation signal to obtain a frequency domain speech excitation signal; The amplitude features of the frequency domain speech excitation signal are enhanced based on the target feature vector; The enhanced speech excitation signal is obtained by performing an inverse Fourier transform on the frequency domain speech excitation signal with enhanced amplitude characteristics.
24. The apparatus according to claim 13, characterized in that, The speech synthesis module is also used for: Based on the linear filtering parameters, the parameters of the linear prediction filter are configured, and the enhanced speech excitation signal is linearly synthesized and filtered by the parameter-configured linear prediction filter to obtain the speech enhancement signal.
25. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 12.
26. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 12.
27. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 12.
Citation Information
Patent Citations
Voice processing method, device and equipment and storage medium
CN111554323A
Artificial intelligence based audio coding
US20210074308A1