A Voice Processing Method, Apparatus, Device, and Storage Medium
By combining traditional signal analysis and processing technology with deep learning technology, the lost voice frames in VoIP system are reconstructed, which solves the shortcomings of traditional PLC technology in the face of continuous multi-frame packet loss, and realizes more efficient voice processing and continuous packet loss compensation.
Patent Information
- Application Number
- CN202010413898.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-05-15
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2040-05-15
AI Technical Summary
In existing VoIP systems, voice signals are prone to sound quality damage during transmission, especially when the network is unstable or sudden packet loss, traditional PLC technology is difficult to effectively handle packet loss in multiple consecutive frames.
A speech processing method combining traditional signal analysis and processing technology and deep learning technology is adopted. By determining the historical speech frame corresponding to the target speech frame, its frequency domain characteristics are obtained, the network model is called for prediction processing to obtain the parameter set of the target speech frame, and the target speech frame is reconstructed through inter-parameter filtering.
This method can effectively improve voice processing capabilities, is suitable for communication scenarios with high real-time requirements, supports continuous packet loss compensation, ensures voice call quality, and can be used in combination with FEC technology to further improve sound quality.
Smart Images

Figure CN111554322B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of Internet technology, specifically to the field of VoIP (Voice over Internet Protocol, voice transmission based on IP) call technology, especially to a voice processing method, a voice processing device, a voice processing equipment and a computer-readable storage medium. Background Art
[0002] The voice signal may be damaged during transmission through the VoIP system. In the prior art, a mainstream solution to the problem of voice quality damage is the classic PLC technology. The main principle is that if the receiving end does not receive the nth (n is a positive integer) voice frame, it will perform signal analysis and processing on the n-1th voice frame to compensate for the nth voice frame. However, practice has found that due to the limited signal analysis and processing capabilities, the voice processing capabilities of the classic PLC technology are limited and cannot be applied to the scenario of sudden packet loss in the existing network. Summary of the invention
[0003] The embodiments of the present application provide a speech processing method, apparatus, device and storage medium, which can make up for the deficiencies of traditional signal analysis and processing technology and improve speech processing capabilities.
[0004] On the one hand, an embodiment of the present application provides a speech processing method, including:
[0005] Determine a historical speech frame corresponding to a target speech frame to be processed;
[0006] Obtain frequency domain features of historical speech frames;
[0007] Calling the network model to predict the frequency domain features of the historical speech frame to obtain a parameter set of the target speech frame; the parameter set includes at least two parameters, the network model includes multiple neural networks, and the number of neural networks is determined according to the number of parameters in the parameter set;
[0008] Reconstruct the target speech frame according to the parameter set.
[0009] On the one hand, an embodiment of the present application provides a speech processing method, including:
[0010] Receive voice signals transmitted via the VoIP system;
[0011] When the target speech frame in the speech signal is lost, the target speech frame is reconstructed using the above method;
[0012] A speech signal is output based on the reconstructed target speech frame.
[0013] On the one hand, an embodiment of the present application provides a speech processing device, including:
[0014] A determination unit, used to determine a historical speech frame corresponding to a target speech frame to be processed;
[0015] An acquisition unit, used for acquiring frequency domain features of historical speech frames;
[0016] A processing unit is used to call a network model to predict the frequency domain features of historical speech frames to obtain a parameter set of a target speech frame; the parameter set includes at least two parameters, the network model includes multiple neural networks, and the number of neural networks is determined according to the number of parameters in the parameter set; and is used to reconstruct the target speech frame according to the parameter set.
[0017] On the one hand, an embodiment of the present application provides another speech processing device, including:
[0018] A receiving unit, used for receiving a voice signal transmitted via the VoIP system;
[0019] A processing unit, used for reconstructing the target speech frame by using the above method when the target speech frame in the speech signal is lost;
[0020] The output unit is used to output a speech signal based on the reconstructed target speech frame.
[0021] On the one hand, an embodiment of the present application provides a speech processing device, the device comprising:
[0022] a processor adapted to implement one or more instructions; and,
[0023] A computer-readable storage medium stores one or more instructions, wherein the one or more instructions are suitable for being loaded by a processor and executing the speech processing method as described above.
[0024] On the one hand, an embodiment of the present application provides a computer-readable storage medium, which stores one or more instructions, and the one or more instructions are suitable for being loaded by a processor and executing the speech processing method as described above.
[0025] In the embodiment of the present application, when it is necessary to reconstruct the target speech frame in the speech signal, the network model can be called to predict the frequency domain features of the historical speech frame corresponding to the target speech frame to obtain the parameter set of the target speech frame, and then the parameter set can be filtered between parameters to achieve the reconstruction of the target speech frame. This speech reconstruction and recovery process combines traditional signal analysis and processing technology with deep learning technology, which makes up for the shortcomings of traditional signal analysis and processing technology and improves the speech processing capability; and based on the prediction of the parameter set of the target speech frame by deep learning of the historical speech frame, the target speech frame is reconstructed according to the parameter set of the target speech frame. The reconstruction process is relatively simple and efficient, and is more suitable for communication scenarios with high real-time requirements; in addition, the parameter set used to reconstruct the target speech frame contains two or more parameters, so that the learning target of the network model is decomposed into several parameters, each parameter corresponds to a different neural network for learning, and different neural networks can be flexibly configured and combined according to different parameter sets to form the structure of the network model. In this way, the network structure can be greatly simplified and the processing complexity can be effectively reduced. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0027] Figure 1 A schematic diagram of the structure of a VoIP system provided by an exemplary embodiment of the present application is shown;
[0028] Figure 2 A schematic diagram of the structure of a speech processing system provided by an exemplary embodiment of the present application is shown;
[0029] Figure 3 A flow chart of a speech processing method provided by an exemplary embodiment of the present application is shown;
[0030] Figure 4 A flowchart of a speech processing method provided by another exemplary embodiment of the present application is shown;
[0031] Figure 5 A flowchart of a speech processing method provided by another exemplary embodiment of the present application is shown;
[0032] Figure 6 A schematic diagram of a STFT provided by an exemplary embodiment of the present application is shown;
[0033] Figure 7 A schematic diagram of the structure of a network model provided by an exemplary embodiment of the present application is shown;
[0034] Figure 8 A schematic diagram of the structure of a speech generation model based on an excitation signal provided by an exemplary embodiment of the present application is shown;
[0035] Fig. 9 A schematic diagram of the structure of a speech processing device provided by an exemplary embodiment of the present application is shown;
[0036] Fig.10 A schematic diagram of the structure of a speech processing device provided by another exemplary embodiment of the present application is shown;
[0037] Fig.11 A structural schematic diagram of a speech processing device provided by an exemplary embodiment of the present application is shown. DETAILED DESCRIPTION
[0038] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0039] The present application embodiment relates to VoIP. VoIP is a voice call technology that uses IP to achieve voice calls and multimedia conferences, that is, to communicate via the Internet. VoIP can also be called IP phone, Internet phone, network phone, broadband phone, and broadband phone service. Figure 1 The structure diagram of a VoIP system provided by an exemplary embodiment of the present application is shown; the system includes a sending end and a receiving end, the sending end refers to a terminal that initiates a voice signal transmitted via the VoIP system; correspondingly, the receiving end refers to a terminal that receives a voice signal transmitted via VoIP; the terminal here may include but is not limited to: a mobile phone, a PC (Personal Computer), a PDA, etc. The processing flow of the voice signal in the VoIP system is roughly as follows:
[0040] On the sender side:
[0041] (1) collecting an input voice signal, which may be collected by a microphone, for example, and the voice signal is an analog signal; performing analog-to-digital conversion on the voice signal to obtain a digital signal;
[0042] (2) Encoding the digital signal to obtain multiple voice frames; here, encoding may refer to OPUS encoding. OPUS is a lossy sound coding format suitable for real-time sound transmission on the network. Its main features include: ① Supporting a sampling rate range from 8000Hz (narrowband signal) to 48000Hz (full-band signal); ② Supporting constant bit rate and variable bit rate; ③ Supporting audio bandwidth from narrowband to full-band; ④ Supporting voice and music; ⑤ Dynamically adjusting the bit rate, audio bandwidth and frame size; ⑤ Having good robustness loss rate and PLC (Packet Loss Concealment) capabilities. Based on OPUS's strong PLC capability and good VoIP sound quality, OPUS encoding is usually used in VoIP systems. The sampling rate Fs in the encoding process can be set according to actual needs. Fs can be 8000Hz (Hertz), 16000Hz, 32000Hz, 48000Hz, etc. Generally, the frame length of a speech frame is determined by the structure of an encoder used in the encoding process. The frame length of a speech frame may be, for example, 10 ms (milliseconds), 20 ms, and the like.
[0043] (3) Encapsulate multiple voice frames into one or more IP packets.
[0044] (4) Send the IP data packet to the receiving end through the network.
[0045] On the receiving side:
[0046] (5) Receive IP data packets transmitted by the network, and decapsulate the received IP data packets to obtain multiple voice frames.
[0047] (6) Decode the speech frame and restore it to digital signal.
[0048] (7) Performing digital-to-analog conversion on the digital signal, restoring it to an analog voice signal and outputting it. The output here may be, for example, played through a speaker.
[0049] The voice signal may be damaged during transmission through the VoIP system. The so-called sound quality damage refers to the phenomenon that after the normal voice signal from the sender is transmitted to the receiver, abnormal conditions such as playback jamming and unsmoothness occur on the receiving end. An important factor causing sound quality damage is network reasons. During the transmission of data packets, the receiving end cannot receive the data packets normally due to network instability or abnormalities, resulting in the loss of voice frames in the data packets, and then the receiving end cannot recover the voice signal, resulting in abnormal conditions such as jamming when outputting the voice signal. In the prior art, there are several mainstream solutions to the phenomenon of sound quality damage:
[0050] One solution involves FEC (Feedforward Error Correction) technology. FEC technology is generally deployed at the transmitting end; its main principle is: after the transmitting end packages and sends the nth (n is a positive integer) voice frame, a certain bandwidth is still allocated in the next data packet to package and send the nth voice frame again. The data packet formed by the repackaging is called a "redundant packet", and the information of the nth voice frame encapsulated in the redundant packet is called the redundant information of the nth voice frame. In order to save transmission bandwidth, the accuracy of the nth voice frame can be reduced, and the information of the low-precision version of the nth voice frame can be packaged into the redundant packet. During the voice transmission process, if the nth voice frame is lost, the receiving end can wait for the redundant packet of the nth voice frame to arrive, and reconstruct the nth voice frame according to the redundant information of the nth voice frame in the redundant packet, and restore the corresponding voice signal. FEC technology can be divided into in-band FEC and out-of-band FEC. The so-called in-band FEC refers to using the idle bytes in a voice frame to store redundant information. Out-of-band FEC refers to the storage of redundant information outside the structure of a voice frame through digital packet encapsulation technology. However, in practice, it is found that the FEC technology has the following shortcomings in solving the problem of sound quality impairment: it requires additional bandwidth to encode redundant information, and the receiving end will increase additional delay while waiting for redundant information; and different encoding mechanisms require specific FEC adaptation, which is costly and inflexible.
[0051] Another solution is the classic PLC technology, which is usually deployed at the receiving end. The main principle of the classic PLC technology is: if the receiving end does not receive the nth voice frame, it will read the n-1th voice frame, and perform signal analysis and processing on the n-1th voice frame to compensate for the nth voice frame. Compared with FEC technology, PLC technology does not require additional bandwidth. However, practice has found that PLC technology still has its shortcomings in solving the problem of sound quality damage: the signal analysis and processing capabilities are limited, and it is only applicable to the case where one voice frame is lost. However, in many cases in the existing network, there is a burst packet loss (that is, multiple consecutive frames are lost). In this case, the above-mentioned PLC-based technology is invalid.
[0052] The embodiment of the present application proposes a voice processing solution, which makes the following improvements on the above-mentioned classic PLC technology: ① Combining traditional signal analysis and processing technology with deep learning technology to improve voice processing capabilities; ② Modeling based on voice signal data, predicting the parameter set of the target voice frame by deep learning the historical voice frame, and then reconstructing the target voice frame according to the parameter set of the target voice frame. The reconstruction process is relatively simple and efficient, and is more suitable for communication scenarios with high real-time requirements; ③ The parameter set used to reconstruct the target voice frame contains two or more parameters, so that the learning target of the network model is decomposed into several parameters, each parameter corresponds to a different neural network for learning, and different neural networks can be flexibly configured and combined according to different parameter sets to form the structure of the network model. In this way, the network structure can be greatly simplified and the processing complexity can be effectively reduced; ④ Support continuous packet loss compensation, that is, in the case of loss of multiple consecutive voice frames, the reconstruction of multiple consecutive voice frames can be achieved to ensure the quality of voice calls; ⑤ Support the combined use with FEC technology, and avoid the adverse effects of sound quality damage in a relatively flexible combination.
[0053] The speech processing solution proposed in the embodiment of the present application will be described in detail below with reference to the accompanying drawings.
[0054] Figure 2 FIG. 1 shows a schematic diagram of a speech processing system provided by an exemplary embodiment of the present application; Figure 2 As shown, the improved PLC technology proposed in the embodiment of the present application is deployed on the downlink receiving end side. The reasons for this deployment are: 1) the receiving end is the last link of the system in end-to-end communication. After the reconstructed target voice frame is restored to a voice signal output (such as played through a speaker, a loudspeaker, etc.), the user can intuitively perceive the voice quality; 2) in the field of mobile communications, the communication link from the downlink air interface to the receiving end is the node most prone to quality problems. Setting a PLC mechanism at this node can achieve a more direct improvement in sound quality.
[0055] Figure 3 A flowchart of a voice processing method provided by an exemplary embodiment of the present application is shown; since the improved PLC technology is deployed at the downlink receiving end, Figure 3 The process shown is Figure 2 The receiving end shown is the execution subject; the method includes the following steps S301-S303.
[0056] S301, receiving a voice signal transmitted via a VoIP system.
[0057] The voice signal is sent from the sending end to the receiving end via the network. As can be seen from the processing flow in the aforementioned VoIP system, the voice signal received by the receiving end is a voice signal in the form of an IP data packet. The receiving end decapsulates the IP data packet to obtain a voice frame.
[0058] S302, when the target speech frame in the speech signal is lost, the target speech frame is reconstructed using the improved PLC technology proposed in the embodiment of the present application. The embodiment of the present application uses the nth speech frame to represent the target speech frame, and the speech processing method involved in the improved PLC technology will be described in detail in subsequent embodiments.
[0059] S303: output a speech signal based on the reconstructed target speech frame.
[0060] After reconstructing the target voice frame, the receiving end will decode and perform digital-to-analog conversion on the target voice frame, and finally play the voice signal through a speaker, etc., thereby realizing the restored output of the voice signal.
[0061] In one embodiment, the improved PLC technology can be used alone. In this case, when the receiving end confirms that the nth voice frame is lost, the packet loss compensation function is activated, and the nth voice frame is reconstructed through the processing flow involved in the improved PLC technology (i.e., the above step S303). In another embodiment, the improved PLC technology can also be used in combination with the FEC technology. In this case, Figure 3 The process shown may also include the following steps S304-S305:
[0062] S304, obtaining redundant information of the target speech frame.
[0063] S305, when the target speech frame in the speech signal is lost, reconstruct the target speech frame according to the redundant information of the target speech frame. If the target speech frame is not reconstructed according to the redundant information of the target speech frame, step S302 is triggered again to reconstruct the target speech frame using the improved PLC technology proposed in the embodiment of the present application.
[0064] In the scenario where the improved PLC technology is used in combination with the FEC technology, an FEC operation is performed at the sending end, that is, not only the n-th voice frame is packaged and sent, but also the redundant information of the n-th voice frame is packaged and sent; when the n-th voice frame is lost, the receiving end first relies on the redundant information of the n-th voice frame to try to reconstruct and restore the n-th voice frame. If the n-th voice frame cannot be successfully restored, the improved PLC function is activated to reconstruct the n-th voice frame through the processing flow involved in the improved PLC technology.
[0065] In an embodiment of the present application, when a target voice frame in a VoIP voice signal is lost, an improved PLC technology can be used to reconstruct the target voice frame. The reconstruction process of the improved PLC technology is relatively simple and efficient, and is more suitable for communication scenarios with high real-time requirements. In addition, it supports continuous packet loss compensation, that is, when multiple consecutive voice frames are lost, it is possible to reconstruct multiple consecutive voice frames to ensure the quality of voice calls. Moreover, the improved PLC technology can also be used in combination with the FEC technology to avoid the adverse effects of sound quality damage in a relatively flexible combination manner.
[0066] The speech processing method involved in the improved PLC technology proposed in the embodiment of the present application will be described in detail below in conjunction with the accompanying drawings.
[0067] Figure 4 A flowchart of a speech processing method provided by another exemplary embodiment of the present application is shown; the method comprises Figure 2 The method is performed by the receiving end shown in the figure; the method includes the following steps S401-S404.
[0068] S401, determining a historical speech frame corresponding to a target speech frame to be processed.
[0069] When there is a voice frame loss in the voice signal transmitted through the VoIP system, the lost voice frame is determined as the target voice frame, and the historical voice frame refers to the voice frame that was transmitted before the target voice frame and can be successfully restored to the voice signal. In the subsequent embodiments of the present application, the target voice frame is the nth (n is a positive integer) frame voice frame in the voice signal transmitted through the VoIP system, and the historical voice frame includes the ntth frame to the n-1th frame in the voice signal transmitted through the VoIP system. A total of t (t is a positive integer) frames of voice frames are used as an example for explanation. The value of t can be set according to actual needs, and the embodiment of the present application does not limit the value of t; for example: if you want to reduce the difficulty of calculation, the value of t can be set to be relatively small, such as t=2, that is, the two adjacent frames before the nth frame are selected as historical voice frames; if you want to obtain more accurate calculation results, the value of t can be set to be relatively large, such as t=n-1, that is, all frames before the nth frame are selected as historical voice frames.
[0070] S402, obtaining frequency domain features of historical speech frames.
[0071] The historical speech frame is a time domain signal. If the frequency domain features of the historical speech frame are to be obtained, the historical speech frame needs to be subjected to time-frequency conversion processing. The time-frequency conversion processing is used to convert the historical speech frame from the time domain space to the frequency domain space, and then the frequency domain features of the historical speech frame can be obtained in the frequency domain space. Here, the time-frequency conversion processing can be implemented by Fourier transform, STFT (Short-Term Fourier Transform) and other operations. Taking the use of STFT operation to perform time-frequency conversion processing on the historical speech frame as an example, the frequency domain features of the historical speech frame may include the STFT coefficients of the historical speech frame. In one embodiment, the frequency domain features of the historical speech frame further include the amplitude spectrum of the STFT coefficients of the historical speech frame to simplify the complexity of the speech processing process.
[0072] S403, calling the network model to predict the frequency domain features of the historical speech frame to obtain a parameter set of the target speech frame; the parameter set includes at least two parameters, the network model includes multiple neural networks, and the number of neural networks is determined according to the number of parameters in the parameter set.
[0073] The parameters in the parameter set refer to the time domain parameters of the target speech frame required for reconstructing and restoring the target speech frame; the parameters in the parameter set may include but are not limited to at least one of the following: short-time correlation parameters, long-time correlation parameters and energy parameters of the target speech frame. The types of target speech frames may include but are not limited to: voiced frames and unvoiced frames; voiced frames belong to quasi-periodic signals, while unvoiced frames belong to non-periodic signals. The types of target speech frames are different, and the parameters required for their reconstruction are also different, so the parameters included in the parameter set of the target speech frame are also different. After determining the parameters in the parameter set according to the type of the target speech frame, the network structure of the network model can be configured accordingly. After the network structure of the network model is configured, the network model can be trained using deep learning methods to obtain an optimized network model. Reuse the optimized network model By predicting the frequency domain features of the historical speech frames, the parameter set Pa(n) of the target speech frame can be obtained.
[0074] S404: Reconstruct the target speech frame according to the parameter set.
[0075] The parameter set Pa(n) contains the predicted time domain parameters of the target speech frame, which are used to represent the time domain characteristics of the time domain signal; then, the target speech frame can be reconstructed and restored using the time domain characteristics of the target speech frame represented by the predicted time domain parameters of the target speech frame. In a specific implementation, the parameters in the parameter set Pa(n) can be subjected to inter-parameter filtering to reconstruct the target speech frame.
[0076] In the embodiment of the present application, when it is necessary to reconstruct the target speech frame in the speech signal, the network model can be called to predict the frequency domain features of the historical speech frame corresponding to the target speech frame to obtain the parameter set of the target speech frame, and then the parameter set can be filtered between parameters to achieve the reconstruction of the target speech frame. This speech reconstruction and recovery process combines traditional signal analysis and processing technology with deep learning technology, which makes up for the shortcomings of traditional signal analysis and processing technology and improves the speech processing capability; and based on the prediction of the parameter set of the target speech frame by deep learning of the historical speech frame, the target speech frame is reconstructed according to the parameter set of the target speech frame. The reconstruction process is relatively simple and efficient, and is more suitable for communication scenarios with high real-time requirements; in addition, the parameter set used to reconstruct the target speech frame contains two or more parameters, so that the learning target of the network model is decomposed into several parameters, each parameter corresponds to a different neural network for learning, and different neural networks can be flexibly configured and combined according to different parameter sets to form the structure of the network model. In this way, the network structure can be greatly simplified and the processing complexity can be effectively reduced.
[0077] For the convenience of description, the following example scenario is used as an example for detailed description in the subsequent embodiments of this application. The example scenario includes the following information: (1) The speech signal is a broadband signal with a sampling rate of Fs=16000 Hz; based on experience, the order of the LPC filter corresponding to the broadband signal with a sampling rate of Fs=16000 Hz is 16; (2) The frame length of the speech frame is 20 ms, and each speech frame contains 320 samples. (3) The 320 sample points of each speech frame are decomposed into two sub-frames, the first sub-frame corresponds to the first 10 ms of the speech frame with a total of 160 sample points, and the second sub-frame corresponds to the last 10 ms of the speech frame with a total of 160 sample points. (4) Each speech frame is divided into 5 ms frames to obtain 4 5 ms sub-frames; based on experience, the order of the LTP filter corresponding to the 5 ms sub-frame is 5. It should be noted that the above-mentioned example scenario is cited only to more clearly describe the process of the speech processing method of the embodiment of the present application, but does not constitute a limitation on the relevant technology of the embodiment of the present application. The speech processing method of the embodiment of the present application is also applicable in other scenarios. For example, in other scenarios, Fs may change accordingly, such as Fs=8000Hz, 32000Hz or 48000Hz; the speech frame may also change accordingly, for example, the frame length may be 10ms, 15ms; the decomposition method of the frame and subframe may change accordingly; for example: when decomposing the speech frame to form a frame, and when dividing the speech frame into a frame to form a subframe, both can be processed according to 5ms, that is, the frame length of the frame and the subframe are both 5ms; and so on. The speech processing process in these other scenarios can refer to the speech processing process in the example scenario of the embodiment of the present application for similar analysis.
[0078] Figure 5 A flowchart of a speech processing method provided by another exemplary embodiment of the present application is shown; the method comprises Figure 2 The method is performed by the receiving end shown in the figure; the method includes the following steps S501-S507.
[0079] S501, determining a historical speech frame corresponding to a target speech frame to be processed.
[0080] The target speech frame refers to the nth speech frame in the speech signal; the historical speech frame includes a total of t speech frames from the nt frame to the n-1th frame in the speech signal, where n and t are both positive integers. The value of t can be set according to actual needs. In this embodiment, t=5. It should be noted that the historical speech frame refers to a speech frame that is transmitted before the target speech frame and can be successfully restored to a speech signal. In one embodiment, the historical speech frame is a speech frame that is completely received by the receiving end and can be normally restored to a speech signal through decoding; in another embodiment, the historical speech frame is a speech frame that has been lost but has been successfully reconstructed through FEC technology, classic PLC technology, the improved PLC technology proposed in the embodiment of the present application, or a combination of the above various technologies. The successfully reconstructed speech frame can be normally decoded to restore the speech signal. Similarly, after the nth speech frame is successfully reconstructed by the speech processing method of the embodiment of the present application, if the n+1th speech frame is lost and needs to be reconstructed, the nth speech frame can be used as a historical speech frame for the n+1th speech frame to help the n+1th speech frame to be reconstructed. Figure 5 As shown, the historical speech frame can be represented as s_prev(n), which represents a sequence of sample points included in the nt frame to the n-1 frame speech frame. In the example shown in this embodiment, it is assumed that t=5, and s_prev(n) has a total of 1600 sample points.
[0081] S502, performing short-time Fourier transform processing on the historical speech frame to obtain frequency domain coefficients corresponding to the historical speech frame.
[0082] S503, extracting the amplitude spectrum from the frequency domain coefficients corresponding to the historical speech frame as the frequency domain feature of the historical speech frame.
[0083] In steps S502-S503, STFT can convert the historical speech frames in the time domain into frequency domain representation. Figure 6 A schematic diagram of a STFT provided by an exemplary embodiment of the present application is shown; Figure 6 In the example shown, t=5, STFT uses a 50% windowing and overlapping operation to eliminate the unevenness between frames. After STFT transformation, the frequency domain coefficients of the historical speech frame are obtained, and the frequency domain coefficients include multiple groups of STFT coefficients; Figure 6As shown, the window function used by STFT can be a Hanning window, and the number of overlapping samples (hop-size) of the window function is 160 points; therefore, this embodiment can obtain 9 groups of STFT coefficients, each group of STFT coefficients includes 320 sample points. In one implementation, the amplitude spectrum can be directly extracted for each group of STFT coefficients, and the extracted amplitude spectrum is composed of an amplitude coefficient sequence and used as the frequency domain feature S_prev(n) of the historical speech frame.
[0084] In another implementation, considering that the STFT coefficients have a symmetrical characteristic, that is, a group of STFT coefficients can be evenly divided into two parts, a part of the STFT coefficients (such as the previous part) can be selected for each group of STFT coefficients to extract the amplitude spectrum, and the extracted amplitude spectrum is composed of an amplitude coefficient sequence and used as the frequency domain feature S_prev(n) of the historical speech frame; in the example shown in this embodiment, the first 161 sample points are selected for each group of STFT coefficients in the 9 groups of STFT coefficients, and the amplitude spectrum corresponding to each selected sample point is calculated, and finally 1449 amplitude coefficients are obtained, and the 1449 amplitude coefficients are composed of an amplitude coefficient sequence and used as the frequency domain feature S_prev(n) of the historical speech frame. In order to simplify the computational complexity, the embodiment of the present application takes the corresponding implementation method when the STFT coefficients have a symmetrical characteristic as an example for explanation.
[0085] In the embodiment of the present application, STFT uses a causal system, that is, it only performs frequency domain feature analysis based on the historical voice frames that have been obtained, and does not use future voice frames (that is, voice frames transmitted after the target voice frames) to perform frequency domain feature analysis. This can ensure real-time communication requirements, making the voice processing solution of the present application suitable for voice call scenarios with high real-time requirements.
[0086] S504, calling the network model to predict the frequency domain features of the historical speech frame to obtain a parameter set of the target speech frame. The parameter set includes at least two parameters, and the network model includes multiple neural networks, the number of which is determined according to the number of types of parameters in the parameter set.
[0087] The definition of each parameter in the parameter set Pa(n) is described in detail below. In the embodiment of the present application, the parameter set Pa(n) includes two or more parameters; further, the parameters in the parameter set Pa(n) are used to establish a reconstruction filter to reconstruct and restore the target speech frame using the reconstruction filter. The core of the reconstruction filter includes an LPC filter and an LTP filter; wherein the LTP filter is responsible for processing parameters related to the long-term correlation of the pitch delay, and the LPC filter is responsible for processing parameters related to the short-term correlation of the linear prediction. Then, the parameters that may be included in the parameter set Pa(n) and the definition of each parameter are as follows:
[0088] (1) Short-term correlation parameters of the target speech frame.
[0089] First, define a p-order filter as shown in Equation 1.1:
[0090] A p (z) = 1 + a 1 z -1 +a 2 z -2 +…+a p z -p Formula 1.1
[0091] In the above formula 1.1, p is the order of the filter. For LPC filter, a j (1≤j≤p) represents the LPC coefficient; for the LTP filter, a j (1≤j≤p) represents the LTP coefficient. z represents the speech signal. Since the LPC filter is responsible for processing parameters related to the short-time correlation of linear prediction, the short-time correlation parameters of the target speech frame can be considered as parameters related to the LPC filter. The LPC filter is implemented based on LP (Linear Prediction) analysis. The so-called LP analysis means that when the target speech frame is filtered by LPC, the filtering result of the nth speech frame is obtained by convolving the p historical speech frames before the nth speech frame with the p-order filter shown in the above formula 1.1; this is in line with the short-time correlation characteristics of speech. According to experience, the order of the LPC filter is p=10 in the scenario with a sampling rate of Fs=8000Hz; in the scenario with a sampling rate of Fs=16000Hz, the order of the LPC filter is p=16.
[0092] In the example shown in this embodiment, the sampling rate Fs=16000 Hz, so p=16 can be taken; the above p-order filter can be further decomposed into the following formula 1.2:
[0093]
[0094] Where P(z) = A p (z)-z -(p+1) A p (z -1 ) Formula 1.3
[0095] Q(z)=A p (z)+z -(p+1) A p (z -1 ) Formula 1.4
[0096] From a physical point of view, P(z) shown in formula 1.3 represents the periodic change law of glottis opening, Q(z) shown in formula 1.4 represents the periodic change law of glottis closing, and P(z) and Q(z) represent the periodic change law of glottis opening and closing.
[0097] The roots formed by the decomposition of the two polynomials P(z) and Q(z) appear alternately in the complex plane, so they are named LSF (Line Spectral Frequency). The LSF is represented by a series of angular frequencies w where the roots of P(z) and Q(z) are distributed on the unit circle in the complex plane. k Assume that the root of P(z) and Q(z) in the complex plane is defined as θ k , then the corresponding angular frequency is defined as follows:
[0098]
[0099] In the above formula 1.5, Re{θ k} represents θ k Real number, Im{θ k} represents θ k imaginary number.
[0100] The line spectrum frequency LSF(n) of the nth speech frame can be calculated by the above formula 1.5. As mentioned above, the line spectrum frequency is a parameter that is strongly correlated with the short-term correlation of speech, so LSF(n) can be used as a parameter in the parameter set Pa(n). In practical applications, speech frames are usually decomposed, that is, the nth speech frame is decomposed into k sub-frames, then the line spectrum frequency LSF(n) of the nth speech frame is decomposed into the line spectrum frequencies LSFk(n) of the k sub-frames; as shown in the example of this embodiment, the target speech frame is divided into two sub-frames, namely the first 10ms and the last 10ms; the LSF(n) of the nth speech frame is decomposed into the line spectrum frequency LSF1(n) of its first sub-frame and the line spectrum frequency LSF2(n) of its second sub-frame. Then, in order to further simplify the computational complexity, in one implementation, the line spectrum frequency LSF2(n) of the second sub-frame of the n-th speech frame can be obtained by the above formula 1.5; then, based on the line spectrum frequency LSF2(n-1) of the second sub-frame of the n-1-th speech frame and the line spectrum frequency LSF2(n) of the second sub-frame of the n-th speech frame, the line spectrum frequency LSF1(n) of the first sub-frame of the n-th speech frame can be obtained by interpolation, and the interpolation factor is expressed as α lsf (n). Thus, the parameter set Pa(n) contains parameter 1 and parameter 2. Parameter 1 refers to the line spectrum frequency LSF2(n) of the second subframe of the target speech frame, which contains 16 LSF coefficients. Parameter 2 refers to the interpolation factor α of the target speech frame. lsf(n), the interpolation factor α lsf (n) can contain 5 candidate values, including 0, 0.25, 0.5, 0.75, and 1.0.
[0101] (2) Long-term correlation parameters of the target speech frame.
[0102] Since the LTP filter is responsible for processing parameters related to the long-term correlation of the pitch delay, the long-term correlation parameters of the target speech frame can be considered as parameters related to the LTP filter. LTP filtering reflects the long-term correlation of speech frames (especially voiced frames), and the long-term correlation is strongly correlated with the pitch delay (Pitch Lag) of the speech frame. The pitch delay reflects the quasi-periodicity of the speech frame, that is, if the pitch delay of the sample point in the target speech frame is to be predicted, it can be obtained by fixing the pitch delay of the sample point in the historical speech frame, and then performing LTP filtering on the fixed pitch delay based on the quasi-periodicity. Therefore, this defines parameters three and four in the parameter set Pa(n). The target speech frame includes m subframes, and the long-term correlation parameters of the target speech frame include the pitch delay and LTP coefficient of each subframe of the target speech frame, and m is a positive integer. In the example shown in this embodiment, m=4, so the parameter set Pa(n) may include parameter three and parameter four. Parameter three refers to the pitch delay of the four subframes of the target speech frame, expressed as pitch(n,0), pitch(n,1), pitch(n,2) and pitch(n,3). Parameter four refers to the LTP coefficients corresponding to the four subframes of the target speech frame. Assuming that the LTP filter is a 5th-order filter, each subframe corresponds to 5 LTP coefficients, so parameter four includes a total of 20 LTP coefficients.
[0103] (3) Energy parameter gain(n) of the target speech frame.
[0104] The energy of different speech frames is also different, and the energy can be reflected by the gain value of each subframe of the speech frame, which defines parameter five in the parameter set Pa(n), and parameter five refers to the energy parameter gain(n) of the target speech frame. In the example shown in this embodiment, the target speech frame includes 4 5ms subframes, and the energy parameter gain(n) of the target speech frame includes the gain values of the 4 5ms subframes, specifically including gain(n,0), gain(n,1), gain(n,2), gain(n,3). Gain(n) is used to amplify the signal of the target speech frame obtained by filtering and reconstruction of the reconstruction filter, so that the reconstructed target speech frame can be amplified to the energy level of the original speech signal, thereby restoring a more accurate target speech frame.
[0105] Referring to step S504, the embodiment of the present application predicts the parameter set Pa(n) of the nth speech frame by calling the network model. Considering the diversity of parameters, different network structures are used for different parameters, that is, the network structure of the network model is determined by the number of parameters included in the parameter set Pa(n). Specifically, the network model includes multiple neural networks, and the number of neural networks is determined according to the number of parameters included in the parameter set Pa(n). Based on the various parameters that may be included in the above parameter set Pa(n); Figure 7 FIG. 1 shows a schematic diagram of a network model provided by an exemplary embodiment of the present application; Figure 7 As shown, the network model includes a first neural network 701 and multiple second neural networks 702, and the second neural network 702 belongs to a subnetwork of the first neural network, that is, the output of the first neural network is used as the input of each second neural network 702. Each second neural network 702 is connected to the first neural network 701; the number of second neural networks 702 corresponds to a parameter in the parameter set, that is, a second neural network 702 can be used to predict a parameter in the parameter set Pa(n). It can be seen that the number of the second neural networks is determined according to the number of parameters in the parameter set. In one embodiment, the first neural network 701 includes a layer of LSTM (Long Short-Term Memory) and three layers of FC (Fully connected layer). The first neural network 701 is used to predict the virtual frequency domain feature S(n) of the target speech frame (i.e., the nth frame of the speech frame). The input of the first neural network 701 is the frequency domain feature S_prev(n) of the historical speech frame obtained in step S503, and the output is the virtual frequency domain feature S(n) of the target speech frame. In the example shown in this embodiment, S(n) is a predicted amplitude coefficient sequence of a virtual 322-dimensional STFT coefficient of the n-th speech frame. In the example shown in this embodiment, the LSTM in the first neural network 701 includes 1 hidden layer and 256 processing units. The first layer FC includes 512 processing units and an activation function. The second layer FC includes 512 processing units and an activation function. The third layer FC includes 322 processing units, which are used to output an amplitude coefficient sequence of a virtual 322-dimensional STFT coefficient of the target speech frame.
[0106] The second neural network 702 is used to predict the parameters of the target speech frame. The input of the second neural network 702 is the virtual frequency domain feature S(n) of the target speech frame output by the first neural network 701, and the output is the various parameters used to reconstruct the target speech frame. In the example shown in this embodiment, each second neural network 702 includes two layers of FC, and the last layer of FC does not include an activation function. The parameters to be predicted by each second neural unit 702 are different, and the structure of FC is also different. Among them, ① in the two-layer FC of the second neural network 702 used to predict parameter one, the first layer FC includes 512 processing units and an activation function; the second layer FC includes 16 processing units, and these 16 processing units are used to output 16 LSF coefficients of parameter one. ② in the two-layer FC of the second neural network 702 used to predict parameter two, the first layer FC includes 256 processing units and an activation function; the second layer FC includes 5 processing units, and these 5 processing units are used to output 5 candidate values of parameter two. ③ In the two-layer FC of the second neural network 702 used to predict parameter three, the first layer FC includes 256 processing units and activation functions; the second layer FC includes 4 processing units, which are used to output the pitch delay of 4 subframes of parameter three. ④ In the two-layer FC of the second neural network 702 used to predict parameter four, the first layer FC includes 512 processing units and activation functions; the second layer FC includes 20 processing units, which are used to output the 20 LTP coefficients included in parameter four.
[0107] based on Figure 7 In the network model shown, in one implementation, step S504 can be refined into the following steps s11-s12:
[0108] s11, calling the first neural network 701 to perform prediction processing on the frequency domain features S_prev(n) of the historical speech frame to obtain the virtual frequency domain features S(n) of the target speech frame.
[0109] s12, inputting the virtual frequency domain features of the target speech frame into at least two second neural networks 702 for prediction processing respectively, and obtaining a parameter set Pa(n) of the target speech frame.
[0110] See also Figure 7 The network model also includes a third neural network 703, which is a parallel network with the first neural network (or the second neural network); the third neural network 703 includes a layer of LSTM and a layer of FC. Figure 7 The network model shown, in another embodiment, the method further includes the following steps s13-s14:
[0111] s13, obtain the energy parameters of the historical speech frames.
[0112] s14, calling the third neural network to predict the energy parameters of the historical speech frames to obtain the energy parameters of the target speech frame, the energy parameters of the target speech frame belong to a parameter set Pa(n) of the target speech frame; the target speech frame includes m subframes, and the energy parameters of the target speech frame include gain values of each subframe of the target speech frame.
[0113] In steps s13-s14, the energy parameters of some or all of the speech frames in the historical speech frames can be used to predict the energy parameters of the target speech frame. In this embodiment, the energy parameters of the historical speech frames are the energy parameters of the n-1th and n-2th speech frames as an example for explanation, and the energy parameter of the n-1th speech frame is expressed as gain(n-1), and the energy parameter of the n-2th speech frame is expressed as gain(n-2). In the example shown in this embodiment, m=4, that is, each speech frame contains 4 5ms subframes; then, the energy parameter gain(n-1) of the n-1th speech frame includes the gain values of the 4 5ms subframes of the n-1th speech frame, specifically including gain(n-1,0), gain(n-1,1), gain(n-1,2), gain(n-1,3); similarly, the energy parameter gain(n-2) of the n-2th speech frame includes the gain values of the 4 5ms subframes of the n-2th speech frame, specifically including gain(n-2,0), gain(n-2,1), gain(n-2,2), gain(n-2,3). Similarly, the energy parameter gain(n) of the nth speech frame includes the gain values of the 4 5mg subframes of the nth speech frame, including gain(n,0), gain(n,1), gain(n,2), gain(n,3). In the example shown in this embodiment, the LSTM in the third neural network includes 128 units; the FC layer includes 4 processing units and an activation function, wherein the 4 processing units are respectively used to output gain values of 4 subframes of the nth speech frame.
[0114] refer to Figure 7 The network structure of the network model shown in the figure can be configured accordingly after the parameters in the parameter set Pa(n) are determined according to actual needs. For example, if the parameter set Pa(n) is determined according to actual needs to include only parameter one, parameter two and parameter five, then the network structure of the network model is composed of a first neural network 701, a second neural network 702 for predicting parameter one, a second neural network 702 for predicting parameter two and a third neural network 703 for predicting parameter five. For another example, if the parameter set Pa(n) is determined according to actual needs to include parameters one to five at the same time, then the network structure of the network model is as follows: Figure 7 After configuring the network structure of the network model, the deep learning method can be used to train the network model to obtain an optimized network model. Reuse the optimized network model The frequency domain features S_prev(n) of the historical speech frames are predicted, and the energy parameters (such as gain(n-1) and gain(n-2)) of the historical speech frames can be further predicted to obtain the parameter set Pa(n) of the target speech frame.
[0115] S505: Establish a reconstruction filter according to the parameter set.
[0116] After obtaining the parameter set Pa(n) of the target speech frame, at least two parameters in the parameter set Pa(n) can be used to establish a reconstruction filter, and the subsequent process of reconstructing the target speech frame can be continued. As mentioned above, the reconstruction filter includes an LTP filter and an LPC filter. The LTP filter can be established using the long-term correlation parameters of the target speech frame (including parameter three and parameter four), and the LPC filter can be established using the short-term correlation parameters of the target speech frame. Referring to formula 1.1 above, the establishment of the filter mainly lies in determining the corresponding coefficients of the filter, and the establishment of the LTP filter lies in determining the LTP coefficients. Parameter four already includes the LTP coefficients, so the LTP filter can be established relatively simply based on parameter four.
[0117] The establishment of the LPC filter is to determine the LPC coefficient; the establishment process of the LPC coefficient is as follows:
[0118] First, parameter 1 refers to the line spectrum frequency LSF2(n) of the second subframe of the target speech frame, which contains a total of 16 LSF coefficients; parameter 2 refers to the interpolation factor α of the target speech frame lsf (n), which may include 5 candidate values, namely 0, 0.25, 0.5, 0.75, and 1.0. Then, the line spectrum frequency LSF1(n) of the first subframe of the target speech frame can be completed by interpolation, and the specific calculation formula is shown in the following formula 1.6:
[0119] LSF1(n)=(1-α LSF (n))·LSF2(n-1)+α LSF (n) LSF2(n) Formula 1.6 The above formula 1.6 indicates that the line spectral frequency LSF1(n) of the first sub-frame of the target speech frame is obtained by weighted summing the line spectral frequency LSF2(n-1) of the second sub-frame of the n-1th frame speech frame and the line spectral frequency LSF2(n) of the second sub-frame of the target speech frame, and the weight is the candidate value of the interpolation factor.
[0120] Secondly, according to the relevant derivation of the above-mentioned formulas 1.1 to 1.5, it can be known that the LPC coefficient and the LSF coefficient are related. Combining formulas 1.1 to 1.5 can respectively obtain the 16th-order LPC coefficient of the first sub-frame 10ms before the target speech frame, that is, LPC1(n); and obtain the 16th-order LPC coefficient of the second sub-frame 10ms after the target speech frame, that is, LPC2(n).
[0121] The LPC coefficients can be determined through the above process, and the LPC filter can be established accordingly.
[0122] S506: Acquire an excitation signal of the target speech frame.
[0123] S507, using a reconstruction filter to filter the excitation signal of the target speech frame to obtain the target speech frame.
[0124] Figure 8 The schematic diagram of the structure of a speech generation model based on an excitation signal provided by an exemplary embodiment of the present application is shown; the physical basis of the speech generation model based on the excitation signal is the human sound generation process, which can be roughly divided into two sub-processes: (1) When a person is speaking, a noise-like impact signal with a certain energy is generated in the human trachea; this impact signal corresponds to the excitation signal, and the excitation signal is a set of random signed noise-like sequences with strong fault tolerance. (2) The impact signal impacts the human vocal cords, producing a quasi-periodic opening and closing; after being amplified by the mouth, a sound is emitted; this process corresponds to a reconstruction filter, and the working principle of the reconstruction filter is to simulate this process to construct a sound. Sounds are divided into unvoiced and voiced sounds. The so-called voiced sound refers to the sound of the vocal cords vibrating when pronouncing; and the unvoiced sound refers to the sound of the vocal cords not vibrating. Taking into account the above characteristics of sound, the above human sound generation process will be further refined: (3) For quasi-periodic signals such as voiced sounds, LTP filters and LPC filters need to be used in the reconstruction process, and the excitation signal will impact the LTP filter and LPC filter respectively; (4) For non-periodic signals such as unvoiced sounds, only LPC filters need to be used in the reconstruction process, and the excitation signal will only impact the LPC filter.
[0125] Based on the above description, it can be known that the excitation signal is a set of random signed noise-like sequences, which is used as a driving source to impact (or excite) the reconstruction filter to generate a target speech frame. In step S506 of the embodiment of the present application, the excitation signal of the historical speech frame can be obtained, and the excitation signal of the target speech frame can be estimated based on the excitation signal of the historical speech frame.
[0126] In one implementation, step S506 may estimate the excitation signal of the target speech frame in a multiplexing manner, and the multiplexing manner may be as shown in the following equation 1.7:
[0127] ex(n)=ex(n-1) Formula 1.7
[0128] In the above formula 1.7, ex(n-1) represents the excitation signal of the n-1th speech frame; ex(n) represents the excitation signal of the target speech frame (ie, the nth speech frame).
[0129] In another implementation, step S506 may estimate the excitation signal of the target speech frame by an average value method, and the average value formula may be expressed as shown in the following formula 1.8:
[0130]
[0131] The above formula 1.8 represents the calculation of the average value of the excitation signal of the historical speech frames from the nt frame to the n-1 frame, and obtains the excitation signal ex(n) of the target speech frame (i.e., the n-th speech frame). In formula 1.8, ex(ni)(1≤i≤t) represents the excitation signal of each speech frame from the nt frame to the n-1 frame.
[0132] In another implementation, step S506 may estimate the excitation signal of the target speech frame by weighted summation, and the weighted summation may be as shown in the following equation 1.9:
[0133]
[0134] The above formula 1.9 represents the weighted summation of the excitation signals of the historical speech frames from the nt frame to the n-1 frame, to obtain the excitation signal ex(n) of the target speech frame (i.e., the nth speech frame). In formula 1.9, ∝ i The weight corresponding to the excitation signal of each speech frame is represented. Taking t=5 as an example, a weight combination can be shown in the following Table 1:
[0135] Table 1: Weight table
[0136] project Weight <![CDATA[∝ 1 ]]> 0.40 <![CDATA[∝ 2 ]]> 0.30 <![CDATA[∝ 3 ]]> 0.15 <![CDATA[∝ 4 ]]> 0.10 <![CDATA[∝ 5 ]]> 0.05
[0137] Combination Figure 8 In one implementation, if the target speech frame is a non-periodic signal such as an unvoiced frame, the reconstruction filter may only include an LPC filter, that is, only the LPC filter is needed to filter the excitation signal of the target speech frame; at this time, the parameter set Pa(n) may include the above-mentioned parameter 1, parameter 2 and parameter 5. Then, the process of generating the target speech frame in step S507 refers to the processing process of the LPC filtering stage, including:
[0138] First, parameter 1 refers to the line spectrum frequency LSF2(n) of the second subframe of the target speech frame, which contains a total of 16 LSF coefficients; parameter 2 refers to the interpolation factor α of the target speech frame lsf(n) may include five candidate values, namely 0, 0.25, 0.5, 0.75, and 1.0. Then the line spectrum frequency LSF1(n) of the first sub-frame of the target speech frame is obtained by calculation according to the above formula 1.6.
[0139] Secondly, according to the relevant derivation of the above-mentioned formulas 1.1 to 1.5, it can be known that the LPC coefficient and the LSF coefficient are related. Combining formulas 1.1 to 1.5 can respectively obtain the 16th-order LPC coefficient of the first sub-frame 10ms before the target speech frame, that is, LPC1(n); and obtain the 16th-order LPC coefficient of the second sub-frame 10ms after the target speech frame, that is, LPC2(n).
[0140] Again, under the impact of the excitation signal of the target speech frame, LPC1(n) is subjected to LPC filtering to reconstruct the first 10ms of the target speech frame, which is a total of 160 sample points, and gain(n,0) and gain(n,1) are called to amplify the first 160 sample points to obtain the first 160 sample points of the reconstructed target speech frame. Similarly, LPC2(n) is subjected to LPC filtering to reconstruct the last 10ms of the target speech frame, which is a total of 160 sample points, and gain(n,2) and gain(n,3) are called to amplify the last 160 sample points to obtain the last 160 sample points of the reconstructed target speech frame. The first 10ms and the last 10ms of the target speech frame are synthesized to obtain a complete target speech frame.
[0141] In the above LPC filtering process, the LPC filtering of the nth speech frame uses the LSF coefficients of the n-1th speech frame. That is to say, the LPC filtering of the nth speech frame needs to be implemented using the historical speech frames adjacent to the nth speech frame, which confirms the short-time correlation characteristics of the LPC filtering.
[0142] In another embodiment, if the target speech frame is a periodic signal such as a voiced frame, the reconstruction filter includes an LPC filter and an LTP filter, that is, the LTP filter and the LPC filter need to be used together to filter the excitation signal of the target speech frame. At this time, the parameter set Pa(n) may include the above-mentioned parameter one, parameter two, parameter three, parameter four and parameter five. Then, the process of generating the target speech frame in step S507 includes:
[0143] (I) LTP filtering stage:
[0144] First, parameter three includes the pitch delays of four subframes, namely pitch(n,0), pitch(n,1), pitch(n,2) and pitch(n,3). The pitch delay of each subframe is processed as follows: ① Compare the pitch delay of the subframe with the preset threshold. If the pitch delay of the subframe is lower than the preset threshold, the pitch delay of the subframe is set to 0, and the LTP filtering step is omitted. ② If the pitch delay of the subframe is not lower than the preset threshold, then take the historical sample point corresponding to the subframe, set the order of the LTP filter to 5, and call the 5th-order LTP filter to perform LTP filtering on the historical sample point corresponding to the subframe to obtain the LTP filtering result of the subframe. Since the LTP filter reflects the long-term correlation of the speech frame, and the long-term correlation is strongly correlated with the pitch delay, the historical sample points corresponding to the subframe in the LTP filter involved in the above step ② are selected with reference to the pitch delay of the subframe. Specifically, the subframe is used as the starting point, and the number of sample points corresponding to the value of the pitch delay is traced back as the historical sample points corresponding to the subframe. For example, if the value of the pitch delay of the subframe is 100, then the historical sample points corresponding to the subframe refer to the 100 sample points traced back from the subframe. It can be seen that setting the historical sample points corresponding to the subframe with reference to the pitch delay of the subframe actually uses the sample points contained in the historical subframe before the subframe (such as the previous 5ms subframe) to perform LTP filtering, which confirms the long-term correlation characteristics of the LTP filter.
[0145] Secondly, the LTP filtering results of each subframe are synthesized, including synthesizing the LTP filtering results of the first subframe and the LTP filtering results of the second subframe to obtain the LTP synthetic signal of the first subframe of the first 10ms of the target speech frame; synthesizing the LTP filtering results of the third subframe and the LTP filtering results of the fourth subframe to obtain the LTP synthetic signal of the second subframe of the last 10ms of the target speech frame; thus completing the processing of the LTP filtering stage.
[0146] (II) LPC filtering stage:
[0147] Referring to the processing process of the LPC filtering stage in the above implementation mode, firstly, the 16th-order LPC coefficients of the first sub-frame 10 ms before the target speech frame are obtained based on parameters one and two, i.e., LPC1(n); and the 16th-order LPC coefficients of the second sub-frame 10 ms after the target speech frame are obtained, i.e., LPC2(n).
[0148] Secondly, the LTP synthetic signal of the first subframe of the first 10ms of the target speech frame obtained in the LTP filtering stage is used together with LPC1(n) to perform LPC filtering, and the first 10ms of the target speech frame is reconstructed with a total of 160 sample points, and gain(n,0) and gain(n,1) are called to amplify the first 160 sample points to obtain the first 160 sample points of the reconstructed target speech frame. Similarly, the LTP synthetic signal of the second subframe of the last 10ms of the target speech frame obtained in the LTP filtering stage is used together with LPC2(n) to perform LPC filtering, and the last 10ms of the target speech frame is reconstructed with a total of 160 sample points, and gain(n,2) and gain(n,3) are called to amplify the last 160 sample points to obtain the last 160 sample points of the reconstructed target speech frame. The first 10ms and the last 10ms of the target speech frame are synthesized to obtain a complete target speech frame.
[0149] Through the above description of this embodiment, when the nth voice frame in the voice signal needs to be PLC, the voice processing method based on this embodiment can reconstruct the nth voice frame. If continuous packet loss occurs, for example, the n+1th voice frame, the n+2th voice frame, etc. are lost, the n+1th voice frame, the n+2th voice frame, etc. can be reconstructed and restored according to the above process to achieve continuous packet loss compensation and ensure voice call quality.
[0150] In an embodiment of the present application, when it is necessary to reconstruct a target speech frame in a speech signal, the network model can be called to predict the frequency domain features of the historical speech frame corresponding to the target speech frame to obtain a parameter set of the target speech frame, and then the parameter set can be inter-parameter filtered to achieve reconstruction of the target speech frame. This speech reconstruction and recovery process combines traditional signal analysis and processing technology with deep learning technology, which makes up for the shortcomings of traditional signal analysis and processing technology and improves speech processing capabilities. It also predicts the parameter set of the target speech frame based on deep learning of historical speech frames, and then reconstructs the target speech frame according to the parameter set of the target speech frame. The reconstruction process is relatively simple and efficient, and is more suitable for communication scenarios with high real-time requirements. In addition, the parameter set used to reconstruct the target speech frame contains two or more parameters, which decomposes the learning target of the network model into several parameters. Each parameter corresponds to a different neural network for learning. Different neural networks can be flexibly configured and combined according to different parameter sets to form the structure of the network model. In this way, the network structure can be greatly simplified, the processing complexity can be effectively reduced, and continuous packet loss compensation is supported. That is, in the case of loss of multiple consecutive speech frames, multiple consecutive speech frames can be reconstructed to ensure the quality of voice calls.
[0151] Fig. 9The structure diagram of a speech processing device provided by an exemplary embodiment of the present application is shown; the speech processing device can be used as a computer program (including program code) running in a terminal, for example, the speech processing device can be an application program in the terminal (such as an App providing a VoIP call function); the terminal running the speech processing device can be used as Figure 1 or Figure 2 The receiving end shown; the speech processing device can be used to perform Figure 4 and Figure 5 Some or all of the steps in the method embodiment shown. Fig. 9 , the speech processing device comprises the following units:
[0152] A determination unit 901, configured to determine a historical speech frame corresponding to a target speech frame to be processed;
[0153] An acquisition unit 902 is used to acquire frequency domain features of historical speech frames;
[0154] Processing unit 903 is used to call the network model to predict the frequency domain features of the historical speech frame to obtain a parameter set of the target speech frame; the parameter set includes at least two parameters, the network model includes multiple neural networks, and the number of neural networks is determined according to the number of parameters in the parameter set; and is used to reconstruct the target speech frame according to the parameter set.
[0155] In one implementation, the acquisition unit 902 is specifically used to perform short-time Fourier transform processing on the historical speech frame to obtain frequency domain coefficients corresponding to the historical speech frame; and extract the amplitude spectrum from the frequency domain coefficients corresponding to the historical speech frame as the frequency domain features of the historical speech frame.
[0156] In one implementation, the network model includes a first neural network and at least two second neural networks, the second neural network is a subnetwork of the first neural network; one second neural network corresponds to one parameter in the parameter set; the processing unit 903 is specifically used to:
[0157] Calling the first neural network to perform prediction processing on the frequency domain features of the historical speech frame to obtain the virtual frequency domain features of the target speech frame; and,
[0158] The virtual frequency domain features of the target speech frame are respectively input into at least two second neural networks for prediction processing to obtain at least two parameters in the parameter set of the target speech frame.
[0159] In one implementation, the processing unit 903 is specifically configured to:
[0160] Establishing a reconstruction filter according to the parameter set;
[0161] Obtaining an excitation signal of a target speech frame;
[0162] The reconstruction filter is used to filter the excitation signal of the target speech frame to obtain the target speech frame.
[0163] In one implementation, the processing unit 903 is specifically configured to:
[0164] Obtaining an excitation signal of a historical speech frame;
[0165] The excitation signal of the target speech frame is estimated according to the excitation signal of the historical speech frame.
[0166] In one implementation, the target voice frame refers to the nth voice frame in the voice signal transmitted via the VoIP system; the historical voice frame includes t voice frames from the nt frame to the n-1th frame in the voice signal transmitted via the VoIP system, where n and t are both positive integers.
[0167] In one implementation, the excitation signal of the historical speech frame includes the excitation signal of the n-1th speech frame; the processing unit 903 is specifically configured to determine the excitation signal of the n-1th speech frame as the excitation signal of the target speech frame.
[0168] In one implementation, the excitation signal of the historical speech frame includes the excitation signal of each speech frame from the nt frame to the n-1 frame; the processing unit 903 is specifically used to: calculate the average value of the excitation signals of the t frames of speech frames from the nt frame to the n-1 frame, to obtain the excitation signal of the target speech frame.
[0169] In one implementation, the excitation signal of the historical speech frame includes the excitation signal of each speech frame from the nt frame to the n-1 frame; the processing unit 903 is specifically used to: perform weighted summation on the excitation signals of the t frames of speech frames from the nt frame to the n-1 frame, to obtain the excitation signal of the target speech frame.
[0170] In one implementation, if the target speech frame is an unvoiced frame, the parameter set includes a short-time correlation parameter of the target speech frame; the reconstruction filter includes a linear predictive coding filter;
[0171] The target speech frame includes k sub-frames, and the short-time correlation parameters of the target speech frame include the line spectrum frequency and the interpolation factor of the kth sub-frame of the target speech frame, where k is an integer greater than 1.
[0172] In one implementation, if the target speech frame is a voiced frame, the parameter set includes a short-time correlation parameter of the target speech frame and a long-time correlation parameter of the target speech frame; the reconstruction filter includes a long-time prediction filter and a linear prediction coding filter;
[0173] The target speech frame includes k sub-frames, and the short-time correlation parameters of the target speech frame include the line spectrum frequency and the interpolation factor of the kth sub-frame of the target speech frame, where k is an integer greater than 1;
[0174] The target speech frame includes m subframes, and the long-term correlation parameters of the target speech frame include the pitch delay and the long-term prediction coefficient of each subframe of the target speech frame, and m is a positive integer.
[0175] In one implementation, the network model further includes a third neural network, and the third neural network and the first neural network are parallel networks; the processing unit 903 is further configured to:
[0176] Obtain energy parameters of historical speech frames;
[0177] Calling a third neural network to predict the energy parameters of the historical speech frames to obtain the energy parameters of the target speech frames, where the energy parameters of the target speech frames belong to a parameter set of the target speech frames;
[0178] The target speech frame includes m subframes, and the energy parameter of the target speech frame includes the gain value of each subframe of the target speech frame.
[0179] In an embodiment of the present application, when it is necessary to reconstruct a target speech frame in a speech signal, the network model can be called to predict the frequency domain features of the historical speech frame corresponding to the target speech frame to obtain a parameter set of the target speech frame, and then the parameter set can be subjected to inter-parameter filtering to achieve reconstruction of the target speech frame. This speech reconstruction and recovery process combines traditional signal analysis and processing technology with deep learning technology, which makes up for the shortcomings of traditional signal analysis and processing technology and improves speech processing capabilities. It also predicts the parameter set of the target speech frame based on deep learning of historical speech frames, and then reconstructs the target speech frame according to the parameter set of the target speech frame. The reconstruction process is relatively simple and efficient, and is more suitable for communication scenarios with high real-time requirements. In addition, the parameter set used to reconstruct the target speech frame contains two or more parameters, which decomposes the learning target of the network model into several parameters. Each parameter corresponds to a different neural network for learning. Different neural networks can be flexibly configured and combined according to different parameter sets to form the structure of the network model. In this way, the network structure can be greatly simplified, the processing complexity can be effectively reduced, and continuous packet loss compensation is supported. That is, in the case of loss of multiple consecutive speech frames, multiple consecutive speech frames can be reconstructed to ensure the quality of voice calls.
[0180] Fig.10 A schematic diagram of the structure of a voice processing device provided by another exemplary embodiment of the present application is shown; the voice processing device can be used as a computer program (including program code) running in a terminal, for example, the voice processing device can be an application program in the terminal (such as an App that provides a VoIP call function); the terminal running the voice processing device can be used as Figure 1 or Figure 2 The receiving end shown; the speech processing device can be used to perform Figure 3 Some or all of the steps in the method embodiment shown. Fig.10 , the speech processing device comprises the following units:
[0181] The receiving unit 1001 is used to receive a voice signal transmitted via the VoIP system;
[0182] Processing unit 1002, used for when the target speech frame in the speech signal is lost, Figure 4 or Figure 5 The method shown reconstructs the target speech frame;
[0183] The output unit 1003 is configured to output a speech signal based on the reconstructed target speech frame.
[0184] In one implementation, the processing unit 1002 is further configured to:
[0185] Obtaining redundant information of a target speech frame;
[0186] When a target speech frame in a speech signal is lost, the target speech frame is reconstructed according to redundant information of the target speech frame;
[0187] If the target speech frame cannot be reconstructed based on the redundant information of the target speech frame, Figure 4 or Figure 5 The method shown reconstructs the target speech frame.
[0188] In an embodiment of the present application, when a target voice frame in a VoIP voice signal is lost, an improved PLC technology can be used to reconstruct the target voice frame. The reconstruction process of the improved PLC technology is relatively simple and efficient, and is more suitable for communication scenarios with high real-time requirements. In addition, it supports continuous packet loss compensation, that is, when multiple consecutive voice frames are lost, it is possible to reconstruct multiple consecutive voice frames to ensure the quality of voice calls. Moreover, the improved PLC technology can also be used in combination with the FEC technology to avoid the adverse effects of sound quality damage in a relatively flexible combination manner.
[0189] Fig.11 FIG. 1 shows a schematic diagram of the structure of a speech processing device provided by an exemplary embodiment of the present application. Fig.11 , the speech processing device may be Figure 1 or Figure 2At the receiving end shown, the speech processing device includes a processor 1101, an input device 1102, an output device 1103 and a computer-readable storage medium 1104. The processor 1101, the input device 1102, the output device 1103 and the computer-readable storage medium 1104 can be connected via a bus or other means. The computer-readable storage medium 1104 can be stored in the memory of the speech processing device, the computer-readable storage medium 1104 is used to store a computer program, the computer program includes program instructions, and the processor 111 is used to execute the program instructions stored in the computer-readable storage medium 1104. The processor 1101 (or CPU (Central Processing Unit)) is the computing core and control core of the speech processing device, which is suitable for implementing one or more instructions, and is specifically suitable for loading and executing one or more instructions to implement the corresponding method flow or corresponding function.
[0190] The embodiment of the present application also provides a computer-readable storage medium (Memory), wherein the computer-readable storage medium is a memory device in a speech processing device for storing programs and data. It is understandable that the computer-readable storage medium here can include both the built-in storage medium in the speech processing device and the extended storage medium supported by the speech processing device. The computer-readable storage medium provides a storage space, which stores the operating system of the speech processing device. In addition, one or more instructions suitable for being loaded and executed by the processor 1101 are also stored in the storage space, and these instructions can be one or more computer programs (including program codes). It should be noted that the computer-readable storage medium here can be a high-speed RAM memory, or a non-volatile memory (non-volatile memory), such as at least one disk storage; optionally, it can also be at least one computer-readable storage medium located away from the aforementioned processor.
[0191] In one embodiment, the computer-readable storage medium stores one or more instructions; the processor 1101 loads and executes the one or more instructions stored in the computer-readable storage medium to implement Figure 4 or Figure 5 Corresponding steps of the speech processing method in the illustrated embodiment; in a specific implementation, one or more instructions in the computer-readable storage medium are loaded by the processor 1101 and execute the following steps:
[0192] Determine a historical speech frame corresponding to a target speech frame to be processed;
[0193] Obtain frequency domain features of historical speech frames;
[0194] Calling the network model to predict the frequency domain features of the historical speech frame to obtain a parameter set of the target speech frame; the parameter set includes at least two parameters, the network model includes multiple neural networks, and the number of neural networks is determined according to the number of parameters in the parameter set;
[0195] Reconstruct the target speech frame according to the parameter set.
[0196] In one implementation, when one or more instructions in the computer-readable storage medium are loaded by the processor 1101 and the step of obtaining the frequency domain features of the historical speech frame is executed, the following steps are specifically performed:
[0197] Performing short-time Fourier transform processing on the historical speech frame to obtain the frequency domain coefficients corresponding to the historical speech frame;
[0198] The amplitude spectrum is extracted from the frequency domain coefficients corresponding to the historical speech frame as the frequency domain feature of the historical speech frame.
[0199] In one implementation, the network model includes a first neural network and at least two second neural networks, the second neural network being a subnetwork of the first neural network; one second neural network corresponds to a parameter in a parameter set; one or more instructions in a computer-readable storage medium are loaded and executed by the processor 1101 to call the network model to perform prediction processing on the frequency domain features of the historical speech frames, and when the parameter set of the target speech frame is obtained, the following steps are specifically performed:
[0200] Calling the first neural network to perform prediction processing on the frequency domain features of the historical speech frame to obtain the virtual frequency domain features of the target speech frame;
[0201] The virtual frequency domain features of the target speech frame are respectively input into at least two second neural networks for prediction processing to obtain at least two parameters in the parameter set of the target speech frame.
[0202] In one implementation, when one or more instructions in the computer-readable storage medium are loaded by the processor 1101 and the step of reconstructing the target speech frame according to the parameter set is executed, the following steps are specifically performed:
[0203] Establishing a reconstruction filter according to the parameter set;
[0204] Obtaining an excitation signal of a target speech frame;
[0205] The reconstruction filter is used to filter the excitation signal of the target speech frame to obtain the target speech frame.
[0206] In one implementation, when one or more instructions in the computer-readable storage medium are loaded by the processor 1101 and the step of obtaining the excitation signal of the target speech frame is executed, the following steps are specifically performed:
[0207] Obtaining an excitation signal of a historical speech frame;
[0208] The excitation signal of the target speech frame is estimated according to the excitation signal of the historical speech frame.
[0209] In one implementation, the target voice frame refers to the nth voice frame in the voice signal transmitted via the VoIP system;
[0210] The historical voice frames include t voice frames, from the nt frame to the n-1 frame, in total, in the voice signal transmitted via the VoIP system, where n and t are both positive integers.
[0211] In one embodiment, the excitation signal of the historical speech frame includes the excitation signal of the n-1th frame speech frame; when one or more instructions in the computer-readable storage medium are loaded by the processor 1101 and execute the step of estimating the excitation signal of the target speech frame based on the excitation signal of the historical speech frame, the following steps are specifically performed: the excitation signal of the n-1th frame speech frame is determined as the excitation signal of the target speech frame.
[0212] In one embodiment, the excitation signal of the historical speech frame includes the excitation signal of each speech frame from the nt frame to the n-1 frame; when one or more instructions in the computer-readable storage medium are loaded by the processor 1101 and execute the step of estimating the excitation signal of the target speech frame based on the excitation signal of the historical speech frame, the following steps are specifically performed: the excitation signal of the t-th frame to the n-1-th frame, a total of t frames of speech frames, is averaged to obtain the excitation signal of the target speech frame.
[0213] In one implementation, the excitation signal of the historical speech frame includes the excitation signal of each speech frame from the nt frame to the n-1 frame; when one or more instructions in the computer-readable storage medium are loaded by the processor 1101 and execute the step of estimating the excitation signal of the target speech frame based on the excitation signal of the historical speech frame, the following steps are specifically performed: weighted summation of the excitation signals of the nt frame to the n-1 frame, a total of t frames of speech frames, is performed to obtain the excitation signal of the target speech frame.
[0214] In one implementation, if the target speech frame is an unvoiced frame, the parameter set includes a short-time correlation parameter of the target speech frame; the reconstruction filter includes a linear predictive coding filter;
[0215] The target speech frame includes k sub-frames, and the short-time correlation parameters of the target speech frame include the line spectrum frequency and the interpolation factor of the kth sub-frame of the target speech frame, where k is an integer greater than 1.
[0216] In one implementation, if the target speech frame is a voiced frame, the parameter set includes a short-time correlation parameter of the target speech frame and a long-time correlation parameter of the target speech frame; the reconstruction filter includes a long-time prediction filter and a linear prediction coding filter;
[0217] The target speech frame includes k sub-frames, and the short-time correlation parameters of the target speech frame include the line spectrum frequency and the interpolation factor of the kth sub-frame of the target speech frame, where k is an integer greater than 1;
[0218] The target speech frame includes m subframes, and the long-term correlation parameters of the target speech frame include the pitch delay and the long-term prediction coefficient of each subframe of the target speech frame, and m is a positive integer.
[0219] In one implementation, the network model further includes a third neural network, and the third neural network and the first neural network are parallel networks; one or more instructions in the computer-readable storage medium are loaded by the processor 1101 and the following steps are also performed:
[0220] Obtain energy parameters of historical speech frames;
[0221] Calling a third neural network to predict the energy parameters of the historical speech frames to obtain the energy parameters of the target speech frames, where the energy parameters of the target speech frames belong to a parameter set of the target speech frames;
[0222] The target speech frame includes m subframes, and the energy parameter of the target speech frame includes the gain value of each subframe of the target speech frame.
[0223] In an embodiment of the present application, when it is necessary to reconstruct a target speech frame in a speech signal, the network model can be called to predict the frequency domain features of the historical speech frame corresponding to the target speech frame to obtain a parameter set of the target speech frame, and then the parameter set can be subjected to inter-parameter filtering to achieve reconstruction of the target speech frame. This speech reconstruction and recovery process combines traditional signal analysis and processing technology with deep learning technology, which makes up for the shortcomings of traditional signal analysis and processing technology and improves speech processing capabilities. It also predicts the parameter set of the target speech frame based on deep learning of historical speech frames, and then reconstructs the target speech frame according to the parameter set of the target speech frame. The reconstruction process is relatively simple and efficient, and is more suitable for communication scenarios with high real-time requirements. In addition, the parameter set used to reconstruct the target speech frame contains two or more parameters, which decomposes the learning target of the network model into several parameters. Each parameter corresponds to a different neural network for learning. Different neural networks can be flexibly configured and combined according to different parameter sets to form the structure of the network model. In this way, the network structure can be greatly simplified, the processing complexity can be effectively reduced, and continuous packet loss compensation is supported. That is, in the case of loss of multiple consecutive speech frames, multiple consecutive speech frames can be reconstructed to ensure the quality of voice calls.
[0224] In another embodiment, the processor 1101 loads and executes one or more instructions stored in a computer-readable storage medium to implement Figure 3 Corresponding steps of the speech processing method in the illustrated embodiment; in a specific implementation, one or more instructions in the computer-readable storage medium are loaded by the processor 1101 and execute the following steps:
[0225] Receive voice signals transmitted via the VoIP system;
[0226] When the target speech frame in the speech signal is lost, the Figure 4 or Figure 5 The method shown reconstructs the target speech frame;
[0227] A speech signal is output based on the reconstructed target speech frame.
[0228] In one embodiment, one or more instructions in the computer readable storage medium are loaded by the processor 1101 and further perform the following steps:
[0229] Obtaining redundant information of a target speech frame;
[0230] When a target speech frame in a speech signal is lost, the target speech frame is reconstructed according to redundant information of the target speech frame;
[0231] If the target speech frame cannot be reconstructed based on the redundant information of the target speech frame, the Figure 4 or Figure 5 The method shown reconstructs the target speech frame.
[0232] In an embodiment of the present application, when a target voice frame in a VoIP voice signal is lost, an improved PLC technology can be used to reconstruct the target voice frame. The reconstruction process of the improved PLC technology is relatively simple and efficient, and is more suitable for communication scenarios with high real-time requirements. In addition, it supports continuous packet loss compensation, that is, when multiple consecutive voice frames are lost, it is possible to reconstruct multiple consecutive voice frames to ensure the quality of voice calls. Moreover, the improved PLC technology can also be used in combination with the FEC technology to avoid the adverse effects of sound quality damage in a relatively flexible combination manner.
[0233] A person skilled in the art can understand that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program, and the program can be stored in a computer-readable storage medium, and when the program is executed, it can include the processes of the embodiments of the above-mentioned methods. The storage medium can be a disk, an optical disk, a read-only memory (ROM) or a random access memory (RAM), etc.
[0234] The above disclosure is only the preferred embodiment of the present application, which certainly cannot be used to limit the scope of rights of the present application. Therefore, equivalent changes made according to the claims of the present application are still within the scope covered by the present application.
Claims
1. A voice processing method, characterized in that, it includes: determining historical speech frames corresponding to a target speech frame to be processed; obtaining frequency domain features of the historical speech frames; invoking a network model to perform prediction processing on the frequency domain features of the historical speech frames to obtain a parameter set of the target speech frame; at least two parameters are included in the parameter set, the network model includes multiple neural networks, and the number of neural networks is determined according to the number of parameters in the parameter set; reconstructing the target speech frame according to the parameter set; wherein, the network model includes a first neural network and at least two second neural networks, and the second neural networks belong to sub-networks of the first neural network; one of the second neural networks corresponds to one type of parameter in the parameter set; the invoking the network model to perform prediction processing on the frequency domain features of the historical speech frames to obtain the parameter set of the target speech frame includes: invoking the first neural network to perform prediction processing on the frequency domain features of the historical speech frames to obtain virtual frequency domain features of the target speech frame; inputting the virtual frequency domain features of the target speech frame into the at least two second neural networks respectively for prediction processing to obtain at least two parameters in the parameter set of the target speech frame.
2. The method according to claim 1, characterized in that, the obtaining the frequency domain features of the historical speech frames includes: performing short-time Fourier transform processing on the historical speech frames to obtain frequency domain coefficients corresponding to the historical speech frames; extracting an amplitude spectrum from the frequency domain coefficients corresponding to the historical speech frames as the frequency domain features of the historical speech frames.
3. The method according to claim 1, characterized in that, the reconstructing the target speech frame according to the parameter set includes: establishing a reconstruction filter according to the parameter set; obtaining an excitation signal of the historical speech frames; estimating an excitation signal of the target speech frame according to the excitation signal of the historical speech frames; performing filtering processing on the excitation signal of the target speech frame by using the reconstruction filter to obtain the target speech frame.
4. The method according to claim 3, characterized in that, the target speech frame refers to the nth speech frame in a voice signal transmitted through a VoIP system; the historical speech frames include a total of t speech frames from the (n - t)th frame to the (n - 1)th frame in the voice signal transmitted through the VoIP system, and both n and t are positive integers; the excitation signal of the historical speech frames includes the excitation signal of the (n - 1)th speech frame; the estimating the excitation signal of the target speech frame according to the excitation signal of the historical speech frames includes: determining the excitation signal of the (n - 1)th speech frame as the excitation signal of the target speech frame.
5. The method according to claim 3, characterized in that, The target speech frame refers to the nth speech frame in the speech signal transmitted through the VoIP system; the historical speech frames include a total of t speech frames from the (n - t)th frame to the (n - 1)th frame in the speech signal transmitted through the VoIP system, where both n and t are positive integers; the excitation signals of the historical speech frames include the excitation signals of each speech frame from the (n - t)th frame to the (n - 1)th frame; estimating the excitation signal of the target speech frame according to the excitation signals of the historical speech frames includes: Calculating the average value of the excitation signals of the t speech frames from the (n - t)th frame to the (n - 1)th frame to obtain the excitation signal of the target speech frame; or, Performing weighted summation on the excitation signals of the t speech frames from the (n - t)th frame to the (n - 1)th frame to obtain the excitation signal of the target speech frame.
6. The method according to claim 3, characterized in that, if the target speech frame is an unvoiced frame, the parameter set includes the short-term correlation parameters of the target speech frame; the reconstruction filter includes a linear predictive coding filter; the target speech frame includes k sub-frames, and the short-term correlation parameters of the target speech frame include the line spectral frequencies and interpolation factors of the kth sub-frame of the target speech frame, where k is an integer greater than 1.
7. The method according to claim 3, characterized in that, if the target speech frame is a voiced frame, the parameter set includes the short-term correlation parameters and the long-term correlation parameters of the target speech frame; the reconstruction filter includes a long-term prediction filter and a linear predictive coding filter; the target speech frame includes k sub-frames, and the short-term correlation parameters of the target speech frame include the line spectral frequencies and interpolation factors of the kth sub-frame of the target speech frame, where k is an integer greater than 1; the target speech frame includes m sub-frames, and the long-term correlation parameters of the target speech frame include the pitch delay and long-term prediction coefficients of each sub-frame of the target speech frame, where m is a positive integer.
8. The method according to claim 1, characterized in that, the network model further includes a third neural network, and the third neural network and the first neural network belong to parallel networks; the method further includes: Obtaining the energy parameters of the historical speech frames; Invoking the third neural network to perform prediction processing on the energy parameters of the historical speech frames to obtain the energy parameters of the target speech frame, and the energy parameters of the target speech frame belong to one of the parameters in the parameter set of the target speech frame; wherein, the target speech frame includes m sub-frames, and the energy parameters of the target speech frame include the gain values of each sub-frame of the target speech frame.
9. A speech processing method, characterized in that, including: Receiving a speech signal transmitted through the VoIP system; When the target speech frame in the speech signal is lost, reconstructing the target speech frame by using the method according to any one of claims 1 - 8; Outputting the speech signal based on the reconstructed target speech frame.
10. A speech processing device, characterized in that, including: A determination unit for determining the historical speech frames corresponding to the target speech frame to be processed; An acquisition unit, configured to acquire the frequency-domain features of the historical speech frame; A processing unit, configured to call a network model to perform prediction processing on the frequency-domain features of the historical speech frame to obtain a parameter set of the target speech frame; at least two parameters are included in the parameter set, the network model includes multiple neural networks, and the number of neural networks is determined according to the number of parameters in the parameter set; and configured to reconstruct the target speech frame according to the parameter set; Wherein, the network model includes a first neural network and at least two second neural networks, and the second neural networks belong to sub-networks of the first neural network; one of the second neural networks corresponds to one type of parameter in the parameter set; specifically, the processing unit is configured to: Call the first neural network to perform prediction processing on the frequency-domain features of the historical speech frame to obtain virtual frequency-domain features of the target speech frame; Input the virtual frequency-domain features of the target speech frame into the at least two second neural networks respectively for prediction processing to obtain at least two parameters in the parameter set of the target speech frame.
11. A speech processing device Characterized in that It includes: A receiving unit, configured to receive a voice signal transmitted through a VoIP system; A processing unit, configured to reconstruct the target speech frame by using the method according to any one of claims 1-8 when the target speech frame in the voice signal is lost; An output unit, configured to output the voice signal based on the reconstructed target speech frame.
12. A speech processing device Characterized in that The device includes: A processor, adapted to implement one or more instructions; and, A computer-readable storage medium, storing one or more instructions, where the one or more instructions are adapted to be loaded and executed by the processor to perform the speech processing method according to any one of claims 1-9.
13. A computer-readable storage medium Characterized in that The computer-readable storage medium stores one or more instructions, where the one or more instructions are adapted to be loaded and executed by a processor to perform the speech processing method according to any one of claims 1-9.
Citation Information
Patent Citations
Frame loss compensation method and frame loss compensation device for transform domain
CN103854649A
Transmission error concealment in an audio signal
US20040010407A1
Method of decoding an audio signal with correction of transmission errors
US6408267B1