Methods, devices, electronic equipment, and storage media for processing speech signals
By predicting and splicing speech data frames, and combining speech prediction models and Wiener filtering, the reverberation problem caused by inter-frame delay in headphone transparency mode was solved, achieving higher quality human voice signal output.
Patent Information
- Application Number
- CN202310621422.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-05-29
- Publication Date
- 2025-11-14
- Estimated Expiration
- 2043-05-29
AI Technical Summary
In headphone transparency mode, the inter-frame processing of the voice signal introduces a delay, resulting in a reverberation effect and reducing the quality of the human voice signal.
By acquiring current and historical speech data, predictions are made and spliced together to obtain future speech data frames, avoiding the time delay of inter-frame overlap processing. Speech prediction models such as LSTM or RNN models are used to predict future speech data, and Wiener filtering is combined to extract human voice signals.
It reduces reverberation, improves the quality of the voice signal, aligns passively transmitted sound with the playback sound, and provides a more natural human voice listening experience.
Smart Images

Figure CN116631419B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of signal processing technology, and in particular to a method, apparatus, electronic device and storage medium for processing speech signals. Background Technology
[0002] Headphones with a transparency mode (ambient sound technology) already exist. When users wear headphones and switch to transparency mode, they can perceive external sounds as if they weren't wearing headphones at all, and clearly hear human voices, enabling clear conversations. However, environments are usually noisy, and we want to hear more human voices and less noise during conversations, hence the development of voice enhancement features.
[0003] In related technologies, the voice signal needs to be processed in frames during the process of voice enhancement. There is a certain delay in the acquisition of the signal when processing the frames, which causes the sound entering the ear in the actual environment and the delayed sound to be superimposed in the ear, resulting in a reverberation effect and reducing the quality of the voice signal. Summary of the Invention
[0004] This application aims to at least partially address one of the technical problems in the related art.
[0005] To this end, this application proposes a method, apparatus, electronic device, and storage medium for processing speech signals, which avoids the time delay introduced by processing inter-frame overlapping data, reduces reverberation, and improves speech quality.
[0006] One embodiment of this application proposes a method for processing speech signals, including:
[0007] Obtain the first subframe audio data of the current environment and at least one historical subframe audio data preceding the first subframe audio data;
[0008] Based on the first subframe speech data and / or the at least one historical subframe speech data, a prediction is made to obtain the second subframe speech data after the first subframe speech data; the first subframe speech data and the second subframe speech data are concatenated to obtain the first frame speech data; and the target historical subframe speech data in the at least one historical subframe speech data is concatenated with the first subframe speech data to obtain the second frame speech data.
[0009] Human voice signal is extracted based on the first frame of speech data and the second frame of speech data to obtain the target human voice signal in the first sub-frame of speech data.
[0010] Another aspect of this application provides a speech signal processing apparatus, comprising:
[0011] The acquisition module is used to acquire the first subframe speech data of the current environment and at least one historical subframe speech data before the first subframe speech data;
[0012] The prediction module is used to predict based on the first subframe speech data and / or at least one historical subframe speech data to obtain the second subframe speech data after the first subframe speech data.
[0013] The splicing module is used to splice the first subframe speech data and the second subframe speech data to obtain the first frame speech data, and to splice the target historical subframe speech data in the at least one historical subframe speech data with the first subframe speech data to obtain the second frame speech data.
[0014] The determining module is used to extract human voice signals based on the first frame of speech data and the second frame of speech data to obtain the target human voice signal in the first sub-frame of speech data.
[0015] Another embodiment of this application provides an electronic device including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, it implements the method described in the foregoing aspect.
[0016] Another embodiment of this application proposes a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method described in the foregoing aspect.
[0017] Another embodiment of this application proposes a computer program product having a computer program stored thereon, which, when executed by a processor, implements the method described in the foregoing aspect.
[0018] The speech signal processing method, apparatus, electronic device, and storage medium proposed in this application acquire first subframe speech data of the current environment and at least one historical subframe speech data preceding the first subframe speech data. Based on the first subframe speech data and / or at least one historical subframe speech data, prediction is performed to obtain second subframe speech data following the first subframe speech data. The first and second subframe speech data are concatenated to obtain a first frame speech data. Furthermore, a target historical subframe speech data from at least one historical subframe speech data is concatenated with the first frame speech data to obtain a second frame speech data. Human voice signal extraction is performed based on the first and second frame speech data to obtain a target human voice signal in the first subframe speech data. By predicting the second subframe speech data following the first subframe speech data of the currently acquired environment, the target human voice signal in the current first subframe speech data can be determined without needing a set delay to wait for the acquisition of the second subframe speech data. This avoids the time delay introduced by processing inter-frame overlapping data, reduces reverberation, and improves speech quality.
[0019] Additional aspects and advantages of this application will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of this application. Attached Figure Description
[0020] The above and / or additional aspects and advantages of this application will become apparent and readily understood from the following description of the embodiments taken in conjunction with the accompanying drawings, wherein:
[0021] Figure 1 This is a schematic diagram of speech signal processing in related technologies;
[0022] Figure 2 A schematic flowchart illustrating a speech signal processing method provided in an embodiment of this application;
[0023] Figure 3 A flowchart illustrating another method for processing voice signals provided in an embodiment of this application;
[0024] Figure 4 This is a schematic diagram of a speech signal processing method provided in an embodiment of this application;
[0025] Figure 5 A schematic diagram of the structure of a speech signal processing device provided in an embodiment of this application;
[0026] Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0027] The embodiments of this application are described in detail below. Examples of these embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain this application, and should not be construed as limiting this application.
[0028] The following description, with reference to the accompanying drawings, outlines a method, apparatus, electronic device, and storage medium for processing voice signals according to embodiments of this application.
[0029] In related technologies, when headphones are in transparency mode, the headphone's feedforward microphone collects ambient speech (sound) signals. These signals are then processed into frames. To avoid sound distortion caused by blockiness, adjacent frames are typically overlapped and added together to obtain the processed speech signal to be played. As an example, Figure 1 This is a schematic diagram of speech signal processing in related technologies, using the example of adding speech frames with 50% overlap between them. Figure 1 As shown, the current half-frame signal captured by the microphone is combined with the adjacent first half-frame signal to form a single frame signal for processing a single frame of speech signal. Due to the 50% overlap, only a half-frame signal can be output after processing, and the signal content corresponds to the first half-frame. In other words, the output after processing is the content of the first half-frame signal. The content of the currently input half-frame signal needs to be superimposed with the captured next half-frame signal before it can be output to the speaker for playback. Therefore, a half-frame delay is introduced, which is the time required to capture the next half-frame signal. This causes the sound passively transmitted to the user's ear and the sound delayed and played back by the headphones through the speaker to overlap in the user's ear, resulting in a reverberation effect and reducing the natural listening experience. To solve this problem, this application proposes a speech signal processing method that predicts the second sub-frame speech data after the first sub-frame speech data of the currently captured environment. This eliminates the need for a set delay to wait for the acquisition of the second sub-frame speech data, and can determine the target human voice signal in the current first sub-frame speech data. This avoids the delay introduced by processing inter-frame overlapping data, reduces reverberation, and improves speech quality.
[0030] Figure 2 This is a schematic flowchart illustrating a method for processing voice signals provided in an embodiment of this application.
[0031] The execution subject of the voice signal processing method in this application embodiment is a voice signal processing device, which can be installed in an earphone.
[0032] like Figure 2 As shown, the method may include the following steps:
[0033] Step 201: Obtain the first subframe speech data of the current environment and at least one historical subframe speech data preceding the first subframe speech data.
[0034] The first subframe speech data is a portion of the speech data in a frame of speech data. In other words, the number of sampling points contained in the first subframe speech data is a first percentage of the number of sampling points in a frame of speech data. The first percentage is 25%, 50%, etc. As an example, the first percentage is 50%, which means that the first subframe speech data is half a frame of speech data. For example, if a frame of speech data contains 1024 sampling points, then the first subframe speech data contains 512 sampling points.
[0035] It should be noted that in related technologies, the audio signal needs to be processed by framing. In order to avoid sound distortion caused by block effect, there is usually an overlap between adjacent frames obtained by framing. For example, the overlap between adjacent frames is 50%.
[0036] As one implementation, the first subframe of speech data in the current environment can be acquired through the microphone of the headphones, for example, through the feedforward microphone of the headphones.
[0037] The historical speech data preceding the first subframe speech data, i.e., at least one historical subframe speech data, is historically stored speech data. Historical subframe speech data is a portion of the speech data within a single frame of speech data; that is, the sum of the frame length of the historical subframe speech data and the frame length of the first subframe speech data equals the frame length of the single frame of speech data, or the frame lengths of the historical subframe speech data and the first subframe speech data are equal. For example, there are two historical subframe speech data: historical subframe speech data A adjacent to the first subframe speech data, and historical subframe speech data B adjacent to historical subframe speech data A. A single frame of speech data includes 1024 sampling points, meaning the frame length of a single frame of speech data is considered to be 1024 sampling points. The first subframe speech data includes 512 sampling points, meaning its frame length is considered to be 512 sampling points. The adjacent historical subframe speech data A includes 512 sampling points, meaning its frame length is considered to be 512 sampling points. Historical subframe speech data B also includes 512 sampling points, meaning its frame length is considered to be 512 sampling points. Alternatively, the first subframe audio data includes 256 sampling points, meaning the frame length is considered to be 256 sampling points; the adjacent historical subframe audio data A includes 768 sampling points, meaning the frame length is considered to be 768 sampling points; and the historical subframe audio data B includes 256 sampling points, meaning the frame length is considered to be 256 sampling points. Other examples are not listed in this embodiment.
[0038] As an example, if the first subframe audio data is half-frame audio data, then the historical audio data contains N half-frames. For example, if N is 3, then the historical audio data and the first subframe audio data can form 2 frames of audio data.
[0039] Step 202: Based on the first subframe speech data and / or at least one historical subframe speech data, predict to obtain the second subframe speech data after the first subframe speech data.
[0040] In the first implementation of this application, prediction is performed based on the first subframe speech data to obtain the second subframe speech data after the first subframe speech data.
[0041] In the second implementation of this application, prediction is performed based on at least one historical subframe speech data to obtain the second subframe speech data after the first subframe speech data.
[0042] In the third implementation of this application, prediction is performed based on the first subframe speech data and at least one historical subframe speech data to obtain the second subframe speech data after the first subframe speech data.
[0043] In this context, the frame length of the second subframe speech data and the frame length of the first subframe speech data are equal to the frame length of one frame of speech data. That is to say, the number of sampling points contained in the second speech data is the second proportion of the number of sampling points in one frame of speech data, and the sum of the first proportion and the second proportion is 1.
[0044] As an example, a frame of speech data includes 1024 sampling points, meaning the frame length of a frame of speech data is considered to be 1024 sampling points. The first subframe of speech data includes 512 sampling points, which is half a frame, meaning the frame length is considered to be 512 sampling points. The predicted second subframe of speech data also includes 512 sampling points, which is half a frame, meaning the frame length is considered to be 512 sampling points. If the first subframe of speech data includes 256 sampling points, which is a quarter frame, meaning the frame length is considered to be 256 sampling points, the predicted second subframe of speech data includes 768 sampling points, which is half a frame, meaning the frame length is considered to be 768 sampling points.
[0045] Step 203: Concatenate the first subframe speech data and the second subframe speech data to obtain the first frame speech data, and concatenate the target historical subframe speech data from at least one historical subframe speech data with the first subframe speech data to obtain the second frame speech data.
[0046] In this embodiment, the frame length of the second subframe audio data and the frame length of the first subframe audio data are equal to the frame length of one frame of audio data. That is, the first subframe audio data and the second subframe audio data are concatenated to obtain the first frame of audio data. The target historical subframe audio data refers to historical subframe audio data whose frame length plus the frame length of the first subframe audio data equals the frame length of one frame of audio data. As one implementation, the target historical subframe audio data is the preceding historical subframe audio data adjacent to the first subframe audio data. Therefore, the target historical subframe audio data and the first audio data are concatenated to obtain the second frame of audio data.
[0047] Step 204: Extract human voice signal based on the first frame of speech data and the second frame of speech data to obtain the target human voice signal in the first subframe of speech data.
[0048] In this embodiment of the application, both the first frame of speech data and the second frame of speech data contain the first subframe of speech data. Therefore, the first subframe of speech data is regarded as the overlapping part between two adjacent frames. A human voice enhancement algorithm is applied to the first frame of speech data and the second frame of speech data respectively to obtain the human voice signal in the first frame of speech data and the human voice signal in the second frame of speech data. Based on the human voice signal in the first frame of speech data and the human voice signal in the second frame of speech data, the target human voice signal in the enhanced first subframe of speech data is obtained.
[0049] In the speech signal processing method of this application embodiment, the following steps are taken: first subframe speech data of the current environment and at least one historical subframe speech data before the first subframe speech data are acquired; prediction is performed based on the first subframe speech data and / or at least one historical subframe speech data to obtain second subframe speech data after the first subframe speech data; the first subframe speech data and the second subframe speech data are concatenated to obtain the first frame speech data; and target historical subframe speech data from at least one historical subframe speech data is concatenated with the first frame speech data to obtain the second frame speech data; human voice signal is extracted based on the first frame speech data and the second subframe speech data to obtain the target human voice signal in the first subframe speech data. By predicting the second subframe speech data after the first subframe speech data of the currently acquired environment, the target human voice signal in the current first subframe speech data can be determined without delaying the acquisition of the second subframe speech data. This avoids the time delay introduced by processing inter-frame overlapping data, reduces reverberation, and improves speech quality.
[0050] Based on the above embodiments, Figure 3 A flowchart illustrating another speech signal processing method provided in an embodiment of this application is shown below. Figure 3 As shown, the method includes the following steps:
[0051] Step 301: Obtain the first subframe speech data of the current environment and at least one historical subframe speech data before the first subframe speech data.
[0052] For details, please refer to the explanations in the foregoing embodiments. The principle is the same, and will not be repeated here.
[0053] Step 302: Based on the first subframe speech data and / or at least one historical subframe speech data, predict to obtain the second subframe speech data after the first subframe speech data.
[0054] In this embodiment, at least one frame of audio data is generated based on the first subframe audio data and at least one historical subframe audio data. One implementation involves acquiring target historical subframe audio data adjacent to the first subframe audio data, as well as historical subframe audio data preceding the target historical subframe audio data. The first subframe audio data and the target historical subframe audio data are concatenated to obtain one frame of audio data. Furthermore, multiple historical subframe audio data preceding the target historical subframe are concatenated into adjacent subframes to obtain at least one frame of audio data. The sum of the frame length of the target historical subframe audio data and the frame length of the first subframe audio data equals the frame length of the first frame of audio data. Then, the second subframe audio signal is predicted based on the at least one frame of audio data. This achieves the prediction of future second subframe audio data based on existing multi-frame audio data, enabling the early acquisition of the second subframe audio data without requiring a set delay to acquire the collected second subframe audio data. This avoids delays, promptly generates the audio signal to be played, aligns the passive sound with the playback sound, and optimizes latency.
[0055] One implementation method for predicting the second subframe speech signal involves inputting at least one frame of speech data into a trained speech prediction model to obtain the predicted second subframe speech signal. The speech prediction model can be, for example, a Long Short-Term Memory Network (LSTM) model or a Recurrent Neural Network (RNN) model. This allows the model to predict future second subframe speech data using existing multi-frame data, improving prediction efficiency and accuracy. The speech prediction model is trained using training samples, which consist of one subframe of speech data and labeled next subframe speech data. The model is trained based on the difference between the predicted next subframe speech data and the labeled next subframe speech data.
[0056] As an example, let's consider each subframe of audio data as a half-frame, meaning the overlap between adjacent frames is considered a half-frame. For example... Figure 4As shown, the first subframe speech data is the currently acquired half-frame speech data, referred to as the current half-frame. The historical speech data contains N historical half-frame speech data, referred to as historical half-frame 1, historical half-frame 2, ... and historical half-frame N, respectively. The future half-frame is predicted by using the first subframe speech data and multiple historical half-frame speech data, that is, the second subframe speech data after the first subframe speech data. The second subframe speech data realizes the prediction of the future half-frame based on multiple frames, which improves the accuracy of prediction.
[0057] Step 303: Concatenate the first subframe speech data and the second subframe speech data to obtain the first frame speech data, and concatenate the target historical subframe speech data from at least one historical subframe speech data with the first subframe speech data to obtain the second frame speech data.
[0058] For details, please refer to the explanations in the foregoing embodiments. The principle is the same, and will not be repeated here.
[0059] like Figure 4 As shown, the speech data of the current half-frame and the future half-frame are concatenated to obtain the first frame of speech data, and the speech data of the historical half-frame 1 and the current half-frame are concatenated to obtain the second frame of speech data. Then, based on the first frame of speech data and the second frame of speech data, signal processing is performed to obtain the speech data of the current half-frame, that is, the target human voice signal corresponding to the current half-frame.
[0060] Step 304: Perform Fourier transform on the first frame of speech data and the second frame of speech data respectively to convert them to the frequency domain to obtain the first frequency domain speech frame and the second frequency domain speech frame.
[0061] In this embodiment, the first and second frames of speech data in the time domain are processed by a window function and then transformed to the frequency domain by Fourier transform to obtain the first and second frequency domain speech frames for signal processing in the frequency domain. The window function can be, for example, a rectangular window, a Hanning window, a flat-top window, or an exponential window.
[0062] Step 305: Perform Wiener filtering on the first frequency domain speech frame to obtain the first frequency domain voice signal, and perform Wiener filtering on the second frequency domain speech frame to obtain the second frequency domain voice signal.
[0063] In one implementation of this application, the filter coefficients of the Wiener filter corresponding to the first frequency domain speech frame are obtained. The first frequency domain speech frame is then filtered using these filter coefficients, i.e., signals other than human voice signals are filtered out to obtain the first frequency domain human voice signal in the first frequency domain speech frame. Similarly, the filter coefficients of the Wiener filter corresponding to the second frequency domain speech frame are obtained. The second frequency domain speech frame is then filtered using these filter coefficients, i.e., signals other than human voice signals are filtered out to obtain the second frequency domain human voice signal in the second frequency domain speech frame.
[0064] The method for determining the filter coefficients of the Wiener filter corresponding to each frequency domain speech frame (the first frequency domain speech frame or the second frequency domain speech frame in this application) is illustrated using the first frequency domain speech frame as an example. As one implementation, the power spectrum corresponding to the first frequency domain speech frame can be determined, and noise estimation can be performed on the first frequency domain speech frame to obtain the power spectrum corresponding to the noise signal in the first frequency domain speech frame. Based on the power spectrum corresponding to the first frequency domain speech frame and the power spectrum corresponding to the noise signal in the first frequency domain speech frame, the posterior signal-to-noise ratio (SNR) corresponding to the first frequency domain speech frame is determined. Based on the posterior SNR and the power spectrum corresponding to the noise signal in the previous frequency domain speech frame of the first frequency domain speech frame, the prior SNR estimate corresponding to the first frequency domain speech frame is determined. Based on the prior SNR estimate, the filter coefficients of the Wiener filter corresponding to the first frequency domain speech frame are generated.
[0065] Step 306: Perform inverse Fourier transform on the first frequency domain voice signal and the second frequency domain voice signal respectively to convert them to the time domain, to obtain the first time domain voice signal and the second time domain voice signal.
[0066] Step 307: Determine the voice signal belonging to the first subframe speech data from the first time-domain voice signal and the second time-domain voice signal to obtain the target voice signal.
[0067] One implementation method involves converting the frequency domain signal to the time domain. Since the first and second time domain voice signals are time sequences in the time domain, i.e., data arranged according to the order of sampling time, the first time domain voice signal contains each sampling point carrying time information. Therefore, the first voice signal belonging to the first subframe speech data in the first time domain voice signal and the second voice signal belonging to the first subframe speech data in the second time domain voice signal can be determined based on the time information of each sampling point. The first and second voice signals are then superimposed to obtain the target voice signal. Another implementation method involves superimposing the first and second voice signals through a synthesis window function to obtain the target voice signal in the first subframe speech data.
[0068] For example, if the historical subframe voice data is "1" and the current first subframe voice data is "2", the related technology outputs "1", while this application outputs "2" voice data. In other words, the voice data output in this application corresponds to the actual voice data collected, there is no delay, the confusion of different data is avoided, and the quality of voice data output is improved.
[0069] Step 308: Play the target human voice signal using the set speaker.
[0070] In this embodiment, the target human voice signal in the currently acquired first subframe speech signal is played through a speaker. The passively transmitted sound and the sound played by the speaker overlap without delay, so that the passive sound and the played sound are aligned, the delay is optimized, and a natural and clear human voice signal can be heard, providing a more natural human voice listening experience.
[0071] Furthermore, in one implementation of this application embodiment, the first frame of speech data and the second frame of speech data are filtered according to a preset transparency filter to obtain first environmental speech data corresponding to the first subframe of speech data in the first frame of speech data, and second environmental speech data corresponding to the second subframe of speech data in the second frame of speech data. The first environmental speech data and the second environmental speech data are superimposed. Optionally, they can be synthesized using a synthesis window function to obtain the target environmental speech data in the first subframe of speech data. Then, a set speaker is used to synchronously play the target human voice signal and the target environmental speech data, so that the human voice part in the transparent environmental speech is enhanced. At the same time, other sound signals in the environment other than the human voice signal can be obtained, but the other sound signals are not enhanced. Only the human voice signal is enhanced, making the human voice signal clearer.
[0072] In the voice data processing method of this application embodiment, existing multi-frame data can be used to predict the future half-frame data. The future half-frame data and the current half-frame data are merged for frame processing, so that the passive sound and the playback sound are aligned and the latency is optimized.
[0073] To implement the above embodiments, this application also proposes a speech signal processing device.
[0074] Figure 5 This is a schematic diagram of the structure of a speech signal processing device provided in an embodiment of this application.
[0075] like Figure 5 As shown, the device may include:
[0076] The acquisition module 51 is used to acquire the first subframe speech data of the current environment and at least one historical subframe speech data before the first subframe speech data;
[0077] Prediction module 52 is used to predict based on the first subframe speech data and / or the at least one historical subframe speech data to obtain the second subframe speech data after the first subframe speech data.
[0078] The splicing module 53 is used to splice the first subframe speech data and the second subframe speech data to obtain the first frame speech data, and to splice the target historical subframe speech data in the at least one historical subframe speech data with the first subframe speech data to obtain the second frame speech data.
[0079] The determining module 54 is used to extract human voice signals based on the first frame of speech data and the second frame of speech data to obtain the target human voice signal in the first sub-frame of speech data.
[0080] Furthermore, in one implementation of this application embodiment, the determining module 54 is specifically used for:
[0081] The first frame of speech data and the second frame of speech data are respectively subjected to Fourier transform to be converted to the frequency domain to obtain the first frequency domain speech frame and the second frequency domain speech frame.
[0082] Wiener filtering is applied to the first frequency domain speech frame to obtain a first frequency domain human voice signal, and Wiener filtering is applied to the second frequency domain speech frame to obtain a second frequency domain human voice signal.
[0083] The first frequency domain human voice signal and the second frequency domain human voice signal are respectively converted to the time domain by inverse Fourier transform to obtain the first time domain human voice signal and the second time domain human voice signal;
[0084] The target voice signal is obtained by determining the voice signal belonging to the first subframe speech data from the first time-domain voice signal and the second time-domain voice signal.
[0085] In one implementation of this application embodiment, the determining module 54 is specifically used for:
[0086] Obtain the filter coefficients of the Wiener filter corresponding to the first frequency domain speech frame;
[0087] The first frequency domain speech frame is filtered using the filtering coefficients to obtain the first frequency domain human voice signal in the first frequency domain speech frame.
[0088] In one implementation of this application embodiment, the determining module 54 is specifically used for:
[0089] Identify the first voice signal in the first time-domain voice signal that belongs to the first subframe speech data, and the second voice signal in the second time-domain voice signal that belongs to the first subframe speech data;
[0090] The first voice signal and the second voice signal are superimposed to obtain the target voice signal.
[0091] In one implementation of this application, the prediction module 54 is specifically used for:
[0092] Based on the first subframe speech data and the at least one historical subframe speech data, at least one frame of speech data is generated; based on the at least one frame of speech data, the second subframe speech signal is predicted.
[0093] In one implementation of this application, the prediction module 54 is specifically used for:
[0094] The at least one frame of speech data is input into the trained speech prediction model to obtain the second subframe speech signal predicted by the speech prediction model.
[0095] In one implementation of this application, the apparatus further includes:
[0096] The target human voice signal is played through a designated speaker.
[0097] It should be noted that the foregoing explanation of the method embodiments also applies to the apparatus of this embodiment, and will not be repeated here.
[0098] In the speech signal processing apparatus of this application embodiment, the first subframe speech data of the current environment and at least one historical subframe speech data preceding the first subframe speech data are acquired. Prediction is performed based on the first subframe speech data and / or at least one historical subframe speech data to obtain a second subframe speech data following the first subframe speech data. The first subframe speech data and the second subframe speech data are concatenated to obtain a first frame speech data. Furthermore, target historical subframe speech data from at least one historical subframe speech data is concatenated with the first subframe speech data to obtain a second frame speech data. Human voice signal extraction is performed based on the first and second frame speech data to obtain the target human voice signal in the first subframe speech data. By predicting the second subframe speech data following the first subframe speech data of the currently acquired environment, the target human voice signal in the current first subframe speech data can be determined without needing to delay for a set duration to wait for the acquisition of the second subframe speech data. This avoids the time delay introduced by processing inter-frame overlapping data, reduces reverberation, and improves speech quality.
[0099] To implement the above embodiments, this application also proposes an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, it implements the method described in the foregoing method embodiments.
[0100] To implement the above embodiments, this application also proposes a non-transitory computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the method described in the foregoing method embodiments.
[0101] To implement the above embodiments, this application also proposes a computer program product having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the method described in the foregoing method embodiments.
[0102] Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. For example, the electronic device 800 may be a mobile phone, computer, digital broadcasting terminal, messaging device, game console, tablet device, medical device, fitness equipment, personal digital assistant, etc.
[0103] Reference Figure 6 The electronic device 800 may include one or more of the following components: processing component 802, memory 804, power component 806, multimedia component 808, audio component 810, input / output (I / O) interface 812, sensor component 814, and communication component 816.
[0104] Processing component 802 typically controls the overall operation of electronic device 800, such as operations associated with display, telephone calls, data communication, camera operation, and recording operations. Processing component 802 may include one or more processors 820 to execute instructions to complete all or part of the steps of the methods described above. Furthermore, processing component 802 may include one or more modules to facilitate interaction between processing component 802 and other components. For example, processing component 802 may include a multimedia module to facilitate interaction between multimedia component 808 and processing component 802.
[0105] Memory 804 is configured to store various types of data to support the operation of electronic device 800. Examples of this data include instructions for any application or method operating on electronic device 800, contact data, phonebook data, messages, pictures, videos, etc. Memory 804 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.
[0106] Power component 806 provides power to various components of electronic device 800. Power component 806 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to electronic device 800.
[0107] Multimedia component 808 includes a screen that provides an output interface between the electronic device 800 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touchscreen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors may sense not only the boundaries of the touch or swipe action but also the duration and pressure associated with the touch or swipe operation. In some embodiments, multimedia component 808 includes a front-facing camera and / or a rear-facing camera. When the electronic device 800 is in an operating mode, such as a shooting mode or a video mode, the front-facing camera and / or the rear-facing camera may receive external multimedia data. Each front-facing camera and rear-facing camera may be a fixed optical lens system or have focal length and optical zoom capabilities.
[0108] Audio component 810 is configured to output and / or input audio signals. For example, audio component 810 includes a microphone (MIC) configured to receive external audio signals when electronic device 800 is in an operating mode, such as call mode, recording mode, and voice recognition mode. The received audio signals may be further stored in memory 804 or transmitted via communication component 816. In some embodiments, audio component 810 also includes a speaker for outputting audio signals.
[0109] I / O interface 812 provides an interface between processing component 802 and peripheral interface modules, such as keyboards, click wheels, buttons, etc. These buttons may include, but are not limited to, home buttons, volume buttons, power buttons, and lock buttons.
[0110] Sensor assembly 814 includes one or more sensors for providing state assessments of various aspects of electronic device 800. For example, sensor assembly 814 can detect the on / off state of electronic device 800, the relative positioning of components such as the display and keypad of electronic device 800, changes in position of electronic device 800 or a component of electronic device 800, the presence or absence of user contact with electronic device 800, orientation or acceleration / deceleration of electronic device 800, and temperature changes of electronic device 800. Sensor assembly 814 may include a proximity sensor configured to detect the presence of nearby objects without any physical contact. Sensor assembly 814 may also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, sensor assembly 814 may also include an accelerometer, gyroscope, magnetometer, pressure sensor, or temperature sensor.
[0111] Communication component 816 is configured to facilitate wired or wireless communication between electronic device 800 and other devices. Electronic device 800 can access wireless networks based on communication standards, such as WiFi, 4G, or 5G, or combinations thereof. In one exemplary embodiment, communication component 816 receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In one exemplary embodiment, communication component 816 also includes a near-field communication (NFC) module to facilitate short-range communication. For example, the NFC module may be implemented based on radio frequency identification (RFID) technology, Infrared Data Association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.
[0112] In an exemplary embodiment, the electronic device 800 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the methods described above.
[0113] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 804 including instructions, which can be executed by a processor 820 of an electronic device 800 to perform the above-described method. For example, the non-transitory computer-readable storage medium may be a ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc.
[0114] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., refer to specific features, structures, materials, or characteristics described in connection with that embodiment or example, which are included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.
[0115] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this application, "multiple" means at least two, such as two, three, etc., unless otherwise explicitly specified.
[0116] Any process or method description in the flowchart or otherwise herein can be understood as representing a module, segment, or portion of code comprising one or more executable instructions for implementing custom logic functions or processes, and the scope of the preferred embodiments of this application includes additional implementations in which functions may be performed not in the order shown or discussed, including substantially simultaneously or in reverse order depending on the functions involved, as should be understood by those skilled in the art to which embodiments of this application pertain.
[0117] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a processor-included system, or other system that can fetch and execute instructions from, an instruction execution system, apparatus, or device). For the purposes of this specification, "computer-readable medium" can be any means that can contain, store, communicate, propagate, or transmit programs for use by, or in conjunction with, an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of computer-readable media include: an electrical connection having one or more wires (electronic device), a portable computer disk drive (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic devices, and portable optical disc read-only memory (CDROM). Alternatively, the computer-readable medium may be paper or other suitable media on which the program can be printed, since the program can be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, interpreting, or otherwise processing as necessary, and then stored in a computer memory.
[0118] It should be understood that various parts of this application can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented using software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.
[0119] Those skilled in the art will understand that all or part of the steps of the methods in the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, the program includes one or a combination of the steps of the method embodiments.
[0120] Furthermore, the functional units in the various embodiments of this application can be integrated into a processing module, or each unit can exist physically separately, or two or more units can be integrated into a module. The integrated module can be implemented in hardware or as a software functional module. If the integrated module is implemented as a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium.
[0121] The storage medium mentioned above can be a read-only memory, a disk, or an optical disk, etc. Although embodiments of this application have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting this application. Those skilled in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of this application.
Claims
1. A method for processing speech signals, characterized in that, include: Acquire the first subframe audio data of the current environment and at least one historical subframe audio data preceding the first subframe audio data; Based on the first subframe speech data and / or the at least one historical subframe speech data, a prediction is made to obtain the second subframe speech data after the first subframe speech data. The first subframe speech data and the second subframe speech data are concatenated to obtain the first frame speech data, and the target historical subframe speech data in the at least one historical subframe speech data is concatenated with the first subframe speech data to obtain the second frame speech data. Human voice signal is extracted based on the first frame of speech data and the second frame of speech data to obtain the target human voice signal in the first sub-frame of speech data.
2. The method as described in claim 1, characterized in that, The step of extracting human voice signals based on the first frame of speech data and the second frame of speech data to obtain the target human voice signal in the first sub-frame of speech data includes: The first frame of speech data and the second frame of speech data are respectively subjected to Fourier transform to be converted to the frequency domain to obtain the first frequency domain speech frame and the second frequency domain speech frame. Wiener filtering is applied to the first frequency domain speech frame to obtain a first frequency domain human voice signal, and Wiener filtering is applied to the second frequency domain speech frame to obtain a second frequency domain human voice signal. The first frequency domain human voice signal and the second frequency domain human voice signal are respectively converted to the time domain by inverse Fourier transform to obtain the first time domain human voice signal and the second time domain human voice signal; The target voice signal is obtained by determining the voice signal belonging to the first subframe speech data from the first time-domain voice signal and the second time-domain voice signal.
3. The method as described in claim 2, characterized in that, The step of performing Wiener filtering on the first frequency domain speech frame to obtain the first frequency domain human voice signal includes: Obtain the filter coefficients of the Wiener filter corresponding to the first frequency domain speech frame; The first frequency domain speech frame is filtered using the filtering coefficients to obtain the first frequency domain human voice signal in the first frequency domain speech frame.
4. The method as described in claim 2, characterized in that, The step of determining the voice signal belonging to the first subframe speech data from the first time-domain voice signal and the second time-domain voice signal to obtain the target voice signal includes: Identify the first voice signal in the first time-domain voice signal that belongs to the first subframe speech data, and the second voice signal in the second time-domain voice signal that belongs to the first subframe speech data; The first voice signal and the second voice signal are superimposed to obtain the target voice signal.
5. The method as described in claim 1, characterized in that, The step of predicting the second subframe speech data based on the first subframe speech data and / or at least one historical subframe speech data to obtain the second subframe speech data includes: Based on the first subframe speech data and the at least one historical subframe speech data, at least one frame of speech data is generated; The second subframe speech signal is predicted based on the at least one frame of speech data.
6. The method as described in claim 5, characterized in that, The step of predicting the second sub-frame speech signal based on the at least one frame of speech data includes: The at least one frame of speech data is input into the trained speech prediction model to obtain the second subframe speech signal predicted by the speech prediction model.
7. The method according to any one of claims 1-6, characterized in that, After extracting the human voice signal based on the first frame of speech data and the second frame of speech data to obtain the target human voice signal in the first sub-frame of speech data, the method further includes: The target human voice signal is played through a designated speaker.
8. A speech signal processing apparatus, characterized in that, include: The acquisition module is used to acquire the first subframe speech data of the current environment and at least one historical subframe speech data before the first subframe speech data; The prediction module is used to predict based on the first subframe speech data and / or at least one historical subframe speech data to obtain the second subframe speech data after the first subframe speech data. The splicing module is used to splice the first subframe speech data and the second subframe speech data to obtain the first frame speech data, and to splice the target historical subframe speech data in the at least one historical subframe speech data with the first subframe speech data to obtain the second frame speech data. The determining module is used to extract human voice signals based on the first frame of speech data and the second frame of speech data to obtain the target human voice signal in the first sub-frame of speech data.
9. An electronic device, characterized in that, It includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, it implements the method as described in any one of claims 1-7.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1-7.
Citation Information
Patent Citations
Excitation signal synthesis during frame erasure or packet loss
CA2142393A1
Frontloading product inventory
CA2799353A1