Speech signal recognition methods, devices, electronic equipment and storage media

By using a multi-channel speech signal recognition method, a trained acoustic model is used to process synchronously acquired multi-channel speech signals, which solves the problem of poor recognition quality in single-channel speech and achieves higher quality speech content recognition.

CN114898736BActive Publication Date: 2025-11-14BEIJING XIAOMI MOBILE SOFTWARE CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210334101.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-30
Publication Date
2025-11-14
Estimated Expiration
2042-03-30

AI Technical Summary

Technical Problem

In existing technologies, recognition methods based on single-channel speech signals result in poor speech content quality.

Method used

A multi-channel speech signal recognition method is adopted, which acquires raw speech signals from multiple channels simultaneously, processes them using a trained first acoustic model, obtains corresponding phoneme sequences, and identifies the phoneme sequences to obtain speech content.

Benefits of technology

It achieves global information recognition based on multiple channel speech signals, reduces signal distortion, and improves the quality and purity of speech content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114898736B_ABST
    Figure CN114898736B_ABST
Patent Text Reader

Abstract

This application proposes a speech signal recognition method, apparatus, electronic device, and storage medium. The method includes: acquiring first speech signals from multiple channels, wherein the first speech signals from each channel are raw speech signals synchronously acquired within a set time period; inputting the first speech signals from multiple channels into a trained first acoustic model to obtain a corresponding first phoneme sequence; and recognizing the first phoneme sequence to obtain speech content. This method realizes recognition based on global information of the first speech signals from multiple channels to obtain speech content, achieving low signal distortion and high purity, and improving the quality of speech content.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field, and more particularly to a speech signal recognition method, apparatus, electronic device, and storage medium. Background Technology

[0002] In real-world scenarios, the acquired speech signals can be multi-channel, such as speech signals from speaker channels or multiple microphone channels in an array microphone. However, in related technologies, speech recognition is based on processing single-channel speech signals, resulting in poor quality of the recognized speech content. Summary of the Invention

[0003] This application proposes a speech signal recognition method, apparatus, electronic device, and storage medium to improve the effect of speech content recognition.

[0004] One embodiment of this application proposes a speech signal recognition method, including:

[0005] Acquire the first speech signals from multiple channels; wherein, the first speech signal of each channel is the original speech signal synchronously acquired within a set time period;

[0006] The first speech signal from the multiple channels is input into the first acoustic model obtained through training to obtain the corresponding first phoneme sequence;

[0007] The speech content is obtained by recognizing the first phoneme sequence.

[0008] Another aspect of this application provides a speech signal recognition device, comprising:

[0009] The acquisition module is used to acquire the first voice signals from multiple channels; wherein, the first voice signal of each channel is the original voice signal synchronously acquired within a set time period;

[0010] The processing module is used to input the first speech signals from the multiple channels into the first acoustic model obtained through training, and obtain the corresponding first phoneme sequence;

[0011] The recognition module is used to recognize the first phoneme sequence to obtain the speech content.

[0012] Another embodiment of this application provides an electronic device including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, it implements the method described in the foregoing aspect.

[0013] Another embodiment of this application proposes a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method described in the foregoing aspect.

[0014] Another embodiment of this application proposes a computer program product having a computer program stored thereon, which, when executed by a processor, implements the method described in the foregoing aspect.

[0015] The speech signal recognition method, apparatus, electronic device, and storage medium proposed in this application acquire first speech signals from multiple channels. The first speech signals of each channel are raw speech signals synchronously acquired within a set time period. The first speech signals from multiple channels are input into a first acoustic model trained to obtain a corresponding first phoneme sequence. The speech content is obtained by recognizing the first phoneme sequence. This realizes the recognition of global information based on the first speech signals from multiple channels to obtain the speech content. It achieves low signal distortion and high purity, thereby improving the quality of the speech content.

[0016] Additional aspects and advantages of this application will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of this application. Attached Figure Description

[0017] The above and / or additional aspects and advantages of this application will become apparent and readily understood from the following description of the embodiments taken in conjunction with the accompanying drawings, wherein:

[0018] Figure 1 A schematic flowchart illustrating a speech signal recognition method provided in an embodiment of this application;

[0019] Figure 2 A flowchart illustrating another speech signal recognition method provided in an embodiment of this application;

[0020] Figure 3 A schematic diagram illustrating a speech content recognition method provided in an embodiment of this application;

[0021] Figure 4 A flowchart illustrating another speech signal recognition method provided in an embodiment of this application;

[0022] Figure 5 This is a schematic diagram of the structure of the first acoustic model provided in the embodiments of this application;

[0023] Figure 6 A flowchart illustrating another speech signal recognition method provided in an embodiment of this application;

[0024] Figure 7 A flowchart illustrating another speech signal recognition method provided in an embodiment of this application;

[0025] Figure 8 This is a schematic diagram of a speech signal enhancement structure provided in an embodiment of this application;

[0026] Figure 9 A flowchart illustrating another speech signal recognition method provided in an embodiment of this application;

[0027] Figure 10 This is a schematic diagram of the structure of a voice signal recognition device provided in an embodiment of this application;

[0028] Figure 11 This is a block diagram of an electronic device provided in an embodiment of this application. Detailed Implementation

[0029] The embodiments of this application are described in detail below. Examples of these embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain this application, and should not be construed as limiting this application.

[0030] The speech signal recognition method, apparatus, electronic device, and storage medium of this application are described below with reference to the accompanying drawings.

[0031] Figure 1 This is a flowchart illustrating a speech signal recognition method provided in an embodiment of this application.

[0032] The execution subject of the voice signal recognition method in this application embodiment is a voice signal recognition device, which can be installed in an electronic device or is an electronic device. The electronic device can be a smart speaker, smart TV, smart set-top box, smartphone, wearable device, etc. The specific form of the electronic device is not limited in this embodiment.

[0033] like Figure 1 As shown, the method may include the following steps:

[0034] Step 101: Acquire the first voice signals from multiple channels, wherein the first voice signal of each channel is the original voice signal synchronously acquired within a set time period.

[0035] In this embodiment, the first voice signal of multiple channels can be the voice signal of the channels of a microphone array and a speaker installed in an electronic device. In a home or vehicle environment, a microphone array consisting of two or more microphones receives the sound emitted by the sound-producing device itself, such as music played by a speaker, as well as the user's voice and environmental noise and echo. Therefore, the first voice signal of multiple channels includes the voice signals of multiple microphone channels, the original signal of the voice signal played by the sound-producing device, the reverberation signal, and the environmental noise signal, wherein the microphone array includes multiple microphones, and each microphone corresponds to the first voice signal of one channel.

[0036] The first speech signal of each channel is the raw speech signal collected within a set time period. The raw speech signal can be any raw speech signal collected within a set time period based on the interval of sound source emission. In other words, the first speech signal of each channel has not undergone frame segmentation processing of the front-end signal. That is, when recognizing the first speech signal of each channel, it is not processed on the partial signals of each frame of speech signal separately, but based on the complete speech signal. The signal is not fragmented, all information of the signal is preserved, the signal distortion is low, and the effect of subsequent speech recognition is improved.

[0037] Step 102: Input the first speech signals from multiple channels into the first acoustic model obtained through training to obtain the corresponding first phoneme sequence.

[0038] The first acoustic model is an acoustic model based on the ASR architecture, while the non-end-to-end multi-channel ASR acoustic model can be implemented based on Kaldi's chain-tdnn.

[0039] In this embodiment, the first acoustic model is trained based on multi-channel speech signals. The trained first acoustic model has learned the relationship between the first speech signals of multiple input channels and the corresponding first phoneme sequences. Since the first acoustic model is also trained based on multi-channel speech signals, it means that the first acoustic model also learns the global semantic information of the speech signals of multiple channels during the training process, and the accuracy of the phoneme sequences corresponding to the speech signals of multiple channels is also high.

[0040] It should be noted that the duration of the first speech signal in each channel is the same. When it is input into the first acoustic model for recognition, the resulting first phoneme sequence contains the phonemes corresponding to each frame of the first speech signal in each channel. This achieves forced alignment of each frame of the first speech signal in multiple channels with the corresponding phonemes. Therefore, the first phoneme sequence is a combination of phonemes at the frame level. For example, if there are three channels, namely channel 1, channel 2 and channel 3, and each channel corresponds to three frames, then the first frame in channel 1, channel 2 and channel 3 all correspond to the phoneme w, the second frame in channel 1, channel 2 and channel 3 all correspond to the phoneme w, and the third frame in channel 1, channel 2 and channel 3 all correspond to the phoneme o (tone 3). Thus, the resulting phoneme sequence is w wo.

[0041] Step 103: Recognize the first phoneme sequence to obtain the speech content.

[0042] Therefore, by recognizing the first phoneme sequence, the corresponding speech content can be obtained. For example, the phoneme sequence "wwo" in step 102 can be recognized as the pronunciation of the word "I," meaning the speech content is "I." As one implementation method, the first phoneme sequence can be input into a trained language model to recognize the corresponding speech content, such as text data.

[0043] In the speech signal recognition method of this application embodiment, first speech signals from multiple channels are acquired. The first speech signals from each channel are original speech signals synchronously collected within a set time period. The first speech signals from multiple channels are input into a first acoustic model trained to obtain a corresponding first phoneme sequence. The speech content is obtained by recognizing the first phoneme sequence. This method realizes the recognition of global information based on the first speech signals from multiple channels to obtain the speech content, achieving low signal distortion and high purity, and improving the quality of the speech content.

[0044] Based on the above embodiments, Figure 2 A flowchart illustrating another speech signal recognition method provided in this application embodiment is shown below. Figure 2 As shown, the method includes the following steps:

[0045] Step 201: Acquire the first voice signals from multiple channels, wherein the first voice signal of each channel is the original voice signal synchronously acquired within a set time period.

[0046] Step 202: Input the first speech signals from multiple channels into the first acoustic model obtained through training to obtain the corresponding first phoneme sequence.

[0047] Steps 201 and 202 can be explained in the foregoing embodiments, as the principle is the same, and will not be repeated in this embodiment.

[0048] Step 203: Merge multiple consecutive identical phonemes in the first phoneme sequence to obtain the second phoneme sequence.

[0049] In this embodiment of the application, the first phoneme sequence is a frame-level phoneme sequence. The first phoneme sequence contains each phoneme of the first speech signal of multiple channels that are forcibly aligned to each frame. Since multiple frames may correspond to the same phoneme, there are multiple instances of the same phoneme appearing consecutively in the first phoneme sequence. Therefore, in order to improve the efficiency and accuracy of recognition, multiple consecutive instances of the same phoneme in the first phoneme sequence can be merged to obtain the second phoneme sequence.

[0050] As one implementation method, at least one phoneme group is determined based on the multiple phonemes arranged sequentially in the first phoneme sequence. Each phoneme group contains multiple adjacent identical phonemes. The identical phonemes in each phoneme group are merged to obtain the second phoneme sequence. For example, the first phoneme sequence is "wwww oo zzz ouououou l eee", and the second phoneme sequence obtained after deduplication is "wo(3)z ou(3)le(1)", where the numbers in parentheses are tones.

[0051] Step 204: Recognize the second phoneme sequence to obtain the speech content.

[0052] Furthermore, recognizing the deduplicated second phoneme sequence can improve the efficiency and accuracy of speech content recognition. One approach is to input the second phoneme sequence into a trained language model to identify the corresponding speech content, such as text data.

[0053] like Figure 3 As shown, Figure 3 The diagram illustrates a speech content recognition method. A first acoustic model, trained from scratch, identifies the first speech signals from multiple channels to obtain the corresponding second phoneme sequence. The obtained second phoneme sequence is then input into a language model for recognition to obtain the corresponding speech content.

[0054] In the speech signal recognition method of this application embodiment, first speech signals from multiple channels are acquired. The first speech signals from each channel are original speech signals synchronously collected within a set time period. The first speech signals from multiple channels are input into a first acoustic model trained to obtain a corresponding first phoneme sequence. This achieves recognition based on global information of the first speech signals from multiple channels, resulting in a corresponding phoneme sequence. This achieves low signal distortion and high purity, improving the accuracy of the phoneme sequence. Furthermore, repeated phonemes in the first phoneme sequence are deduplicated before recognition to obtain the speech content, thereby improving the recognition efficiency and quality of the speech content.

[0055] The above embodiments utilize the trained first acoustic model. Based on the above embodiments, Figure 4 This is a flowchart illustrating another speech signal recognition method provided in an embodiment of this application, specifically explaining the training method of the first acoustic model, such as... Figure 4 As shown, the method includes the following steps:

[0056] Step 401: Obtain the first training sample set.

[0057] Each first training sample in the first training sample set contains second speech signals from multiple channels. The second speech signals from each channel are original sample speech signals synchronously collected within a set time period. Each first training sample is labeled with a corresponding third phoneme sequence. The third phoneme sequence can be a phoneme sequence corresponding to the second speech signals of each channel determined manually, or it can be obtained based on other models. This will be described in detail in subsequent embodiments.

[0058] The description of the first voice signal of each channel in the aforementioned embodiments also applies to the second voice signal of each channel, and the principle is the same, so it will not be repeated in this embodiment.

[0059] It should be noted that the first speech signal with multiple channels and the second speech signal with multiple channels are only used for differentiation, and the first training sample may also contain the first speech signal with multiple channels.

[0060] like Figure 5 As shown, taking a first training sample as an example, the second speech signal of multiple channels includes the second speech signal of the sound source channel and the second speech signal of each recording channel, as well as the labeled third phoneme sequence. Among them, Ch0 is the second speech signal of the sound source channel, such as the speaker channel, and Ch1, Ch2, ..., ChN are the second speech signals of each recording channel, such as the microphone channel.

[0061] Step 402: For each first training sample, input the first training sample into the first acoustic model to obtain the fourth phoneme sequence corresponding to the first training sample.

[0062] Step 403: Adjust the parameters of the first acoustic model based on the difference between the fourth phoneme sequence and the labeled third phoneme sequence.

[0063] In this embodiment, for each first training sample, the corresponding first training sample is input into the first acoustic model to obtain the recognized fourth phoneme sequence corresponding to the first training sample. This fourth phoneme sequence is merely a marker of different phoneme sequences. Then, based on the difference between the recognized fourth phoneme sequence and the labeled third phoneme sequence, a corresponding loss function is determined. Based on the loss function, the parameters of the first acoustic model are adjusted. The parameters of the first acoustic model are continuously adjusted using multiple training samples in the training sample set until the difference between the recognized fourth phoneme sequence and the labeled third phoneme sequence is minimized, at which point the model training is complete.

[0064] like Figure 5 As shown, the training network of the first acoustic model is used to train the multi-channel of the first acoustic model to obtain the trained second acoustic model.

[0065] In the speech signal recognition method of this application embodiment, the first acoustic model is trained using the second speech signal from multiple channels as training samples. This realizes the training of the first acoustic model based on the original second speech signal from multiple channels without framing signal processing. Since the signal is not segmented, the global information of the speech signal is used, which improves the training effect of the first acoustic model.

[0066] Based on the above embodiments, Figure 6 A flowchart illustrating another speech signal recognition method provided in this application embodiment is shown below. Figure 6 As shown, this illustrates how the annotation information for the training samples of the first acoustic model is determined to improve the efficiency of training sample generation. Before step 401, the method includes the following steps:

[0067] Step 501: Acquire multiple sets of second speech signals from multiple channels.

[0068] In this context, the second speech signal from each group of multiple channels is used to generate the first training sample.

[0069] Step 502: For the second speech signals of multiple channels in each group, perform speech signal processing based on the second speech signals of multiple channels to obtain an enhanced single-channel first target speech signal.

[0070] As one implementation method, for the second speech signals of multiple channels in each group, beamforming is used to obtain a single-channel second speech signal. Then, the single-channel second speech signal is enhanced by a post-filter and converted into the time domain to obtain an enhanced single-channel first target speech signal.

[0071] Step 503: Input the first target speech signal from a single channel into the trained second acoustic model to obtain the corresponding third phoneme sequence.

[0072] The second acoustic model, for example, is a mixture of Gaussian hidden Markov models (GMM-HMM).

[0073] The second acoustic model has learned the correspondence between the enhanced single-channel first target speech signal and the corresponding third phoneme sequence through training. The training method of the second acoustic model will be described in detail in subsequent embodiments and will not be repeated here.

[0074] Step 504: Generate a first training sample set based on multiple sets of second speech signals from multiple channels and the corresponding third phoneme sequences.

[0075] In this embodiment of the application, the second acoustic model obtained through training is used to identify the second speech signals of multiple channels, thereby obtaining the third phoneme sequence corresponding to the second speech signals of multiple channels. The third phoneme sequence is used as the annotation information of the corresponding second speech signals of multiple channels, that is, the standard phoneme sequence. Compared with the manual annotation method, the annotation efficiency is improved.

[0076] In the speech signal recognition method of this application embodiment, the second acoustic model obtained through training is used to recognize the second speech signals of multiple channels and obtain the annotation information of the second speech signals of multiple channels, that is, the standard phoneme sequence. Compared with the manual annotation method, the annotation efficiency is improved.

[0077] Based on the above embodiments, Figure 7 A flowchart illustrating another speech signal recognition method provided in this application embodiment is shown below. Figure 7 As shown, this illustrates how to process multiple channels of speech signals to obtain an enhanced single-channel first target speech signal. Step 502 includes the following steps:

[0078] Step 601: Based on the second speech signal of the sound source channel, perform echo cancellation on the second speech signal of each recording channel to obtain the echo-cancelled second speech signal of each recording channel.

[0079] As an example, Figure 8 This is a schematic diagram of a speech signal enhancement structure provided in an embodiment of this application. Figure 8 As shown, Ch0' is the second speech signal of the sound source channel, and Ch1', Ch2', ..., ChN' are the second speech signals of each recording channel.

[0080] In this embodiment, the first speech signals of multiple recording channels contain interference signals of acoustic echo. Acoustic echo is caused by the sound from the speaker repeatedly feeding back to the microphone in hands-free or conferencing applications. Therefore, echo cancellation is required for the first speech signals of multiple recording channels. One implementation involves determining the acoustic transfer function for the transmission of the second speech signal from the sound source channel. Echo estimation is performed on the second speech signal from the sound source channel based on the acoustic transfer function to obtain the estimated echo signal. Based on the second speech signal and echo signal of each recording channel, the echo-cancelled second speech signal for each recording channel is obtained. Specifically, the acoustic transfer function including the reflection path from the sound-generating device to the recording device, such as from the speaker to the microphone, is estimated. Then, the Wienerhof equation for echo cancellation can be constructed, and the acoustic transfer function is solved by inversion. Furthermore, the second speech signal from the incoming sound source channel is filtered using the estimated acoustic transfer function to obtain the estimated echo signal. Finally, this estimated echo signal is subtracted from the second speech signal of each recording channel to obtain the echo-removed second speech signal for each recording channel. Echo cancellation improves the accuracy of the second speech signals of multiple recording channels.

[0081] It should be noted that when performing echo cancellation on the second speech signals of multiple channels, the second speech signals of each channel are not processed by frame segmentation. That is, this application does not process each frame of speech signal obtained from speech signal processing. In other words, when processing the second speech signals of each channel, this application uses the original speech signals collected within a set time period to utilize the global information of the second speech signal of that channel. Compared with the streaming signal processing method that uses local information of speech frames, the global information carries the complete context information of each frame, which can improve the effect of speech signal recognition.

[0082] Step 602: Beamforming is performed on the second speech signals from the multiple recording channels with echo cancellation to obtain a single-channel second speech signal.

[0083] In this embodiment, by adjusting the basic unit parameters of the phase array, signals at certain angles achieve constructive interference, while signals at other angles achieve destructive interference. A beammap is generated from the second speech signals of each recording channel. The azimuth (angle) of the main lobe or peak of the beammap is determined. The second speech signal with the largest signal response in each recording channel is identified, indicating that the beam output power corresponding to that channel is 1, meaning the estimated signal power arriving in that direction is 1. Then, an adaptive beamforming method, namely Minimum Variance Distortionless Response (MVDR), is used to determine the weights corresponding to the second speech signals of each recording channel. The second speech signals of each recording channel are weighted, summed, and filtered to finally output the speech signal in the desired direction, effectively forming a "beam." This beamforming of the echo-cancelled second speech signals from multiple recording channels yields a single-channel second speech signal. In this application, weighted merging of the second speech signals from multiple recording channels suppresses interference signals from non-target directions, thereby enhancing the single-channel second speech signal obtained after beamforming.

[0084] Step 603: The single-channel second speech signal is enhanced by a post-filter.

[0085] In this embodiment, the single-channel second speech signal also contains interference signals from non-target directions that have not been completely suppressed, resulting in residual noise or interference in the single-channel second speech signal. Therefore, it is necessary to perform filtering processing again. By setting the filtering parameters of the post-Wiener filter, speech enhancement is performed on the single-channel second speech signal to obtain a cleaner enhanced single-channel second speech signal.

[0086] Step 604: Perform an inverse Fourier transform on the enhanced single-channel second speech signal to obtain the single-channel first target speech signal.

[0087] In this application, the enhanced single-channel second speech signal is transformed from the frequency domain to the time domain by performing an inverse Fourier transform, so as to facilitate subsequent data processing.

[0088] In the speech signal recognition method of this application embodiment, the second speech signals of multiple channels to be processed are acquired, and echo cancellation is performed to obtain the second speech signals of multiple recording channels with echo cancellation. Based on the second speech signals of the multiple recording channels with echo cancellation, beamforming is performed to obtain a single-channel second speech signal. The single-channel second speech signal is then enhanced to obtain an enhanced single-channel second speech signal. In this method, the second speech signals of each channel are not processed by frame segmentation, thereby realizing the acquisition of a single-channel target first speech signal based on the global information of the speech signals of multiple recording channels. This achieves lower signal distortion and higher purity, and improves the quality of the target first speech signal.

[0089] The above embodiments utilize a trained second acoustic model to identify and annotate the third phoneme sequence from the first training samples used in the training process of the first acoustic model, thereby improving annotation efficiency. Based on the above embodiments, Figure 9 This is a flowchart illustrating another speech signal recognition method provided in an embodiment of this application, specifically illustrating the training method of the second acoustic model, such as... Figure 9 As shown, the method includes the following steps:

[0090] Step 801: Obtain the second training sample set.

[0091] The second training sample set contains multiple second training samples. Each second training sample contains an enhanced single-channel second target speech signal and a corresponding standard phoneme sequence. The enhanced single-channel second target speech signal is obtained by processing the third speech signals of multiple channels. The third speech signals of each channel are the original speech signals synchronously collected within a set time period.

[0092] As an example, a standard phoneme sequence can be manually obtained by mapping the corresponding standard text using a predefined pronunciation dictionary. For example, if the standard text is "Dragon Ball: Earthlings Are the Strongest", the phoneme sequence obtained by mapping the text using a predefined pronunciation dictionary is: long2 zh u1 zh ix1 d i4 q iu2 r en2 z ui4 q iang2. The numbers represent the tones of the pronunciation. For example, 2 is the second tone in pinyin, and 4 is the fourth tone in pinyin. Another example is "Dragon Ball: The Battle of the Strongest". The phoneme sequence obtained by mapping the text using a predefined pronunciation dictionary is: l ong2 zh u1 zh ix1 qiang2 zh e3 zh eng1 b a4.

[0093] It should be noted that the explanation of the enhanced single-channel first target speech signal in the foregoing embodiments also applies to the enhanced single-channel second target speech signal in this embodiment. The principle is the same, and it will not be repeated in this embodiment.

[0094] Step 802: For each second training sample, input the second training sample into the second acoustic model to predict the fifth phoneme sequence corresponding to the second training sample.

[0095] In this embodiment, the second acoustic model, for example, is a mixture Gaussian Hidden Markov Model (GMM-HMM). The training samples of this model are enhanced single-channel second target speech signals. Through training, more accurate phoneme sequences can be obtained. Then, the more accurate phoneme sequences are used as the annotation information of the training samples of the first acoustic model to train the first acoustic model, which can improve the training efficiency and effect of the first acoustic model. At the same time, the trained first acoustic model can directly recognize speech signals from multiple channels without the need for speech signal processing, thus improving the recognition efficiency.

[0096] The fifth phoneme sequence is also a frame-level phoneme sequence, indicating the phonemes corresponding to each frame. The explanation of the first phoneme sequence in the preceding embodiments is similar in principle and will not be repeated here.

[0097] Step 803: Adjust the parameters of the second acoustic model based on the accuracy of the fifth phoneme sequence.

[0098] In this embodiment, for each second training sample, when the second recognition model performs phoneme forced alignment on each frame of the third speech signal across multiple channels, it can determine which phonemes correspond to each other and their order based on the labeled fifth phoneme sequence. This determines the fifth phoneme sequence corresponding to the third speech signal across multiple channels. Furthermore, as one implementation, the accuracy of the fifth phoneme sequence can be assessed, and the parameters of the second acoustic model can be adjusted based on the accuracy. When the accuracy meets the set requirements, the model parameter adjustment is complete, and the second acoustic model training is finished. Alternatively, the parameters of the second acoustic model can be adjusted according to the accuracy of the fifth phoneme sequence based on a set number of model iterations until the required number of iterations is reached, at which point the second acoustic model training is complete.

[0099] When performing frame segmentation processing on the third audio signal of multiple channels, the frame segmentation processing can be performed according to a 25ms time window and a 10ms frame shift. In addition, the audio signals of other multiple channels appearing in this embodiment can also be processed according to the corresponding time window and frame shift to achieve frame segmentation.

[0100] It is important to understand that during the training process, the second acoustic model can determine the phonemes that appear and their order of appearance based on the labeled phoneme sequence. However, since the starting frame of each phoneme and how many consecutive frames each phoneme appears are unknown, the phoneme label for each frame is also unknown. Based on the labeled phoneme sequence, it utilizes some prior knowledge and is therefore an incomplete unsupervised training.

[0101] In the training method of the acoustic model in this application embodiment, the enhanced single-channel second target speech signal used as the training sample is the original third target speech signal with multiple channels that has not undergone frame segmentation. That is, it is an enhanced speech signal obtained based on global information processing. It is used to train the second acoustic model, which improves the training effect of the model and thus obtains a high-quality phoneme sequence.

[0102] To achieve the above embodiments, this application also proposes a voice signal recognition device.

[0103] Figure 10 This is a schematic diagram of the structure of a voice signal recognition device provided in an embodiment of this application.

[0104] like Figure 10 As shown, the device includes:

[0105] The acquisition module 91 is used to acquire the first voice signals of multiple channels; wherein the first voice signal of each channel is the original voice signal synchronously acquired within a set time period.

[0106] The processing module 92 is used to input the first speech signal of the multiple channels into the first acoustic model obtained by training, and obtain the corresponding first phoneme sequence.

[0107] The recognition module 93 is used to recognize the first phoneme sequence to obtain the speech content.

[0108] Furthermore, in one implementation of this application embodiment, the identification module 93 is specifically used for:

[0109] Multiple consecutive identical phonemes in the first phoneme sequence are merged to obtain a second phoneme sequence; the second phoneme sequence is then recognized to obtain the speech content.

[0110] In one implementation of this application embodiment, the identification module 93 is specifically used for:

[0111] Based on the multiple phonemes arranged sequentially in the first phoneme sequence, at least one phoneme group is determined; the phoneme group contains multiple adjacent identical phonemes; the identical phonemes in each of the phoneme groups are merged to obtain the second phoneme sequence.

[0112] In one implementation of this application, the method further includes a first training module, wherein the first acoustic model is obtained in the following manner:

[0113] A first training module is used to acquire a first training sample set; each first training sample in the first training sample set contains a second speech signal from multiple channels, and the second speech signal from each channel is an original sample speech signal synchronously acquired within a set time period; each first training sample is labeled with a corresponding third phoneme sequence; for each first training sample, the first training sample is input into the first acoustic model to obtain a fourth phoneme sequence corresponding to the first training sample; the parameters of the first acoustic model are adjusted according to the difference between the fourth phoneme sequence and the labeled third phoneme sequence.

[0114] In one implementation of this application, the method further includes:

[0115] An enhancement module is used to acquire multiple sets of second speech signals from the multiple channels; and to perform speech signal processing on each set of second speech signals from the multiple channels to obtain an enhanced single-channel first target speech signal.

[0116] The generation module is used to input the first target speech signal of the single channel into the trained second acoustic model to obtain the corresponding third phoneme sequence; and to generate the first training sample set based on the multiple sets of second speech signals of the multiple channels and the corresponding third phoneme sequences.

[0117] As one implementation, the second speech signal from multiple channels is sampled from the sound source channel and multiple recording channels. The enhancement module is specifically used for:

[0118] Based on the second speech signal of the sound source channel, echo cancellation is performed on the second speech signal of each recording channel to obtain the echo-cancelled second speech signal of each recording channel.

[0119] Beamforming is performed on the second speech signals from the multiple recording channels of the echo cancellation to obtain a single-channel second speech signal;

[0120] The single-channel second speech signal is enhanced by a post-filter;

[0121] The enhanced single-channel second speech signal is subjected to an inverse Fourier transform to obtain the single-channel first target speech signal.

[0122] As one implementation method, the enhancement module is also specifically used for:

[0123] Determine the acoustic transfer function for the transmission of the second speech signal in the sound source channel;

[0124] The second speech signal of the sound source channel is estimated based on the acoustic transfer function to obtain the estimated echo signal;

[0125] Based on the second speech signal of each of the recording channels and the echo signal, the second speech signal of each recording channel with echo cancellation is obtained.

[0126] As one implementation, the device further includes a second training module, wherein the second acoustic model is obtained in the following manner:

[0127] The second training module is used to acquire a second training sample set; wherein the second training sample set contains multiple second training samples, each second training sample containing an enhanced single-channel second target speech signal and a corresponding standard phoneme sequence; the enhanced single-channel second target speech signal is obtained by processing the third speech signals of multiple channels, wherein the third speech signals of each channel are the original speech signals synchronously acquired within a set duration; for each second training sample, the second training sample is input into the second acoustic model to predict the fifth phoneme sequence corresponding to the second training sample; the parameters of the second acoustic model are adjusted according to the accuracy of the fifth phoneme sequence.

[0128] It should be noted that the foregoing explanation of the method embodiments also applies to the apparatus of this embodiment, and will not be repeated here.

[0129] In the speech signal recognition device of this application embodiment, first speech signals from multiple channels are acquired. The first speech signals from each channel are original speech signals synchronously collected within a set time period. The first speech signals from multiple channels are input into a first acoustic model trained to obtain a corresponding first phoneme sequence. The speech content is obtained by recognizing the first phoneme sequence. This realizes the recognition of global information based on the first speech signals from multiple channels to obtain the speech content. It achieves low signal distortion and high purity, thereby improving the quality of the speech content.

[0130] To implement the above embodiments, this application also proposes an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, it implements the method described in the foregoing method embodiments.

[0131] To implement the above embodiments, this application also proposes a non-transitory computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the method described in the foregoing method embodiments.

[0132] To implement the above embodiments, this application also proposes a computer program product having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the method described in the foregoing method embodiments.

[0133] Figure 11 This is a block diagram of an electronic device provided in an embodiment of this application. For example, the electronic device 800 may be a mobile phone, computer, digital broadcasting terminal, messaging device, game console, tablet device, medical device, fitness equipment, personal digital assistant, etc.

[0134] Reference Figure 11 The electronic device 800 may include one or more of the following components: a processing component 818, a memory 804, a power component 806, a multimedia component 808, an audio component 810, an input / output (I / O) interface 812, a sensor component 814, and a communication component 816.

[0135] Processing component 818 typically controls the overall operation of electronic device 800, such as operations associated with display, telephone calls, data communication, camera operation, and recording operations. Processing component 818 may include one or more processors 820 to execute instructions to complete all or part of the steps of the methods described above. Furthermore, processing component 818 may include one or more modules to facilitate interaction between processing component 818 and other components. For example, processing component 818 may include a multimedia module to facilitate interaction between multimedia component 808 and processing component 818.

[0136] Memory 804 is configured to store various types of data to support the operation of electronic device 800. Examples of this data include instructions for any application or method operating on electronic device 800, contact data, phonebook data, messages, pictures, videos, etc. Memory 804 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.

[0137] Power component 806 provides power to various components of electronic device 800. Power component 806 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to electronic device 800.

[0138] Multimedia component 808 includes a screen that provides an output interface between the electronic device 800 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touchscreen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors may sense not only the boundaries of the touch or swipe action but also the duration and pressure associated with the touch or swipe operation. In some embodiments, multimedia component 808 includes a front-facing camera and / or a rear-facing camera. When the electronic device 800 is in an operating mode, such as a shooting mode or a video mode, the front-facing camera and / or the rear-facing camera may receive external multimedia data. Each front-facing camera and rear-facing camera may be a fixed optical lens system or have focal length and optical zoom capabilities.

[0139] Audio component 810 is configured to output and / or input audio signals. For example, audio component 810 includes a microphone (MIC) configured to receive external audio signals when electronic device 800 is in an operating mode, such as call mode, recording mode, and voice recognition mode. The received audio signals may be further stored in memory 804 or transmitted via communication component 816. In some embodiments, audio component 810 also includes a speaker for outputting audio signals.

[0140] I / O interface 812 provides an interface between processing component 818 and peripheral interface modules, such as keyboards, click wheels, buttons, etc. These buttons may include, but are not limited to, home buttons, volume buttons, power buttons, and lock buttons.

[0141] Sensor assembly 814 includes one or more sensors for providing state assessments of various aspects of electronic device 800. For example, sensor assembly 814 can detect the on / off state of electronic device 800, the relative positioning of components such as the display and keypad of electronic device 800, changes in position of electronic device 800 or a component of electronic device 800, the presence or absence of user contact with electronic device 800, orientation or acceleration / deceleration of electronic device 800, and temperature changes of electronic device 800. Sensor assembly 814 may include a proximity sensor configured to detect the presence of nearby objects without any physical contact. Sensor assembly 814 may also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, sensor assembly 814 may also include an accelerometer, gyroscope, magnetometer, pressure sensor, or temperature sensor.

[0142] Communication component 816 is configured to facilitate wired or wireless communication between electronic device 800 and other devices. Electronic device 800 can access wireless networks based on communication standards, such as WiFi, 4G, or 5G, or combinations thereof. In one exemplary embodiment, communication component 816 receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In one exemplary embodiment, communication component 816 also includes a near-field communication (NFC) module to facilitate short-range communication. For example, the NFC module may be implemented based on radio frequency identification (RFID) technology, Infrared Data Association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.

[0143] In an exemplary embodiment, the electronic device 800 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the methods described above.

[0144] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 804 including instructions, which can be executed by a processor 820 of an electronic device 800 to perform the above-described method. For example, the non-transitory computer-readable storage medium may be a ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc.

[0145] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., refer to specific features, structures, materials, or characteristics described in connection with that embodiment or example, which are included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.

[0146] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this application, "multiple" means at least two, such as two, three, etc., unless otherwise explicitly specified.

[0147] Any process or method description in the flowchart or otherwise herein can be understood as representing a module, segment, or portion of code comprising one or more executable instructions for implementing custom logic functions or processes, and the scope of the preferred embodiments of this application includes additional implementations in which functions may be performed not in the order shown or discussed, including substantially simultaneously or in reverse order depending on the functions involved, as should be understood by those skilled in the art to which embodiments of this application pertain.

[0148] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a processor-included system, or other system that can fetch and execute instructions from, an instruction execution system, apparatus, or device). For the purposes of this specification, "computer-readable medium" can be any means that can contain, store, communicate, propagate, or transmit programs for use by, or in conjunction with, an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of computer-readable media include: an electrical connection having one or more wires (electronic device), a portable computer disk drive (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic devices, and portable optical disc read-only memory (CDROM). Alternatively, the computer-readable medium may be paper or other suitable media on which the program can be printed, since the program can be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, interpreting, or otherwise processing as necessary, and then stored in a computer memory.

[0149] It should be understood that various parts of this application can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented using software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.

[0150] Those skilled in the art will understand that all or part of the steps of the methods in the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, the program includes one or a combination of the steps of the method embodiments.

[0151] Furthermore, the functional units in the various embodiments of this application can be integrated into a processing module, or each unit can exist physically separately, or two or more units can be integrated into a module. The integrated module can be implemented in hardware or as a software functional module. If the integrated module is implemented as a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium.

[0152] The storage medium mentioned above can be a read-only memory, a disk, or an optical disk, etc. Although embodiments of this application have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting this application. Those skilled in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of this application.

Claims

1. A speech signal recognition method, characterized in that, include: Acquire the first speech signals from multiple channels; wherein, the first speech signal of each channel is the original speech signal synchronously acquired within a set time period; The first speech signal from the multiple channels is input into the first acoustic model obtained through training to obtain the corresponding first phoneme sequence; The speech content is obtained by recognizing the first phoneme sequence; The first acoustic model was obtained in the following way: Acquire multiple sets of second voice signals from the multiple channels; For the second speech signals of the multiple channels in each group, speech signal processing is performed based on the second speech signals of the multiple channels to obtain an enhanced single-channel first target speech signal; The first target speech signal of the single channel is input into the trained second acoustic model to obtain the corresponding third phoneme sequence; A first training sample set is generated based on the second speech signals from the multiple channels and the corresponding third phoneme sequences. For each of the first training samples in the first training sample set, the first training sample is input into the first acoustic model to obtain the fourth phoneme sequence corresponding to the first training sample. The parameters of the first acoustic model are adjusted based on the difference between the fourth phoneme sequence and the labeled third phoneme sequence.

2. The method as described in claim 1, characterized in that, The speech content is obtained by recognizing the first phoneme sequence, including: Multiple consecutive identical phonemes in the first phoneme sequence are merged to obtain the second phoneme sequence; The second phoneme sequence is identified to obtain the speech content.

3. The method as described in claim 2, characterized in that, The step of merging multiple consecutive identical phonemes in the first phoneme sequence to obtain a second phoneme sequence includes: Based on the multiple phonemes arranged sequentially in the first phoneme sequence, at least one phoneme group is determined; the phoneme group contains multiple adjacent identical phonemes; The same phoneme in each of the phoneme groups is merged to obtain the second phoneme sequence.

4. The method as described in claim 1, characterized in that, The second speech signals from the multiple channels are sampled from the sound source channel and multiple recording channels. The speech signal processing based on the second speech signals from the multiple channels to obtain an enhanced single-channel first target speech signal includes: Based on the second speech signal of the sound source channel, echo cancellation is performed on the second speech signal of each recording channel to obtain the echo-cancelled second speech signal of each recording channel. Beamforming is performed on the second speech signals from the multiple recording channels of the echo cancellation to obtain a single-channel second speech signal; The single-channel second speech signal is enhanced by a post-filter; The enhanced single-channel second speech signal is subjected to an inverse Fourier transform to obtain the single-channel first target speech signal.

5. The method as described in claim 4, characterized in that, The step of performing echo cancellation on the second speech signals of each recording channel based on the second speech signal of the sound source channel to obtain echo-cancelled second speech signals of each recording channel includes: Determine the acoustic transfer function for the transmission of the second speech signal in the sound source channel; The second speech signal of the sound source channel is estimated based on the acoustic transfer function to obtain the estimated echo signal; Based on the second speech signal of each of the recording channels and the echo signal, the second speech signal of each recording channel with echo cancellation is obtained.

6. The method as described in claim 1, characterized in that, The second acoustic model was obtained in the following way: Obtain a second training sample set; wherein the second training sample set contains multiple second training samples, each of which contains an enhanced single-channel second target speech signal and a corresponding standard phoneme sequence; the enhanced single-channel second target speech signal is obtained by processing the third speech signals of multiple channels, wherein the third speech signal of each channel is the original speech signal synchronously acquired within a set time period; For each of the second training samples, the second training sample is input into the second acoustic model to predict the fifth phoneme sequence corresponding to the second training sample; The parameters of the second acoustic model are adjusted based on the accuracy of the fifth phoneme sequence.

7. A voice signal recognition device, characterized in that, include: The acquisition module is used to acquire the first voice signals from multiple channels; wherein, the first voice signal of each channel is the original voice signal synchronously acquired within a set time period; The processing module is used to input the first speech signals from the multiple channels into the first acoustic model obtained through training, and obtain the corresponding first phoneme sequence; The recognition module is used to recognize the first phoneme sequence to obtain the speech content; The first acoustic model was obtained in the following way: Acquire multiple sets of second voice signals from the multiple channels; For the second speech signals of the multiple channels in each group, speech signal processing is performed based on the second speech signals of the multiple channels to obtain an enhanced single-channel first target speech signal; The first target speech signal of the single channel is input into the trained second acoustic model to obtain the corresponding third phoneme sequence; A first training sample set is generated based on the second speech signals from the multiple channels and the corresponding third phoneme sequences. For each of the first training samples in the first training sample set, the first training sample is input into the first acoustic model to obtain the fourth phoneme sequence corresponding to the first training sample. The parameters of the first acoustic model are adjusted based on the difference between the fourth phoneme sequence and the labeled third phoneme sequence.

8. An electronic device, characterized in that, It includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, it implements the method as described in any one of claims 1-6.

9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1-6.

Citation Information

Patent Citations

  • Voice awakening method and device, electronic equipment and storage medium

    CN111933111A

  • Methods and apparatus for reducing spurious insertions in speech recognition

    US20040199385A1