Speech recognition method, device, equipment, and storage medium
The speech recognition method enhances target speaker identification in mixed voice environments by using a feature extraction model and joint training of models, improving accuracy and user experience.
Patent Information
- Application Number
- JP2024514680
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-09-07
- Filing Date
- 2021-11-10
- Publication Date
- 2025-12-16
- Estimated Expiration
- 2041-11-10
AI Technical Summary
Existing speech recognition technologies struggle to accurately identify the speech content of a target speaker in mixed voice environments, where multiple speakers and background noise are present, leading to reduced effectiveness and user experience.
A speech recognition method that involves extracting target speaker features from mixed speech using a feature extraction model trained on enrollment speech, followed by a joint training of the feature extraction and speech recognition models to enhance accuracy, utilizing multi-scale voiceprint features and a pre-created model to improve feature extraction and recognition.
The method achieves relatively accurate speech recognition results and improves user experience by effectively separating target speaker speech from mixed speech, addressing the limitations of independent model training and cascading errors.
Smart Images

Figure 0007786691000024 
Figure 0007786691000025 
Figure 0007786691000026
Abstract
Description
[Technical Field]
[0001] This application claims priority to a Chinese patent application with application number CN202111042821.8, filed with the China Patent Office on September 7, 2021, and titled "Speech recognition method, device, equipment and storage medium," the entire contents of which are hereby incorporated by reference into this application.
[0002] The present application belongs to the field of speech recognition technology, and in particular to a speech recognition method, device, equipment and storage medium. [Background technology]
[0003] With the rapid development of artificial intelligence technology, smart devices are playing an increasingly important role in people's lives, and voice interaction is favored by users as the most convenient and natural human-machine interaction method.
[0004] When a user uses a smart device, he or she may be in a complex environment where other people's voices are present, and in this case, the voice collected by the smart device is a mixed voice. In order to achieve a good user experience during voice interaction, it is necessary to recognize the voice content of the target speaker from the mixed voice. Therefore, how to recognize the voice content of the target speaker from the mixed voice has become an urgent issue. Summary of the Invention [Problem to be solved by the invention]
[0005] In view of this, the present application provides a speech recognition method, device, equipment and storage medium for accurately recognizing the speech content of a target speaker from mixed speech, and its technical solutions are as follows: [Means for solving the problem]
[0006] 1. A speech recognition method, comprising: Acquiring speech characteristics of a target speech mixture and speaker characteristics of a target speaker; an extraction direction is to approach target speech features as speech features used to obtain speech recognition results that match the actual speech content of the target speaker, and extract speech features of the target speaker from the speech features of the target mixed speech based on the speech features of the target mixed speech and the speaker features of the target speaker, thereby obtaining extracted speech features of the target speaker; and obtaining a speech recognition result for the target speaker based on the extracted speech features of the target speaker.
[0007] Optionally, obtaining speaker features of the target speaker comprises: acquiring an enrollment speech of the target speaker; The method includes extracting short-term and long-term voiceprint features from the enrollment speech of the target speaker, and using the obtained multi-scale voiceprint features as speaker features of the target speaker.
[0008] As an option, the extraction direction of the target speech feature is to extract speech features of the target speaker from the speech features of the target speech mixture based on the speech features of the target speech mixture and the speaker features of the target speaker, extracting speech features of the target speaker from the speech features of the target speech mixture based on speech features of the target speech mixture and speaker features of the target speaker using a pre-created feature extraction model; Here, the feature extraction model is a model of speech features of a training speech mixture including speech of a specified speaker. and obtained by training speech recognition results obtained based on extracted speech features of the designated speaker as an optimization target using speaker features of the designated speaker, wherein the extracted speech features of the designated speaker are speech features of the designated speaker extracted from speech features of the training speech mixture.
[0009] Optionally, the feature extraction model is obtained by training the extracted speech features of the designated speaker and the speech recognition results obtained based on the extracted speech features of the designated speaker as optimization targets.
[0010] As an option, using the pre-created feature extraction model and extracting speech features of the target speaker from speech features of the target speech mixture based on speech features of the target speech mixture and speaker features of the target speaker, inputting the speech features of the target speech mixture and the speaker features of the target speaker into the feature extraction model to obtain a feature mask corresponding to the target speaker; and extracting speech features of the target speaker from the speech features of the target speech mixture based on the speech features of the target speech mixture and a feature mask corresponding to the target speaker.
[0011] Optionally, obtaining a speech recognition result for the target speaker based on the extracted speech features of the target speaker includes: obtaining a speech recognition result for the target speaker based on the extracted speech features of the target speaker and the enrolled speech features of the target speaker; Here, the enrollment speech features of the target speaker are speech features of the enrollment speech of the target speaker.
[0012] Optionally, obtaining a speech recognition result for the target speaker based on the extracted speech features of the target speaker includes: inputting speech recognition input features including at least the extracted speech features of the target speaker into a pre-created speech recognition model to obtain a speech recognition result for the target speaker; The speech recognition model is obtained by joint training with the feature extraction model, and the speech recognition model is obtained by training using the extracted speech features of the specified speaker and using speech recognition results obtained based on the extracted speech features of the specified speaker as an optimization target.
[0013] Optionally, inputting the speech recognition input features into the speech recognition model to obtain a speech recognition result for the target speaker includes: Encoding the speech recognition input features based on an encoder module of the speech recognition model to obtain an encoding result; extracting an audio-related feature vector required for decoding at a decoding time from the encoding result based on an attention module of the speech recognition model; and decoding the audio-related feature vector extracted from the encoding result based on a decoder module of the speech recognition model to obtain the recognition result at the decoding time.
[0014] Optionally, the process of jointly training the speech recognition model and the feature extraction model comprises: extracting speech features of the designated speaker from speech features of the training speech mixture using a feature extraction model to obtain extracted speech features of the designated speaker; obtaining a speech recognition result for the designated speaker using a speech recognition model and the extracted speech features of the designated speaker; and The parameters of a feature extraction model are updated based on the extracted voice features of the designated speaker and the voice recognition result of the designated speaker, and the voice recognition is performed based on the voice recognition result of the designated speaker. and updating the parameters of the voice recognition model.
[0015] Optionally, the training speech mixtures correspond to speech of the designated speaker; updating parameters of a feature extraction model based on the extracted speech features of the designated speaker and the speech recognition result of the designated speaker, and updating parameters of a speech recognition model based on the speech recognition result of the designated speaker; Obtaining annotated text of the speech of the designated speaker, and obtaining speech features of the speech of the designated speaker as standard speech features of the designated speaker; determining a first predicted loss based on the extracted speech features of the designated speaker and the standard speech features of the designated speaker, and determining a second predicted loss based on the speech recognition result of the designated speaker and the annotated text of the speech of the designated speaker; updating parameters of a feature extraction model based on the first prediction loss and the second prediction loss, and updating parameters of a speech recognition model based on the second prediction loss.
[0016] Optionally, the training speech mixtures and the designated speaker's speech corresponding to the training speech mixtures are obtained from a pre-created training data set; The training dataset construction process includes: obtaining a plurality of speeches from a plurality of speakers, each of which comprises a single speaker's speech annotated with text; one of the plurality of speeches is a speech of a designated speaker, and one or more speeches of other speakers among the plurality of speeches are mixed with the speech of the designated speaker to obtain a training speech mixture, and the training speech mixture obtained by mixing the speech of the designated speaker is used as training data; All training data obtained constitutes the training data set.
[0017] A speech recognition device, a feature acquisition module used to acquire speech features of the target speech mixture and speaker features of the target speaker; a feature extraction module that extracts target speech features from the speech features of the target speech mixture based on the speech features of the target speech mixture and the speaker features of the target speaker, and obtains extracted speech features of the target speaker; a speech recognition module used to obtain a speech recognition result for the target speaker based on the extracted speech features of the target speaker; Here, the target speech features are speech features used to obtain speech recognition results that match the actual speech content of the target speaker.
[0018] Optionally, the feature acquisition module: The system includes a speaker feature acquisition module that is used to acquire an enrollment speech of the target speaker, extract short-term voiceprint features and long-term voiceprint features from the enrollment speech of the target speaker, and use the obtained multi-scale voiceprint features as speaker features of the target speaker.
[0019] As an option, the feature extraction module is specifically used to extract the speech features of the target speaker from the speech features of the target speech mixture based on the speech features of the target speech mixture and the speaker features of the target speaker using a pre-made feature extraction model; Here, the feature extraction model is obtained by training using speech features of a training speech mixture including speech of a designated speaker and speaker features of the designated speaker, with speech recognition results obtained based on the extracted speech features of the designated speaker as an optimization target, and the extracted speech features of the designated speaker are the speech features of the designated speaker extracted from the speech features of the training speech mixture.
[0020] Optionally, the speech recognition module is specifically used to obtain a speech recognition result of the target speaker based on the extracted speech features of the target speaker and the enrolled speech features of the target speaker; Here, the enrollment speech features of the target speaker are speech features of the enrollment speech of the target speaker.
[0021] As an option, the speech recognition module is specifically used to input speech recognition input features including at least the extracted speech features of the target speaker into a pre-made speech recognition model to obtain a speech recognition result of the target speaker; Here, the speech recognition model is obtained by joint training with the feature extraction model, and the speech recognition model is obtained by training using the extracted speech features of the specified speaker and using the speech recognition results obtained based on the extracted speech features of the specified speaker as an optimization target.
[0022] A voice recognition device, a memory used to store a program; and a processor used to execute the program and realize each step of the speech recognition method described in any one of the above.
[0023] The readable storage medium stores a computer program, which, when executed by a processor, implements each step of the speech recognition method described in any one of the above.
[0024]
[0009] From the above-mentioned solutions, the speech recognition method, device, equipment, and storage medium according to the present application can extract speech features of a target speaker from the speech features of a target mixed speech based on the speech features of a target mixed speech and the speaker features of the target speaker, thereby obtaining a speech recognition result for the target speaker based on the extracted speech features of the target speaker. Furthermore, in the present application, when extracting speech features of a target speaker from the speech features of a target mixed speech, the extraction direction is directed toward the target speech features (speech features used to obtain speech recognition results that match the actual speech content of the target speaker). Therefore, the extracted speech features are the target speech features or speech features close to the target speech features. Thus, the speech features extracted using the above method are useful features for speech recognition. Performing speech recognition based on the extracted speech features produces favorable effects in speech recognition, i.e., relatively accurate speech recognition results and a good user experience. [Brief explanation of the drawings]
[0025] In order to more clearly explain the technical solutions of the embodiments of the present application or the prior art, the drawings necessary for explaining the embodiments or the prior art will be briefly described below. Obviously, the drawings described below are only embodiments of the present application, and those skilled in the art can obtain other drawings according to the provided drawings without any creative efforts.
[0026] [Figure 1] FIG. 1 is a flowchart of a speech recognition method according to an embodiment of the present application. [Figure 2] FIG. 2 is a flowchart of jointly training a feature extraction model and a speech recognition model according to an embodiment of the present application. [Figure 3] FIG. 3 is a process schematic diagram of jointly training a feature extraction model and a speech recognition model according to an embodiment of the present application. [Figure 4] FIG. 4 is a diagram showing the structure of a speech recognition device according to an embodiment of the present application. [Figure 5] FIG. 5 is a diagram illustrating the structure of a voice recognition device according to an embodiment of the present application. DETAILED DESCRIPTION OF THE INVENTION
[0027] The following clearly and in detail describes the technical solutions in the embodiments of the present application, together with the drawings of the embodiments of the present application. It is clear that the described embodiments are only a part of the embodiments of the present application, but not all of them. Based on the embodiments of the present application, a person skilled in the art can easily obtain other embodiments without creative work, all of which are within the scope of protection of the present application.
[0028] In the external environment, people are surrounded by various sound sources, such as multiple people speaking simultaneously, traffic noise, and natural noise. Thanks to the tireless efforts of researchers, the problem of separating the background noise mentioned above, i.e., the problem of speech enhancement in the usual sense, has been well solved. However, when multiple people speak simultaneously, how to recognize the speech content of a target speaker, i.e., how to recognize the speech content of a target speaker from a mixed speech, is a more difficult problem and merits further research.
[0029] In order to recognize the speech content of a target speaker from a mixed speech, the applicant has conducted research and found that the initial concept includes first training a feature extraction model, then training a speech recognition model; obtaining the target speaker's enrollment speech, and extracting a d-vector as the target speaker's speaker features from the target speaker's enrollment speech; extracting the target speaker's speech features from the target mixed speech based on the speaker features of the target speaker and the speech features of the target mixed speech based on the pre-trained feature extraction model; performing a series of conversion processes on the extracted target speaker's speech features to obtain the target speaker's speech; inputting the target speaker's speech into the pre-trained speech recognition model to perform speech recognition and obtain the target speaker's speech recognition result.
[0030] After conducting research into the above concept, the applicant has found that it has a number of defects, including: First, the voiceprint information contained in the d-vector extracted from the target speaker's enrollment speech is insufficient, which affects the effectiveness of subsequent feature extraction. Second, the feature extraction model and the speech recognition model are trained independently and are completely decoupled, making effective joint optimization difficult. When two independently trained models are cascaded to perform speech recognition, cascading errors exist, which affect the effectiveness of speech recognition. Third, if the features extracted from the front-end feature extraction unit are poor, the effectiveness of speech recognition may be reduced without any remedial measures being taken in the back-end speech recognition unit.
[0031] After further research in light of the above concepts and the deficiencies present in the concepts, the applicant proposes a speech recognition method that completely overcomes the above deficiencies. The speech recognition method can accurately recognize the speech content of a target speaker from a mixed speech. The speech recognition method is applied to a terminal with data processing capabilities, and the terminal can recognize the speech content of a target speaker from a target mixed speech using the speech recognition method of the present application. The terminal may include a processing component, memory, input / output ports, and a power supply component, and may optionally include a multimedia component, an audio component, a sensor component, and a communication component. Here, the processing component is used for data processing, can perform the speech synthesis process of the present application, and may include one or more processors. The processing component may also include one or more modules for interacting with other components.
[0032] The memory is configured to store various types of data and may be implemented by any type of volatile or non-volatile storage device, or combinations thereof, such as one or more combinations of static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disks, optical disks, etc.
[0033] The power supply component supplies power to each component of the device and manages the power management system and and one or more power sources.
[0034] The multimedia component may include a display, preferably a touch display that receives input signals from a user, and may include a front and / or rear camera.
[0035] The audio component is configured to output and / or input audio signals and may include, for example, a microphone configured to receive external audio signals, and a speaker configured to output audio signals and voices synthesized by the terminal.
[0036] The input / output port is a port between a processing component and a peripheral port module, which may include a keyboard, buttons, etc. Here, the buttons may include, but are not limited to, a home page button, volume buttons, start button, lock button, etc.
[0037] The sensor component may include one or more sensors to provide the device with various state assessments, for example, the sensor component may detect whether the device is open or closed, whether a user is touching the device, the device's orientation, speed, temperature, etc. The sensor component may include, but is not limited to, one or more combinations of an image sensor, an acceleration sensor, a gyroscope sensor, a pressure sensor, a temperature sensor, etc.
[0038] The communication component is configured to communicate with the terminal via wired or wireless communication with another device, and the terminal may access a wireless network based on a communication standard such as one or a combination of WiFi, 2G, 3G, 4G, and 5G.
[0039] Alternatively, the terminal may be implemented by one or more application specific integrated circuits (ASICs), digital signal processors (ASPs), digital signal processor devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors or other electronic components to perform the simultaneous interpretation method according to the present application.
[0040] The speech recognition method according to the present application can also be used in a server. The server can recognize the speech content of a target speaker from a target mixed speech using the speech recognition method according to the present application. In one embodiment, the server is connected to a terminal via a network, and the terminal acquires the target mixed speech and transmits the target mixed speech to the server via the network connected to the server. The server recognizes the speech content of the target speaker from the target mixed speech using the speech recognition method according to the present application and transmits the speech content of the target speaker to the terminal via the network. The server can include one or more processors and memory, where the memory is configured to store various types of data and can be implemented by any type of volatile or non-volatile storage device, such as one or more combinations of static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk, optical disk, etc., or a combination thereof. The server can also include one or more power sources, one or more wired network ports and / or one or more wireless network ports, and one or more operating systems.
[0041] Next, the speech recognition method according to the present application will be described with reference to the following examples.
[0042] (First Example) FIG. 1 is a flowchart of a speech recognition method according to an embodiment of the present application, which includes the following steps:
[0043] S101: Acquire speech features of a target mixed speech and speaker features of a target speaker.
[0044] Here, the target mixed speech is speech from multiple speakers, and includes speech from the target speaker as well as speech from other speakers. The present application aims to accurately recognize the speech content of the target speaker in the presence of speech from other speakers.
[0045] Here, the process of acquiring the speech features of the target mixed speech is to acquire a feature vector (such as a spectral feature) of each speech frame from the target mixed speech, obtain a sequence of feature vectors, and use the sequence of feature vectors as the speech features of the target mixed speech. If the target mixed speech contains K speech frames, and the feature vector of the k-th speech frame is x k The speech features of the target speech mixture are expressed as [x1,x2,…,x k ,…,x K ] is expressed as follows.
[0046] There are several embodiments for obtaining speaker features of a target speaker, and this embodiment provides the following two options: In one possible embodiment, the enrollment speech of the target speaker is obtained, a d-vector is extracted from the enrollment speech of the target speaker, and the extracted d-vector is used as the speaker feature of the target speaker; Considering that the voiceprint information contained in the d-vector is simple and not rich, and in order to improve the effect of subsequent feature extraction, this embodiment provides another preferred embodiment, that is, the enrollment speech of the target speaker is obtained, and short-term and long-term voiceprint features are extracted from the enrollment speech of the target speaker to obtain multi-scale voiceprint features, and the multi-scale voiceprint features are used as the speaker feature of the target speaker.
[0047] Compared with the speaker features obtained from the first embodiment, the speaker features obtained from the second embodiment have richer voiceprint information. Therefore, if feature extraction is subsequently performed using the speaker features obtained from the second embodiment, better feature extraction effects can be obtained.
[0048] Next, a specific implementation process of "extracting short-term and long-term voiceprint features from the enrolled voice of a designated speaker" in the second embodiment will be described.
[0049] The process of extracting short-term and long-term voiceprint features from the enrollment speech of the target speaker may include: using a pre-built speaker characterization extraction model to extract short-term and long-term voiceprint features from the enrollment speech of the target speaker. Specifically, obtaining a voice feature sequence of the enrollment speech of the target speaker, inputting the voice feature sequence of the enrollment speech of the target speaker into the pre-built speaker characterization extraction model to obtain the short-term and long-term voiceprint features of the target speaker.
[0050] Alternatively, the speaker feature extraction model can use a convolutional neural network, where the speech feature sequence of the target speaker's enrollment speech is input to the convolutional neural network for feature extraction, resulting in shallow layer features and deep layer features. Here, the shallow layer features are used as short-term voiceprint features because they have a small receptive field and can effectively characterize short-term voiceprints. However, the deep layer features are used as long-term voiceprint features because they have a large receptive field and can effectively characterize long-term voiceprints.
[0051] The speaker feature extraction model in this embodiment is obtained by training using a large amount of training speech (preferably from a single speaker) with actual speaker labels, where the actual speaker labels of the training speech represent the speaker corresponding to the training speech. Optionally, the Cross Entropy (CE) method or Metric Learning (ML) can be used. The rules can be used to train a speaker characterization extraction model.
[0052] S102: The target speech features are extracted from the speech features of the target mixed speech based on the speech features of the target mixed speech and the speaker features of the target speaker, and the extracted speech features of the target speaker are obtained.
[0053] Here, the target speech features are speech features that are used to obtain speech recognition results that match the actual speech content of the target speaker.
[0054] By taking the approach to the target speech features as the extraction direction, it is possible to extract target speech features or speech features similar to the target speech features from the speech features of the target mixed speech; in other words, by taking the approach to the target speech features as the extraction direction, it is possible to extract speech features that are useful for subsequent speech recognition from the speech features of the target mixed speech; and by performing speech recognition based on speech features that are useful for speech recognition, good speech recognition results can be obtained.
[0055] As an option, the process of taking the approach to the target speech features as an extraction direction, extracting speech features of the target speaker from the speech features of the target speech mixture based on the speech features of the target speech mixture and the speaker features of the target speaker, and obtaining extracted speech features of the target speaker can include using a pre-created feature extraction model to extract speech features of a specified speaker from the target speech mixture based on the target speech features and the target speaker features, and obtaining extracted speech features of the target speaker.
[0056] Here, the feature extraction model is obtained by training using speech features of a training speech mixture containing speech of a designated speaker and speaker features of the designated speaker, with the speech recognition results obtained based on the extracted speech features of the designated speaker as the optimization target. It should be noted that in the training stage, the speech features of the training speech mixture and the speaker features of the designated speaker are used as inputs to the feature extraction model, and the speech features of the designated speaker extracted from the training speech features are used as output.
[0057] In one possible embodiment, when training a feature extraction model, a speech recognition result obtained based on extracted speech features of a specified speaker is set as an optimization target. By training a feature extraction model using a speech recognition result obtained based on extracted speech features of a specified speaker as an optimization target, speech features useful for speech recognition can be extracted from mixed speech features based on the feature extraction model.
[0058] In order to improve the effect of feature extraction, in another possible embodiment, when training a feature extraction model, the extracted speech features of a designated speaker and the speech recognition results obtained based on the extracted speech features of the designated speaker are set as optimization targets. Also, by setting the extracted speech features of a designated speaker and the speech recognition results obtained based on the extracted speech features of the designated speaker as optimization targets, speech features that are useful for speech recognition and are close to the standard speech features of the designated speaker can be extracted from the mixed speech features based on the feature extraction model. For the sake of explanation, the standard speech features of the target speaker refer to speech features obtained based on the speech (clean speech) of the designated speaker.
[0059] S103: Obtain a speech recognition result for the target speaker based on the extracted speech features of the target speaker.
[0060] There are several embodiments for obtaining a speech recognition result for a target speaker based on the extracted speech features of the target speaker. In one possible embodiment, the speech recognition result for a target speaker can be obtained based only on the extracted speech features of the target speaker. In order to improve the effect of speech recognition, in another possible embodiment, the speech recognition result for a target speaker can be obtained based on the extracted speech features of the target speaker and the enrolled speech features of the target speaker (the enrolled speech features of the target speaker refer to the speech features of the enrolled speech of the target speaker). The speech recognition result of the speaker can be obtained, and the registered speech features of the target speaker can be used as auxiliary information to improve the effectiveness of speech recognition.
[0061] Specifically, a pre-made speech recognition model can be used to obtain the speech recognition result of the target speaker. More specifically, the extracted speech features of the target speaker are used as the speech recognition input features, or the extracted speech features of the target speaker and the registered speech features of the target speaker are used as the speech recognition input features, and the speech recognition input features are input into the pre-made speech recognition model to obtain the speech recognition result of the target speaker.
[0062] It should be noted that when the extracted speech features of the target speaker and the registered speech features of the target speaker are input into a speech recognition model as speech recognition input features, the registered speech features of the target speaker can assist the speech recognition model in speech recognition if the extracted speech features of the target speaker are inaccurate, thereby improving the effectiveness of speech recognition.
[0063] Preferably, the speech recognition model is obtained by joint training with the feature extraction model, and the speech recognition model is obtained by training the above-mentioned "extracted speech features of a specified speaker" as training samples, and by using speech recognition results obtained based on the extracted speech features of the specified speaker as optimization targets. The feature extraction model is jointly trained with the speech recognition model so that the feature extraction model can be optimized in a direction advantageous for speech recognition.
[0064] The speech recognition method according to the embodiment of the present application can extract speech features of a target speaker from speech features of a target mixed speech, and obtain speech recognition results for the target speaker based on the extracted speech features of the target speaker. Furthermore, in the embodiment of the present application, when extracting speech features of the target speaker from speech features of a target mixed speech, the extraction direction is directed toward the target speech features (speech features used to obtain speech recognition results that match the actual speech content of the target speaker). Therefore, the extracted speech features are the target speech features or speech features close to the target speech features. Performing speech recognition based on these speech features produces favorable effects in speech recognition, i.e., relatively accurate speech recognition results and a good user experience.
[0065] (Second Example) In the above-described embodiment, the feature extraction model used to extract speech features of a target speaker from speech features of a target mixed speech, and the speech recognition model used to obtain speech recognition results for the target speaker based on the features extracted by the feature extraction model, are trained by a joint training method. This embodiment will mainly describe the joint training process of the feature extraction model and the speech recognition model.
[0066] Next, based on FIG. 2 and referring to FIG. 3, the co-training process of the feature extraction model and the speech recognition model will be described. The co-training process of the feature extraction model and the speech recognition model can include the following.
[0067] S201: Obtain the training mixed speech s m from the pre-made training dataset S.
[0068] Here, the training dataset S includes a plurality of training data, each of which includes the speech (clean speech) of a designated speaker, and also includes the training mixed speech of the speech of the designated speaker. Here, there is annotation text (the annotation text is the content of the speech of the designated speaker) for the speech of the designated speaker.
[0069] The construction process of the training dataset S is
[0070] Step a1: Obtain a plurality of speeches of a plurality of speakers.
[0071] Each of the plurality of speeches obtained in this step is the speech of one speaker, and each speech has annotation text. Assuming the content of the speech of one speaker is "The weather is nice today", the annotation text of the speech is " <s> , now, filling, heaven, qi, non, illusion,< / s> ". Here, " <s> " is the beginning symbol of a sentence, and "< / s> " is the end symbol of the sentence. <Step a2: One of some or all of the multiple voices is taken as the voice of a designated speaker, and one or more voices of other speakers from among the other voices are mixed with the voice of the designated speaker to obtain a training mixed voice, and the training mixed voice and the voice of the designated speaker are taken as one training data.
[0074] For example, the acquired multiple speeches include one speech from speaker a, one speech from speaker b, one speech from speaker c, and one speech from speaker d, where each speech is a clean speech from one speaker. The speech from speaker a is designated as the speech of the specified speaker, and speeches from another speaker (one or more speakers) are mixed with the speech from speaker a to obtain a training mixed speech. For example, the speech from speaker b is mixed with the speech from speaker a, or the speeches from speakers b and c are mixed with the speech from speaker a, and the training mixed speech obtained by mixing the speech from speaker a with the speech from another speaker is considered to be one training data. Similarly, the speech from speaker b is designated as the speech of the specified speaker, and speeches from another speaker (one or more speakers) are mixed with the speech from speaker b to obtain a training mixed speech. Furthermore, the training mixed speech obtained by mixing the speech from speaker b with the speech from another speaker is considered to be one training data. Multiple training data are obtained in this manner.
[0075] It should be noted that when mixing a designated speaker's voice with another speaker's voice, if the length of the other speaker's voice is different from that of the designated speaker, the other speaker's voice must be processed to have the same length as that of the designated speaker. If the designated speaker's voice contains K voice frames (i.e., the length of the designated speaker's voice is K), and the length of the other speaker's voice is greater than K, the K+1th voice frame and the following voice frames can be deleted from the other speaker's voice, i.e., the previous K voice frames are kept, and if the length of the other speaker's voice is less than K and is L, then the previous KL voice frames are copied and supplemented.
[0076] Step a3: All the obtained training data constitute a training dataset.
[0077] S202: Training mixed speech s m The speech features of the training mixture speech features X m and acquires the speaker features of the designated speaker as training speaker features.
[0078] As in the first embodiment, a speaker feature extraction model is first created, and then the previously created speaker feature extraction model is used to extract speaker features from the enrolled speech of a specified speaker, and the extracted speaker features are set as training speaker features. As shown in Figure 3, a speaker feature extraction model 300 is used to extract short-term and long-term voiceprint features from the enrolled speech of a specified speaker, and the extracted short-term and long-term voiceprint features are set as the speaker features of the specified speaker.
[0079] It should be noted that the speaker feature extraction model is pre-trained before the joint training of the feature extraction model and the speech recognition model, and the feature extraction model and the speech recognition model In the joint training stage, its parameters are invariant and are not updated by the parameters of the feature extraction model and the speech recognition model.
[0080] S203: Using the feature extraction model, train the mixed speech feature X m and based on the training speaker features, the training mixed speech features X m Extract the speech features of the specified speaker from
number
[0081] Specifically, first, the training mixture of speech features X m and the training speaker features are input into a feature extraction model to obtain a feature mask M corresponding to the designated speaker, and then the training mixed speech features X are extracted based on the feature mask M corresponding to the designated speaker. m Extract the speech features of the specified speaker from
number
[0082] As shown in Figure 3, the training mixture speech features X m and training speaker features are input to a feature extraction model 301, which extracts the input training mixed speech features X m and determining and outputting a feature mask M corresponding to the specified speaker based on the training speaker features.
[0083] Examples of the feature extraction model 301 in this embodiment include a recurrent neural network (RNN), a convolution neural network (CNN), and a deep neural network (DNN).
[0084] For clarity, the training speech mixture feature Xm is a feature vector sequence [x m1 ,x m2 ,…,x mk ,…,x mK ] (K is the total number of speech frames in the training mixture speech). When the training mixture speech features Xm and the training speaker features are input to the feature extraction model 301, the training speaker features can be combined with the feature vectors of each speech frame in the training mixture speech, and then input to the feature extraction model 301. For example, if the feature vector of each speech frame in the training mixture speech is 40-dimensional and the short-term and long-term voiceprint features of the training speaker features are both 40-dimensional, then after combining the short-term and long-term voiceprint features with the feature vector of each speech frame in the training mixture speech, a combined feature vector of 120 dimensions is obtained. When extracting the speech features of a designated speaker, adding rich input information to the short-term and long-term voiceprint features allows the feature extraction model to effectively extract the speech features of the designated speaker.
[0085] In this embodiment, the feature mask M corresponding to a specified speaker is generated by the training mixture speech features X m The proportion of speech features of a specified speaker in the training speech mixture X can be characterized. m [x m1 ,x m2 ,…,x mk ,…,x mK ], and the feature mask M corresponding to a specified speaker is denoted as [m1,m2,……,m k ,……,m K ], m1 is x m1 represents the proportion of the speech features of a given speaker in x m2 By this analogy, m K x mK represents the proportion of the speech features of a specified speaker in K is a value in [0,1]. After obtaining the feature mask M corresponding to the specified speaker, we use the training mixture speech features X m It is specified as By multiplying the feature mask M corresponding to the selected speaker for each frame, the training mixture speech feature X m Speech features of a specified speaker extracted from
number
[0086] S204: Extract speech features of designated speaker
number
number
[0087] Preferably, in order to improve the recognition effect of the speech recognition model, the enrollment speech features of the designated speaker (the enrollment speech features of the designated speaker mean the speech features of the enrollment speech of the designated speaker) Xe=[x e1 ,x e2 ,……,x ek ,……,x eK ] and extract the speech features of the specified speaker.
number
[0088] Optionally, the speech recognition model of this embodiment may include an encoder module, an attention module, and a decoder module, where the encoder module extracts speech features of a designated speaker.
number
number
number
number
[0089] The input of the encoder module is the extracted speech features of a specified speaker.
number
number
number
number
[0090] The attention module extracts speech features for each specified speaker.
number
[0091] The decoding module is used to decode the audio-related feature vector extracted from the attention module and obtain the recognition result of the decoding time.
[0092] As shown in FIG. 3, the attention module 3023 performs the following steps at each decoding time based on the attention mechanism:
number
[0093] For the sake of explanation, the attention mechanism refers to taking a vector as a query term (query), performing the attention mechanism operation on a series of feature vector sequences, and outputting the feature vector that best matches the query term. Specifically, it calculates the matching coefficient between the query term and each feature vector in the feature vector sequence, and then multiplies the corresponding feature vector by these matching coefficients to calculate the sum, and the resulting new feature vector is the feature vector that best matches the query term.
[0094] At the t-th decoding time, the attention module 3023 receives the state feature vector d t Let be the query item, and d t and H x =[h1 x ,h2 x ,……,h K x ] with each feature vector, and the matching coefficient w1 x , w2 x , ……, w K x Calculate the matching coefficient w1 x , w2 x , ……, w K x H x =[h1 x ,h2 x ,……,h K x ], and sum the resulting feature vector to the audio-related feature vector c t x Similarly, the attention module 3023 t and H e =[h1 e ,h2 e ,……,h K e ] Matching coefficient w1 with each feature vector in e , w2 e , ……, w K eCalculate the matching coefficient w1 e , w2 e , ……, w K e H e =[h1 e ,h2 e ,……,h K e ], and sum the resulting feature vector to the audio-related feature vector c t e Let the audio-related feature vector c t x and c t e After obtaining the audio-related feature vector c t x and c t e is input to the decoder module 3024 for decoding, and the recognition result at the t-th decoding time is obtained.
[0095] Here, the state feature vector d of the decoder module 3024 t is the recognition result y at the t-1th decoding time. t-1 and c output from the attention module t-1 x and c t-1 e Optionally, the decoder module 3024 may include multiple neural network layers, for example, in the case of two unidirectional long-short-term memory layers, at the t-th decoding time, the first long-short-term memory layer of the decoder module 3024 determines the recognition result y at the t-1-th decoding time. t-1 and c output from the attention module 3023 t-1 x and c t-1 e is used as input, and the decoder state feature vector d t d t is input to the attention module 3023, and the c t x and c t e Then, c tx and c t e and the combined vector is fed as the input of the second long-short-term memory layer of the decoder module 3024 (e.g., c t x and c t e are both 128-dimensional vectors, and c t x and c t e By combining these, a 256-dimensional combined vector is obtained. The 256-dimensional combined vector is input to the second long-short-term memory layer of the decoder module 3024, and the decoder output h t d Finally, h t d The posterior probability of the output character is calculated from the posterior probability of the output character, and the recognition result at the t-th decoding time is determined based on the posterior probability of the output character.
[0096] S205: Extract speech features of designated speaker
number
number
number
[0097] Specifically, the implementation process of S205 may include:
[0098] S2051: Training mixed speech s m The specified speaker voice s corresponding to t Annotated text T of (specified speaker's speech) t and obtain the specified speaker voice st The speech features of the specified speaker are converted into standard speech features X t Obtain as.
[0099] For points to be explained, please refer to the designated speaker voices here. t The enrollment voice of the designated speaker is the voice of each designated speaker.
[0100] S2052: Extract speech features for a specified speaker
number
number
[0101] Optionally, extract speech features for a specified speaker
number
number
[0102] S2053: Update the parameters of the feature extraction model based on the first prediction loss Loss1 and the second prediction loss Loss2, and update the parameters of the speech recognition model based on the second prediction loss Loss2.
[0103] By updating the parameters of the feature extraction model based on the first prediction loss Loss1 and the second prediction loss Loss2, the feature extraction model can extract speech features that are close to the standard speech features of a specified speaker and are useful for speech recognition from the training mixed speech features.By inputting these speech features into the speech recognition model and performing speech recognition, good speech recognition results can be achieved.
[0104] (Third Example) In this embodiment, based on the third embodiment, the process in the first embodiment will be described, which is "using a pre-created feature extraction model, based on the target mixed speech features and target speaker features, extracting speech features of a specified speaker from the target mixed speech features, and obtaining extracted speech features of the target speaker."
[0105] The process of using a pre-made feature extraction model, extracting speech features of a specified speaker from the target mixture speech features based on the target mixture speech features and the target speaker features, and obtaining extracted speech features of the target speaker may include the following steps:
[0106] Step b1: The speech features of the target mixed speech and the speaker features of the target speaker are input into a feature extraction model to obtain a feature mask corresponding to the target speaker.
[0107] Here, the feature mask corresponding to the target speaker can characterize the proportion of the speech features of the target speaker in the speech features of the target speech mixture.
[0108] Step b2: Extract speech features of the target speaker from speech features of the target mixed speech based on the feature mask corresponding to the target speaker, and obtain extracted speech features of the target speaker.
[0109] Specifically, the speech features of the target mixed speech are multiplied by a feature mask corresponding to the target speaker for each frame to obtain the extracted speech features of the target speaker.
[0110] After the extracted speech features of the target speaker are obtained, the extracted speech features of the target speaker and the enrolled speech features of the target speaker are input into a speech recognition model to obtain a speech recognition result for the target speaker. Specifically, the process of inputting the extracted speech features of the target speaker and the enrolled speech features of the target speaker into a speech recognition model to obtain a speech recognition result for the target speaker may include the following steps:
[0111] Step c1: Based on the encoder module of the speech recognition model, the extracted speech features of the target speaker and the registered speech features of the target speaker are encoded respectively to obtain two encoding results.
[0112] Step c2: Based on the attention module of the speech recognition model, extract audio-related feature vectors required for decoding at the decoding time from the two encoding results respectively.
[0113] Step c3: According to the decoder module of the speech recognition model, the audio-related feature vectors extracted from the two encoding results are decoded to obtain the recognition result at the decoding time.
[0114] It should be noted that the process of inputting the extracted speech features of a target speaker into a speech recognition model and obtaining the speech recognition result of the target speaker is the same as the process of inputting the extracted speech features of a designated speaker and the registered speech features of a designated speaker into a speech recognition model in a training stage and obtaining the speech recognition result of the designated speaker. Similar to the realization process of obtaining the recognition result, the specific realization process of steps c1 to c3 can be referred to the description of the encoder module, attention module, and decoder module in the second embodiment, and therefore will be omitted in this embodiment.
[0115] As can be seen from the above first to third embodiments, the speech recognition method according to the present application has the following advantages. First, in the present application, multi-scale voiceprint features are extracted from the enrolled speech of the target speaker and input into the feature extraction model, thereby increasing the richness of information input to the feature extraction model and improving the feature extraction effect of the feature extraction model. Second, the feature extraction model and the speech recognition model are jointly trained, and the prediction loss of the speech recognition model is applied to the feature extraction model, allowing the feature extraction model to extract speech features useful for speech recognition and improving the accuracy of speech recognition results. Third, the speech features of the enrolled speech of the target speaker are used as additional input to the speech recognition model, and when the speech features extracted by the feature extraction model are poor, the speech recognition model is assisted in speech recognition, thereby achieving relatively accurate speech recognition results. From the above, the speech recognition method according to the present application can accurately recognize the speech content of the target speaker even when there is interference from complex human voices.
[0116] (Fourth Example) In addition, an embodiment of the present application provides a voice recognition device, and the voice recognition device according to the embodiment of the present application is described as follows. The voice recognition device described below may be cross-referenced with the voice recognition method described above.
[0117] FIG. 4 is a diagram showing the structure of a speech recognition device according to an embodiment of the present application; a feature acquisition module 401 used to acquire speech features of the target mixed speech and speaker features of the target speaker; a feature extraction module 402 that extracts the target speaker's voice features from the target mixed voice features based on the target mixed voice features and the speaker features of the target speaker, and obtains the extracted target speaker voice features; a speech recognition module 403 used to obtain a speech recognition result of the target speaker based on the extracted speech features of the target speaker; Here, the target speech features are speech features used to obtain speech recognition results that match the actual speech content of the target speaker.
[0118] Optionally, the feature acquisition module 401: a speech feature acquisition module used to acquire speech features of the target mixed speech; and a speaker feature acquisition module used to acquire speaker features of the target speaker.
[0119] As an option, when acquiring the speaker features of the target speaker, the speaker feature acquisition module is specifically used to acquire the enrollment speech of the target speaker, extract short-term voiceprint features and long-term voiceprint features from the enrollment speech of the target speaker, and use the obtained multi-scale voiceprint features as the speaker features of the target speaker.
[0120] As an option, the feature extraction module 402 is specifically used to use a pre-made feature extraction model to extract the voice features of the target speaker from the voice features of the target mixed voice based on the voice features of the target mixed voice and the speaker features of the target speaker.
[0121] Here, the feature extraction model is obtained by training using speech features of a training speech mixture including speech of a designated speaker and speaker features of the designated speaker, with speech recognition results obtained based on the extracted speech features of the designated speaker as an optimization target, and the extracted speech features of the designated speaker are the speech features of the designated speaker extracted from the speech features of the training speech mixture.
[0122] Optionally, the feature extraction model is obtained by training the extracted speech features of the designated speaker and the speech recognition results obtained based on the extracted speech features of the designated speaker as optimization targets.
[0123] Optionally, the feature extraction module 402: a feature mask determination submodule, which is used to input the speech features of the target speech mixture and the speaker features of the target speaker into the feature extraction model to obtain a feature mask corresponding to the target speaker; a speech feature extraction sub-module, which is used to extract speech features of the target speaker from the speech features of the target speech mixture based on the speech features of the target speech mixture and a feature mask corresponding to the target speaker; Here, the feature mask can characterize the proportion of the speech features of the corresponding speaker in the speech features of the target speech mixture.
[0124] Optionally, the speech recognition module 403 is specifically used to obtain a speech recognition result of the target speaker based on the extracted speech features of the target speaker and the enrolled speech features of the target speaker, where the enrolled speech features of the target speaker are the speech features of the enrolled speech of the target speaker.
[0125] As an option, the speech recognition module 403 is specifically used to input speech recognition input features including at least the extracted speech features of the target speaker into a pre-made speech recognition model to obtain a speech recognition result of the target speaker.
[0126] Here, the speech recognition model is obtained by joint training with the feature extraction model, and the speech recognition model is obtained by training using the extracted speech features of the specified speaker and using the speech recognition results obtained based on the extracted speech features of the specified speaker as an optimization target.
[0127] As an option, the speech recognition module 403 inputs speech recognition input features including at least the extracted speech features of the target speaker into a pre-created speech recognition model to obtain a speech recognition result for the target speaker, specifically by encoding the speech recognition input features based on an encoder module of the speech recognition model to obtain an encoding result, extracting audio-related feature vectors required for decoding at the decoding time from the encoding result based on an attention module of the speech recognition model, and decoding the audio-related feature vectors extracted from the encoding result based on a decoder module of the speech recognition model to obtain a recognition result at the decoding time.
[0128] Optionally, the speech recognition device according to the embodiment of the present application may further include a model training module, which may include an extracted speech feature obtaining module, a speech recognition result obtaining module, and a parameter updating module.
[0129] The extracted speech feature acquisition module is used to extract speech features of the designated speaker from speech features of the training speech mixture using a feature extraction model to obtain extracted speech features of the designated speaker.
[0130] The speech recognition result obtaining module is used to obtain a speech recognition result for the designated speaker using a speech recognition model and the extracted speech features of the designated speaker.
[0131] The model update module updates parameters of a feature extraction model based on the extracted speech features of the designated speaker and the speech recognition result of the designated speaker, and It is used to update the parameters of the speech recognition model based on the speech recognition results of the user.
[0132] Optionally, the model updating module may include an annotation text acquisition module, a standard audio feature acquisition module, a prediction loss determination module, and a parameter updating module.
[0133] The training speech mixtures correspond to the speech of the designated speaker.
[0134] The standard voice feature acquisition module is used for acquiring voice features of the voice of the designated speaker as the standard voice features of the designated speaker.
[0135] The annotation text acquisition module is used to acquire annotation text of the designated speaker's speech.
[0136] The predicted loss determination module is used to determine a first predicted loss based on the extracted speech features of the designated speaker and the standard speech features of the designated speaker, and to determine a second predicted loss based on the speech recognition result of the designated speaker and the annotated text of the speech of the designated speaker.
[0137] The parameter updating module is used to update parameters of a feature extraction model based on the first prediction loss and the second prediction loss, and to update parameters of a speech recognition model based on the second prediction loss.
[0138] Optionally, the training mixture speech and the designated speaker's speech corresponding to the training mixture speech are obtained from a pre-created training dataset. The speech recognition device according to the embodiment of the present application may further include a training dataset construction module.
[0139] The training dataset construction module is used to acquire a plurality of speeches from a plurality of speakers, each consisting of a single speaker's speech accompanied by annotated text; select one of some or all of the plurality of speeches as the speech of a designated speaker; mix one or more speeches of other speakers from among the other speeches with the speech of the designated speaker to obtain a training speech mixture; and use the training speech mixture obtained by mixing with the speech of the designated speaker as a piece of training data; and construct the training dataset from all of the acquired training data.
[0140] The speech recognition device according to the embodiment of the present application can extract speech features of a target speaker from speech features of a target mixed speech, and obtain speech recognition results for the target speaker based on the extracted speech features of the target speaker. Furthermore, in the embodiment of the present application, when extracting speech features of the target speaker from speech features of a target mixed speech, the extraction direction is directed toward the target speech features (speech features used to obtain speech recognition results that match the actual speech content of the target speaker). Therefore, the extracted speech features are the target speech features or speech features close to the target speech features. Performing speech recognition based on these speech features produces favorable effects in speech recognition, i.e., relatively accurate speech recognition results and a good user experience.
[0141] (Fifth Example) An embodiment of the present application also provides a speech recognition device. Figure 5 shows a structural diagram of a speech recognition device. The speech recognition device may include at least one processor 501, at least one communication port 502, at least one memory 503, and at least one communication bus 504.
[0142] In the embodiment of the present application, the number of the processor 501, the communication port 502, the memory 503, and the communication bus 504 is at least one, and the processor 501, the communication port 502, and the memory 503 are connected to the communication bus. They communicate with each other via the service 504.
[0143] Processor 501 may be a single central processor CPU, or an Application Specific Integrated Circuit ASIC (Application Specific Integrated Circuit), or one or more integrated circuits configured to implement an embodiment of the present invention.
[0144] The memory 503 may include a high-speed RAM memory, and may also include non-volatile memory, such as at least one disk memory.
[0145] Among these, the memory stores a program, and the processor can call the program stored in the memory. The program is used to acquire speech features of a target mixed speech and speaker features of a target speaker, and to extract an approach to the target speech features as speech features used to acquire a speech recognition result that matches the actual speech content of the target speaker, extracting speech features of the target speaker from the speech features of the target mixed speech based on the speech features of the target mixed speech and the speaker features of the target speaker to obtain extracted speech features of the target speaker, and to acquire a speech recognition result of the target speaker based on the extracted speech features of the target speaker.
[0146] Alternatively, the subdivision and extension functions of the program may refer to the above description.
[0147] (Sixth Example) An embodiment of the present application also provides a readable storage medium, which can store a program adapted to be executed by a processor, for acquiring speech features of a target mixed speech and speaker features of a target speaker, extracting speech features of the target speaker from the speech features of the target mixed speech based on the speech features of the target mixed speech and the speaker features of the target speaker, and obtaining extracted speech features of the target speaker based on the extracted speech features of the target speaker, with the target speaker being the extraction-oriented approach to the target speech features as speech features used to obtain a speech recognition result of the actual speech content match of the target speaker.
[0148] Alternatively, the subdivision and extension functions of the program may refer to the above description.
[0149] Finally, it should be clarified that, in this specification, related terms such as "first" and "second" are used to distinguish one entity or operation from another, and do not necessarily require or imply any actual relationship or ordering between those entities or operations. Furthermore, the terms "comprise," "include," or other variations thereof are intended to cover a non-exclusive inclusion, such that a process, method, article, or facility that includes a set of elements includes not only those elements but also other elements not expressly listed or inherent in such process, method, article, or facility. Absent further limitations, an element qualified by the phrase "comprises ..." does not exclude the presence of other identical elements in the process, method, article, or facility that includes said element.
[0150] Each embodiment in this specification is described in a step-by-step manner, with each embodiment being described with an emphasis on the differences from other embodiments, and the same and similar parts between the embodiments may be referred to each other.
[0151] The above description of the disclosed embodiments will enable one skilled in the art to make or use the present invention. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein may be incorporated in other embodiments without departing from the spirit or scope of the present invention. Thus, the present invention is not intended to be limited to the embodiments shown herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. 1. A speech recognition method, comprising: Acquiring speech characteristics of a target speech mixture and speaker characteristics of a target speaker; an extraction direction is to approach target speech features as speech features used to obtain speech recognition results that match the actual speech content of the target speaker, and extract speech features of the target speaker from the speech features of the target mixed speech based on the speech features of the target mixed speech and the speaker features of the target speaker, thereby obtaining extracted speech features of the target speaker; obtaining a speech recognition result for the target speaker based on the extracted speech features of the target speaker; Including, The acquiring of the speech characteristics of the target speech mixture includes: obtaining a feature vector for each audio frame from the target mixed audio to obtain a sequence of feature vectors, and defining the obtained sequence of feature vectors as audio features of the target mixed audio; Acquiring speaker characteristics of the target speaker includes: acquiring an enrollment speech of the target speaker; inputting the speech feature sequence of the enrollment speech of the target speaker into a convolutional neural network, performing feature extraction, obtaining shallow layer features and deep layer features, using the shallow layer features as short-term voiceprint features and the deep layer features as long-term voiceprint features to obtain multi-scale voiceprint features, and using the obtained multi-scale voiceprint features as speaker features of the target speaker; A speech recognition method comprising:
2. Extracting the speech features of the target speaker from the speech features of the target mixed speech based on the speech features of the target mixed speech and the speaker features of the target speaker, with the approach to the target speech features as the extraction direction, extracting speech features of the target speaker from the speech features of the target speech mixture based on speech features of the target speech mixture and speaker features of the target speaker using a pre-created feature extraction model; wherein the feature extraction model is obtained by training using speech features of a training speech mixture including speech of a designated speaker and speaker features of the designated speaker, with a speech recognition result obtained based on the extracted speech features of the designated speaker as an optimization target, and the extracted speech features of the designated speaker are speech features of the designated speaker extracted from the speech features of the training speech mixture.
2. The speech recognition method according to claim 1.
3. 3. The speech recognition method according to claim 2, wherein the feature extraction model is obtained by training the extracted speech features of the specified speaker and the speech recognition results obtained based on the extracted speech features of the specified speaker as optimization targets.
4. Extracting speech features of the target speaker from the speech features of the target speech mixture using the pre-created feature extraction model based on speech features of the target speech mixture and speaker features of the target speaker, inputting the speech features of the target speech mixture and the speaker features of the target speaker into the feature extraction model to obtain a feature mask corresponding to the target speaker; extracting speech features of the target speaker from the speech features of the target speech mixture based on the speech features of the target speech mixture and a feature mask corresponding to the target speaker; 3. The speech recognition method according to claim 2, further comprising:
5. obtaining a speech recognition result for the target speaker based on the extracted speech features of the target speaker, obtaining a speech recognition result for the target speaker based on the extracted speech features of the target speaker and the enrolled speech features of the target speaker; 2. The speech recognition method according to claim 1, wherein the enrollment speech features of the target speaker are speech features of the enrollment speech of the target speaker.
6. obtaining a speech recognition result for the target speaker based on the extracted speech features of the target speaker, inputting speech recognition input features including at least the extracted speech features of the target speaker into a pre-created speech recognition model to obtain a speech recognition result for the target speaker; the speech recognition model is obtained by joint training with the feature extraction model, and the speech recognition model is obtained by training using the extracted speech features of the designated speaker and using speech recognition results obtained based on the extracted speech features of the designated speaker as an optimization target; 3. The speech recognition method according to claim 2, further comprising:
7. inputting the speech recognition input features into the speech recognition model to obtain a speech recognition result for the target speaker, Encoding the speech recognition input features based on an encoder module of the speech recognition model to obtain an encoding result; extracting an audio-related feature vector required for decoding at a decoding time from the encoding result based on an attention module of the speech recognition model; Decoding the audio-related feature vector extracted from the encoding result based on a decoder module of the speech recognition model to obtain a recognition result at the decoding time; 7. The speech recognition method according to claim 6, further comprising:
8. The process of jointly training the speech recognition model and the feature extraction model includes: extracting speech features of the designated speaker from speech features of the training speech mixture using a feature extraction model to obtain extracted speech features of the designated speaker; obtaining a speech recognition result for the designated speaker using a speech recognition model and the extracted speech features of the designated speaker; and updating parameters of a feature extraction model based on the extracted speech features of the designated speaker and the speech recognition result of the designated speaker, and updating parameters of a speech recognition model based on the speech recognition result of the designated speaker; 7. The speech recognition method according to claim 6, further comprising:
9. the training speech mixtures correspond to speech of the designated speakers; updating parameters of a feature extraction model based on the extracted speech features of the designated speaker and the speech recognition result of the designated speaker, and updating parameters of a speech recognition model based on the speech recognition result of the designated speaker; Obtaining annotated text of the speech of the designated speaker, and obtaining speech features of the speech of the designated speaker as standard speech features of the designated speaker; A first predicted loss is determined based on the extracted speech features of the designated speaker and the standard speech features of the designated speaker, and a speech recognition result of the designated speaker and a speech recognition result of the designated speaker are calculated. determining a second prediction loss based on the annotated text of the speech; and updating parameters of a feature extraction model based on the first prediction loss and the second prediction loss, and updating parameters of a speech recognition model based on the second prediction loss; 9. The speech recognition method according to claim 8, further comprising:
10. the training mixtures and the designated speaker's speech corresponding to the training mixtures are obtained from a pre-created training data set; The training dataset construction process includes: obtaining a plurality of speeches from a plurality of speakers, each of which comprises a single speaker's speech annotated with text; one of the plurality of speeches is a speech of a designated speaker, and one or more speeches of other speakers among the plurality of speeches are mixed with the speech of the designated speaker to obtain a training speech mixture, and the training speech mixture obtained by mixing the speech of the designated speaker is used as training data; all the obtained training data constitutes said training data set; 10. The speech recognition method according to claim 9, further comprising:
11. A speech recognition device, a feature acquisition module used to acquire speech features of the target speech mixture and speaker features of the target speaker; a feature extraction module that extracts target speech features from the speech features of the target speech mixture based on the speech features of the target speech mixture and the speaker features of the target speaker, and obtains extracted speech features of the target speaker; a speech recognition module used to obtain a speech recognition result for the target speaker based on the extracted speech features of the target speaker; Including, wherein the target speech features are speech features used to obtain speech recognition results that match the actual speech content of the target speaker, The feature acquisition module: a speech feature acquisition module that acquires a feature vector of each speech frame from the target mixed speech, obtains a sequence of feature vectors, and defines the obtained sequence of feature vectors as speech features of the target mixed speech; a speaker feature acquisition module for acquiring an enrollment speech of the target speaker, inputting the speech feature sequence of the enrollment speech of the target speaker into a convolutional neural network to perform feature extraction, obtaining shallow layer features and deep layer features, using the shallow layer features as short-term voiceprint features and the deep layer features as long-term voiceprint features to obtain multi-scale voiceprint features, and using the obtained multi-scale voiceprint features as speaker features of the target speaker. A speech recognition device characterized by:
12. The feature extraction module is specifically used to extract the speech features of the target speaker from the speech features of the target speech mixture based on the speech features of the target speech mixture and the speaker features of the target speaker, using a pre-created feature extraction model; wherein the feature extraction model is obtained by training using speech features of a training speech mixture including speech of a designated speaker and speaker features of the designated speaker, with the speech recognition results obtained based on the extracted speech features of the designated speaker as an optimization target, and the extracted speech features of the designated speaker are the speech features of the designated speaker extracted from the speech features of the training speech mixture.
12. The speech recognition device according to claim 11.
13. The speech recognition module is specifically used to obtain a speech recognition result of the target speaker according to the extracted speech features of the target speaker and the enrolled speech features of the target speaker; 12. The speech recognition device according to claim 11, wherein the enrollment speech features of the target speaker are speech features of the enrollment speech of the target speaker.
14. the speech recognition module is used to input speech recognition input features including at least the extracted speech features of the target speaker into a pre-created speech recognition model to obtain a speech recognition result for the target speaker; wherein the speech recognition model is obtained by joint training with the feature extraction model, and the speech recognition model is obtained by training using the extracted speech features of the designated speaker and using the speech recognition result obtained based on the extracted speech features of the designated speaker as an optimization target.
13. The speech recognition device according to claim 12.
15. A voice recognition device, a memory used to store a program; a processor used to execute the program and realize each step of the speech recognition method according to any one of claims 1 to 10; A speech recognition facility comprising:
16. A computer-readable storage medium storing a computer program, When the computer program is executed by a processor, it implements each step of the speech recognition method according to any one of claims 1 to 10. A computer-readable storage medium comprising:
Citation Information
Patent Citations
Speech recognition method and related equipment
CN111145736A
Speech recognition method, device and apparatus and storage medium
CN111583916A
Speech signal processing device, speech signal processing method, speech signal process program, learning device, learning method, and learning program
JP2021039219A