Training Method, Device, Equipment and Storage Medium for Streaming Speech Recognition Model
By introducing potential feature extraction and context feature extraction models into the streaming speech recognition model, combining mask processing and loss function, the problem of low recognition accuracy of the streaming speech recognition model without relying on labeled text data is solved, and a more efficient speech recognition effect is achieved.
Patent Information
- Application Number
- CN202210449547.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-26
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2042-04-26
AI Technical Summary
The existing streaming speech recognition model has the problem of low recognition accuracy during training, especially when it does not rely on labeled text data.
The potential feature extraction model and the context feature extraction model are used to extract features by obtaining the prefixed speech signal at the current moment, and adjusting the model parameters in combination with mask processing and loss function to realize the pre-training process.
It improves the accuracy and response speed of speech recognition, reduces dependence on labeled data, and enhances the generalization ability of the model.
Smart Images

Figure CN114898742B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technologies, specifically to artificial intelligence fields such as speech recognition and deep learning, and particularly to a method, apparatus, device, and storage medium for training a streaming speech recognition model. Background Art
[0002] Speech recognition refers to converting speech into text. Speech recognition can be divided into streaming speech recognition and non-streaming speech recognition. Non-streaming speech recognition waits for the entire speech input and then performs speech recognition, outputting the text corresponding to the entire speech input at once. Streaming speech recognition performs speech recognition on the input speech in real time and outputs the speech recognition results in real time. Summary of the Invention
[0003] The present disclosure provides a method, apparatus, device, and storage medium for training a streaming speech recognition model.
[0004] According to one aspect of the present disclosure, there is provided a method for training a streaming speech recognition model, where the streaming speech recognition model includes a latent feature extraction model and a context feature extraction model, and the method includes: based on the entire speech sample, obtaining the prefix speech signal at the current moment, where the prefix speech signal at the current moment includes: the speech signal before the current moment in the entire speech sample; using the latent feature extraction model to perform feature extraction processing on the input prefix speech signal at the current moment to output latent features; performing masking processing on the latent features to obtain masked latent features; using the context feature extraction model to perform feature extraction processing on the input masked latent features to output context features; constructing a loss function based on the context features; and adjusting the model parameters of the latent feature extraction model and the model parameters of the context feature extraction model based on the loss function.
[0005] According to another aspect of the present disclosure, there is provided a training device for a streaming speech recognition model. The streaming speech recognition model includes a latent feature extraction model and a context feature extraction model. The device includes: an acquisition module configured to acquire a prefix speech signal at a current moment based on an entire speech sample, where the prefix speech signal at the current moment includes: the speech signal before the current moment in the entire speech sample; a first feature extraction module configured to perform feature extraction processing on the input prefix speech signal at the current moment by using the latent feature extraction model to output latent features; a masking processing module configured to perform masking processing on the latent features to obtain masked latent features; a second feature extraction module configured to perform feature extraction processing on the input masked latent features by using the context feature extraction model to output context features; a construction module configured to construct a loss function based on the context features; and an adjustment module configured to adjust the model parameters of the latent feature extraction model and the model parameters of the context feature extraction model based on the loss function.
[0006] According to another aspect of the present disclosure, there is provided an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the method according to any one of the above aspects.
[0007] According to another aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute the method according to any one of the above aspects.
[0008] According to another aspect of the present disclosure, there is provided a computer program product, including a computer program, where the computer program implements the method according to any one of the above aspects when executed by a processor.
[0009] According to the technical solution of the present disclosure, the accuracy of speech recognition can be improved.
[0010] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] The drawings are used to better understand the solution and do not limit the present disclosure. Among them:
[0012] Figure 1 is a schematic diagram according to the first embodiment of the present disclosure;
[0013] Figure 2 It is a schematic diagram of the composition of the streaming speech recognition model in the embodiments of the present disclosure;
[0014] Figure 3 It is a schematic diagram of the application scenario for implementing the training method of the speech recognition model in the embodiments of the present disclosure;
[0015] Figure 4 It is a schematic diagram according to the second embodiment of the present disclosure;
[0016] Figure 5 is Figure 4 The corresponding system architecture diagram;
[0017] Figure 6 It is a schematic diagram according to the third embodiment of the present disclosure;
[0018] Figure 7 It is a schematic diagram of the electronic device for implementing the training method of the speech recognition model in the embodiments of the present disclosure. Detailed implementation manners
[0019] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for clarity and conciseness, descriptions of well-known functions and structures are omitted below.
[0020] In the related art, the streaming speech recognition model usually adopts a training method, that is, using speech and text as a sample pair for training. However, it has the problem of low character accuracy.
[0021] To improve the accuracy of speech recognition, the present disclosure provides the following embodiments.
[0022] Figure 1 It is a schematic diagram according to the first embodiment of the present disclosure. This embodiment provides a training method for a streaming speech recognition model. The streaming speech recognition model includes a latent feature extraction model and a context feature extraction model. The method includes:
[0023] 101. Based on the entire speech sample, obtain the prefix speech signal at the current moment. The prefix speech signal at the current moment includes: the speech signal before the current moment in the entire speech sample.
[0024] 102. Use the latent feature extraction model to perform feature extraction processing on the input prefix speech signal at the current moment to output latent features.
[0025] 103. Mask the potential features to obtain masked potential features.
[0026] 104. Use the context feature extraction model to perform feature extraction processing on the input masked potential features to output context features.
[0027] 105. Construct a loss function based on the context features.
[0028] 106. Based on the loss function, adjust the model parameters of the potential feature extraction model and the model parameters of the context feature extraction model.
[0029] In this embodiment, overall, the loss function is constructed based on context representations, and the context representations are obtained based on speech signals. Since only speech signals are utilized without the samples corresponding to the speech signals, the training method of this embodiment is different from the usual training methods and can be called a pre-training method.
[0030] In the usual training methods, when training a streaming speech recognition model using speech and its text, since the text is data that needs to be annotated and the annotation cost is very high, in contrast, unannotated data is easily obtainable. For this reason, since the pre-training method of this embodiment does not require annotated text, large-scale speech samples can be relatively easily obtained, and a streaming speech recognition model is trained based on the large-scale speech samples.
[0031] As Figure 2 shown, the streaming speech recognition model 200 includes: a latent representations extraction model 201 and a context feature extraction model 202. Among them, since it is applied to the field of speech recognition, the latent features can specifically be latent speech representations. The input of the latent feature extraction model 201 is a speech signal, and the output is latent features. The input of the context feature extraction model 202 is the masked latent features, and the output is context features.
[0032] Since it is applied to the field of streaming speech recognition, the speech signal input to the latent feature extraction model 201 can specifically be a prefix speech signal rather than the entire speech signal.
[0033] The prefix speech signal corresponds to a moment, and the prefix speech signal at the current moment includes: the speech signal before the current moment in the entire speech signal. "Before" includes the current moment.
[0034] For example, the entire speech signal can be represented by X, X = {x1, x2,..., xN}, where N is a positive integer, and x i represents the speech signal at the i-th moment, where i = 1, 2,..., N.
[0035] For the i-th moment, its prefix speech signal includes: {x1, x2,..., x i};
[0036] Among them, the prefix speech signal at the current moment includes the speech signal before the current moment in the whole speech signal, which can mean only including the speech signal before the current moment; or, it can also mean that in addition to including the speech signal before the current moment, it also includes the speech signal within a preset time difference after the current moment. Here, "after" does not include the current moment.
[0037] Still taking the i-th moment as an example, its prefix speech signal can be: {x1, x2,..., x i}; or,
[0038] Assuming that the preset time difference is one time period, the prefix speech signal at the i-th moment can also be: {x1, x2,..., x i , x i+1}.
[0039] It can be understood that for the excess part, it can be considered empty. For example, for the i = N-th moment, x i+1 can be set to be empty.
[0040] After obtaining the prefix speech signal, a streaming speech recognition model can be trained based on the prefix speech signal, that is, relevant features of the prefix speech signal are used to construct a loss function, and the constructed loss function is used to adjust the model parameters. When adjusting the model parameters, the usual Back Propagation (BP) algorithm can be used to adjust the model parameters until the training ends after reaching a predetermined number of iterations, and the model parameters at the end of the training are used as the final model parameters.
[0041] In this embodiment, by processing the prefix speech signal to obtain context features, constructing a loss function based on the context features, and adjusting the model parameters based on the loss function, since it is not necessary to label the text corresponding to the speech signal, the pre-training is applied to the training process of the streaming speech recognition model, thereby improving the speech recognition accuracy.
[0042] To better understand the embodiments of the present disclosure, an application scenario applicable to the embodiments of the present disclosure is described.
[0043] Figure 3It is a schematic diagram of an application scenario for implementing the training method of the speech recognition model according to an embodiment of the present disclosure. In this embodiment, speech recognition is taken as an example in a server.
[0044] As Figure 3 shown, the application scenario may include: a user device 301 and a server 302. The user device 301 and the server 302 interact with each other through a communication network. The user device may include a mobile device (such as a mobile phone, a portable computer, etc.), a smart home device (such as a smart speaker, a smart TV, etc.), a smart wearable device (such as a smart watch, a smart bracelet, etc.), etc. The server may be a local server or a cloud server. The communication network may be a wide area network, a local area network, the Internet, or any other public or private network or a combination of the above.
[0045] During speech recognition, the user device 301 may send a speech signal to the server 302, and the server 302 uses a speech recognition model to recognize the speech signal to obtain a speech recognition result. The speech recognition result is the text corresponding to the speech signal. After that, the server 302 feeds back the speech recognition result to the user device 301, and the user device 301 may display the speech recognition result to the user through a user interface (UI).
[0046] For streaming speech recognition, the speech recognition model used by the server 302 may be referred to as a streaming speech recognition model. During streaming speech recognition, the speech signal is recognized in real time, and the speech recognition result is output in real time.
[0047] For example, for the speech signal "I think I might be hungry", as Figure 3 shown, the speech recognition result will be output word by word.
[0048] It can be understood that in this embodiment, speech recognition is taken as an example in a server. However, if the user device has speech recognition capabilities, speech recognition can also be performed locally on the user device.
[0049] Combined with Figure 3 the application scenario shown, the embodiments of the present disclosure are described as follows.
[0050] Figure 4 It is a schematic diagram according to the second embodiment of the present disclosure, Figure 5 is Figure 4 the corresponding system architecture diagram.
[0051] As Figure 4 shown, the method of this embodiment includes:
[0052] 401. Based on the entire voice sample, obtain the prefix voice signal at the current moment, where the prefix voice signal at the current moment includes: the voice signal before the current moment in the entire voice sample.
[0053] Among them, in the entire voice sample, the voice signal before the current moment can be selected as the prefix voice signal at the current moment; or, in the entire voice sample, the voice signal before the current moment and the voice signal within a preset time difference after the current moment can be selected as the prefix voice signal at the current moment.
[0054] Exemplarily, assuming the current moment is the i-th moment, the prefix voice signal at the current moment refers to: {x1, x2,..., x i , x i+1}.
[0055] In this embodiment, since streaming speech recognition is character-by-character recognition and does not need to wait until the entire voice signal is input before recognition, therefore, by including the voice signal before the current moment in the prefix voice signal instead of the entire voice signal, it can be applicable to the streaming speech recognition scenario and improve the response speed of speech recognition. By including a segment of voice signal after the current moment in the prefix voice signal at the current moment, future information of the current moment can be referred to, thereby improving the accuracy of speech recognition.
[0056] 402. Use the potential feature extraction model to perform feature extraction processing on the input prefix voice signal at the current moment to output potential features.
[0057] Among them, referring to Figure 5 , taking the potential feature extraction model as a Convolutional Neural Network (CNN) model as an example.
[0058] After the prefix voice signal is input into the CNN model and processed by the CNN model, potential features can be output.
[0059] Among them, the prefix feature is represented by Z, corresponding to the prefix voice signal X = {x1, x2,..., x i , x i+1}, and the potential feature Z = {z1, z2,..., z i , z i+1}.
[0060] Under normal circumstances, speech recognition is completed based on the speech spectrum after time-frequency analysis, and the speech time-frequency spectrum has structural characteristics. To improve the speech recognition rate, it is necessary to overcome various diversities faced by the speech signal, including the diversity of speakers (speakers themselves and among speakers), the diversity of the environment, etc. Since CNN provides convolutional invariance in time and space, applying the idea of convolutional neural networks to speech recognition can utilize the invariance of convolution to overcome the diversity of the speech signal itself. From this perspective, it can be considered that the time-frequency spectrum obtained by analyzing the entire speech signal is treated as an image, and the widely used deep convolutional network in images is used to identify it. Therefore, in this embodiment, using CNN to extract the features of the speech signal can overcome the diversity features of the speech signal and improve the accuracy of speech recognition.
[0061] 403. Perform a masking process on the latent feature to obtain the masked latent feature.
[0062] Among them, since the latent feature is a vector, that is, a vector including multiple elements, during the masking process, specifically, a latent feature can be randomly selected for the masking process.
[0063] For example, the latent feature is Z = {z1, z2,..., z i , z i+1}, assuming that a randomly selected latent feature is z i , then the masked latent feature is: Z’ = {z1, z2,..., [MASK], z i+1}. Among them, [MSAK] is a masking character, and the masking character is, for example, a randomly generated character. Specifically, it can be determined by using the masking method in related technologies. For example, the masking method in the Encoder of the Bidirectional Encoder Representations from Transformers (BERT) can be used to generate [MASK].
[0064] In this embodiment, by randomly selecting a latent feature for the masking process, since it is randomly selected, the generalization ability of the speech recognition model can be improved.
[0065] 404. Use the context feature extraction model to perform feature extraction processing on the input masked latent feature to output context features.
[0066] Among them, referring to Figure 5 , taking the context feature extraction model as the Transformer model as an example.
[0067] After the masked latent feature Z' is input into the Transformer model and processed by the Transformer model, the context feature C can be output.
[0068] Among them, corresponding to the above-mentioned prefix voice signal and prefix feature, the context feature C can be expressed as: C = {c1, c2,..., c i , c i+1}.
[0069] It can be understood that when specifically using the Transformer model, various variants of the Transformer model can be adopted. For example, an encoder-decoder structure can be adopted, or only an encoder structure can be adopted, or only a decoder structure can be adopted, etc.
[0070] The Transformer model can well solve the sequence-to-sequence problem, and can obtain better speech recognition results while reducing the amount of calculation and improving the parallel efficiency. Therefore, in this embodiment, by using the Transformer model to extract features, better speech recognition results can be obtained while reducing the amount of calculation and improving the parallel efficiency.
[0071] 405. Perform quantization processing on the latent feature to obtain a quantization feature.
[0072] Among them, the quantization feature can be represented by Q, Q = {q1, q2,..., q i , q i+1}.
[0073] The quantization processing can specifically refer to product quantization. Product quantization refers to the Cartesian product, which means decomposing the original vector space into the Cartesian product of several low-dimensional vector spaces, and performing quantization on the decomposed low-dimensional vector spaces respectively. In this way, each vector can be represented by a combination of quantization codes of multiple low-dimensional spaces. Here, the quantization is to quantize the continuous space into a finite space. In addition, in order to ensure differentiability, a gumble softmax operation can be performed after quantization. The specific content of product quantization and gumble softmax operation can refer to related technologies.
[0074] It can be understood that there is no temporal sequence limit relationship between 404 and 405.
[0075] 406. Construct a loss function based on the context feature and the quantization feature.
[0076] Among them, asFigure 5 As shown, the loss function L is constructed based on the contextual features C and the quantitative features Q.
[0077] Specifically, the context feature is the context feature corresponding to the masked latent feature. That is, a latent feature of the random mask is z i , then the context feature used to construct the loss function is z i , the corresponding context feature c i .
[0078] The goal of the loss function is to make the contextual features corresponding to the masked latent features as consistent as possible with the quantitative features corresponding to the masked latent features, that is, to make c i =q i .
[0079] Specifically, the loss function may be a contrastive loss function. The core idea of contrastive learning is to bring positive samples closer together and increase the distance between positive and negative samples. To this end, a calculation formula that reflects this idea may be used.
[0080] For example, the contrastive loss function
[0081] Among them, τ is a hyperparameter, which is a preset value; K is the total number of quantitative features corresponding to the current moment; sim() is a similarity calculation function, which can be cosine similarity, or Euclidean distance, etc.
[0082] In addition, other factors may also be referenced in the loss function. For example, the loss function L=L1+L2, where L1 is a contrast loss function and L2 is a function constructed based on the quantitative feature Q. For details, please refer to the construction content of the loss function for the quantitative feature.
[0083] In this embodiment, a loss function is constructed based on context features and quantitative features, which eliminates the need for labeled data and improves speech recognition effects.
[0084] 407. Based on the loss function, adjust the model parameters of the latent feature extraction model and the model parameters of the context feature extraction model.
[0085] Among them, a common model parameter adjustment method can be adopted, for example, the BP algorithm can be used to adjust the model parameters until the training ends after reaching a predetermined number of iterations, and the model parameters at the end of the training are used as the final model parameters.
[0086] In this embodiment, the streaming speech recognition model takes the CNN model and the Transformer model as an example. It can be understood that other structures can also be adopted, such as using only the CNN model, or using the CNN model and the Recurrent Neural Network (RNN) model, etc.
[0087] In this embodiment, obtaining the prefix speech signal at the current moment based on the entire speech sample is applicable to streaming speech recognition and can improve the speech recognition efficiency. Using the CNN model to extract the latent features of the prefix speech signal can utilize the advantage of the CNN model in effectively solving diversity and improve the speech recognition effect. Using the Transformer model to extract the context features can utilize the advantages of the Transformer model and improve the speech recognition effect. By performing quantization processing on the latent features, the infinite continuous space can be converted into a finite discrete space, making the features more robust and not affected by a small amount of perturbations, thereby improving the representational ability of the features. By constructing a contrastive loss function based on the context features and the quantized features, self-supervised learning can be achieved. When constructing the loss function, the corresponding text of the speech is not required, reducing the dependence on labeled data, achieving pre-training, and improving the speech recognition accuracy.
[0088] Figure 6 It is a schematic diagram according to the third embodiment of the present disclosure. This embodiment provides a training device for a streaming speech recognition model. As Figure 6 shown, the streaming speech recognition model includes a latent feature extraction model and a context feature extraction model. The training device 600 of the streaming speech recognition model includes: an acquisition module 601, a first feature extraction module 602, a masking processing module 603, a second feature extraction module 604, a construction module 605, and an adjustment module 606.
[0089] The acquisition module 601 is used to obtain the prefix speech signal at the current moment based on the entire speech sample. The prefix speech signal at the current moment includes: the speech signal before the current moment in the entire speech sample. The first feature extraction module 602 is used to perform feature extraction processing on the input prefix speech signal at the current moment by using the latent feature extraction model to output latent features. The masking processing module 603 is used to perform masking processing on the latent features to obtain the masked latent features. The second feature extraction module 604 is used to perform feature extraction processing on the input masked latent features by using the context feature extraction model to output context features. The construction module 605 is used to construct a loss function based on the context features. The adjustment module 606 is used to adjust the model parameters of the latent feature extraction model and the model parameters of the context feature extraction model based on the loss function.
[0090] In this embodiment, by processing the prefix voice signal to obtain context features, constructing a loss function based on the context features, and adjusting the model parameters based on the loss function, since it is not necessary to label the text corresponding to the voice signal, the pre-training is applied to the training process of the streaming voice recognition model, thereby improving the voice recognition accuracy.
[0091] In some embodiments, the constructing module 605 is further configured to: perform quantization processing on the latent features to obtain quantization features; and construct a loss function based on the context features and the quantization features.
[0092] In this embodiment, constructing a loss function based on the context features and the quantization features can improve the voice recognition effect without the need for labeled data.
[0093] In some embodiments, the obtaining module 601 is further configured to: select, from the entire voice sample, the voice signal before the current moment as the prefix voice signal of the current moment; or, select, from the entire voice sample, the voice signal before the current moment and the voice signal within a preset time difference after the current moment as the prefix voice signal of the current moment.
[0094] In this embodiment, since streaming voice recognition is character-by-character recognition and does not need to wait until the entire voice signal is input before recognition, therefore, by including the voice signal before the current moment instead of the entire voice signal in the prefix voice signal, it can be applied to the streaming voice recognition scenario and improve the response speed of voice recognition. By including a segment of the voice signal after the current moment in the prefix voice signal of the current moment, future information of the current moment can be referred to, thereby improving the voice recognition accuracy.
[0095] In some embodiments, the masking processing module 603 is further configured to: randomly select a latent feature for masking processing.
[0096] In this embodiment, by randomly selecting a latent feature for masking processing, since it is randomly selected, the generalization ability of the voice recognition model can be improved.
[0097] In some embodiments, the latent feature extraction model is: a CNN model; and / or, the context feature extraction model is: a Transformer model.
[0098] Under normal circumstances, speech recognition is completed based on the speech spectrum after time-frequency analysis, and the speech time-frequency spectrum has structural characteristics. To improve the speech recognition rate, it is necessary to overcome the various diversities faced by the speech signal, including the diversity of speakers (speakers themselves and among speakers), the diversity of the environment, etc. Since CNN provides translation-invariant convolutions in time and space, applying the idea of convolutional neural networks to speech recognition can utilize the invariance of convolutions to overcome the diversity of the speech signal itself. From this perspective, it can be considered that the time-frequency spectrum obtained by analyzing the entire speech signal is treated as an image, and the widely used deep convolutional network in images is used to identify it. Therefore, in this embodiment, using CNN to extract the features of the speech signal can overcome the diversity characteristics of the speech signal and improve the accuracy of speech recognition.
[0099] The Transformer model can well solve the sequence-to-sequence problem and achieve better speech recognition results while reducing the computational complexity and improving the parallel efficiency. Therefore, in this embodiment, by using the Transformer model to extract features, better speech recognition results can be achieved while reducing the computational complexity and improving the parallel efficiency.
[0100] It can be understood that in the embodiments of the present disclosure, the same or similar content in different embodiments can be referred to each other.
[0101] It can be understood that the "first", "second", etc. in the embodiments of the present disclosure are only used for distinction and do not represent the level of importance, the order of time sequence, etc.
[0102] In the technical solution of the present disclosure, the collection, storage, use, processing, transmission, provision, and disclosure of the user's personal information comply with the provisions of relevant laws and regulations and do not violate public order and good customs.
[0103] According to the embodiments of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0104] Figure 7FIG. shows a schematic block diagram of an exemplary electronic device 700 that can be used to implement embodiments of the present disclosure. The electronic device 700 is intended to represent various forms of digital computers, such as, for example, laptop computers, desktop computers, workstations, servers, blade servers, mainframe computers, and other suitable computers. The electronic device 700 can also represent various forms of mobile devices, such as, for example, personal digital assistants, cellular telephones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely exemplary and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0105] As Figure 7 shown, the electronic device 700 includes a computing unit 701 that can perform various appropriate actions and processes in accordance with a computer program stored in a read-only memory (ROM) 702 or a computer program loaded from a storage unit 708 into a random access memory (RAM) 703. In the RAM 703, various programs and data required for the operation of the electronic device 700 can also be stored. The computing unit 701, the ROM 702, and the RAM 703 are connected to each other via a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.
[0106] A plurality of components in the electronic device 700 are connected to the I / O interface 705, including: an input unit 706, such as a keyboard, a mouse, etc.; an output unit 707, such as various types of displays, speakers, etc.; a storage unit 708, such as a magnetic disk, an optical disk, etc.; and a communication unit 709, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 709 allows the electronic device 700 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0107] The computing unit 701 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 701 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 701 executes the various methods and processes described above, such as the training method of the streaming speech recognition model. For example, in some embodiments, the training method of the streaming speech recognition model can be implemented as a computer software program that is tangibly contained in a machine-readable medium, such as the storage unit 708. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 700 via the ROM 702 and / or the communication unit 709. When the computer program is loaded into the RAM 703 and executed by the computing unit 701, one or more steps of the training method of the streaming speech recognition model described above can be executed. Alternatively, in other embodiments, the computing unit 701 can be configured to execute the training method of the streaming speech recognition model by any other suitable means (e.g., by means of firmware).
[0108] Various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGA), application-specific integrated circuits (ASIC), application-specific standard products (ASSP), system-on-a-chip (SOC), complex programmable logic devices (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a dedicated or general-purpose programmable processor, and can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0109] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the program code is executed by the processor or controller, the functions / operations specified in the flowchart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as an independent software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0110] In the context of this disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0111] For purposes of providing an interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide for interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic, speech, or tactile input).
[0112] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of a communication network include: a local area network (LAN), a wide area network (WAN), and the Internet.
[0113] A computer system may include a client and a server. The client and the server are generally far from each other and usually interact via a communication network. The relationship between the client and the server is created by computer programs running on respective computers and having a client-server relationship with each other. The server may be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, solving the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services ("Virtual Private Server", or simply "VPS"). The server may also be a server of a distributed system, or a server combined with a blockchain.
[0114] It should be understood that various forms of processes shown above can be used, steps can be reordered, added or deleted. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and no limitations are imposed herein.
[0115] The above specific embodiments do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the protection scope of this disclosure.
Claims
1. A training method for a streaming speech recognition model, the streaming speech recognition model including a latent feature extraction model and a context feature extraction model, the method comprising: Based on the entire speech sample, obtaining the prefix speech signal at the current moment, the prefix speech signal at the current moment including: the speech signal before the current moment in the entire speech sample; Using the latent feature extraction model to perform feature extraction processing on the input prefix speech signal at the current moment to output latent features; Performing a masking process on the latent features to obtain masked latent features; Using the context feature extraction model to perform feature extraction processing on the input masked latent features to output context features; Constructing a loss function based on the context features; Adjusting the model parameters of the latent feature extraction model and the model parameters of the context feature extraction model based on the loss function; The constructing a loss function based on the context features includes: Performing quantization processing on the latent features to obtain quantized features; Constructing a loss function based on the context features and the quantized features.
2. The method according to claim 1, wherein, The obtaining the prefix speech signal at the current moment based on the entire speech sample includes: In the entire speech sample, selecting the speech signal before the current moment as the prefix speech signal at the current moment; or, In the entire speech sample, selecting the speech signal before the current moment and the speech signal within a preset time difference after the current moment as the prefix speech signal at the current moment.
3. The method according to claim 1, wherein The performing a masking process on the latent features to obtain masked latent features includes: Randomly selecting a latent feature for masking processing.
4. According to the method of any one of claims 1-3, wherein, The latent feature extraction model is: a convolutional neural network CNN model; and / or, The context feature extraction model is: a Transformer model.
5. A training device for a streaming speech recognition model, the streaming speech recognition model including a latent feature extraction model and a context feature extraction model, the device comprising: An obtaining module, configured to obtain the prefix speech signal at the current moment based on the entire speech sample, the prefix speech signal at the current moment including: the speech signal before the current moment in the entire speech sample; A first feature extraction module, configured to use the latent feature extraction model to perform feature extraction processing on the input prefix speech signal at the current moment to output latent features; A masking processing module, configured to perform a masking process on the latent features to obtain masked latent features; A second feature extraction module, configured to use the context feature extraction model to perform feature extraction processing on the input masked latent features to output context features; A constructing module, configured to construct a loss function based on the context features; An adjusting module, configured to adjust the model parameters of the latent feature extraction model and the model parameters of the context feature extraction model based on the loss function; The constructing module is further configured to: Quantify the potential features to obtain quantified features; Construct a loss function based on the context features and the quantified features.
6. The apparatus according to claim 5, wherein, The obtaining module is further configured to: In the entire speech sample, select the speech signal before the current moment as the prefix speech signal at the current moment; or, In the entire speech sample, select the speech signal before the current moment and the speech signal within a preset time difference after the current moment as the prefix speech signal at the current moment.
7. The device according to claim 5, wherein, The masking processing module is further configured to: Randomly select a potential feature for masking processing.
8. The apparatus according to any one of claims 5-7, wherein The potential feature extraction model is: a convolutional neural network CNN model; and / or, The context feature extraction model is: a Transformer model.
9. An electronic device, comprising: At least one processor; And A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method according to any one of claims 1-4.
10. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to execute the method according to any one of claims 1-4.
11. A computer program product, comprising a computer program which, when executed by a processor, implements the method according to any one of claims 1-4.
Citation Information
Patent Citations
Systems and methods for training dual-mode machine-learned speech recognition models
WO2022072801A2