Speech recognition method and device, computer device and storage medium
By combining speech encoders and speech word embedders, the accuracy problem of speech recognition models in different scenarios is solved, the accuracy of speech recognition is improved, and the efficiency and effectiveness of intelligent diagnosis and remote consultation are enhanced.
Patent Information
- Application Number
- CN202310688282.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-09
- Publication Date
- 2026-02-06
- Estimated Expiration
- 2043-06-09
AI Technical Summary
Existing speech recognition technology has low accuracy in different application scenarios, especially in intelligent diagnosis and treatment and remote consultation, resulting in low consultation efficiency and poor results.
The speech encoder of the speech recognition model encodes the speech data to be recognized to obtain speech features, and then performs word embedding processing through a speech word embedder to obtain word embedding features. Combining the speech features and word embedding features for speech recognition improves the accuracy of speech recognition.
It improves the accuracy of speech recognition, thereby enhancing the efficiency and effectiveness of intelligent diagnosis and remote consultation.
Smart Images

Figure CN116597823B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of speech recognition and digital medical treatment, and in particular to a speech recognition method and device, a computer device and a storage medium. BACKGROUND
[0002] With the continuous development of speech technology and the application of speech recognition technology to more and more industries, the research on speech recognition technology by the industry is becoming more and more in-depth. For example, in the context of digital medical treatment, such as intelligent diagnosis and treatment and remote consultation, speech recognition technology is often used.
[0003] At present, in the field of speech recognition technology, speech data is generally input into a trained speech recognition model for speech recognition to obtain a text. The speech recognition model includes an acoustic model and a language model. In the process of end-to-end training and learning of the speech recognition model, the training of the speech recognition model includes information of the acoustic model and information of the language model, so that the speech recognition model is coupled with the language information in the training data. Since the speech recognition model is coupled with the language information in the training data, when the trained speech recognition model is applied to a scenario different from the training data, it is easy to cause the recognition accuracy of the trained speech recognition model to be low. Therefore, the existing speech recognition technology has the problem of low recognition accuracy when applied to a speech recognition scenario different from the training data, and in the context of intelligent diagnosis and treatment and remote consultation, low speech recognition accuracy will result in low efficiency and poor effect of inquiry. SUMMARY
[0004] Therefore, it is necessary to provide a speech recognition method, device, computer device and storage medium to solve the problem of low recognition accuracy in the existing speech recognition technology.
[0005] A speech recognition method comprises:
[0006] obtaining to-be-recognized speech data;
[0007] encoding the to-be-recognized speech data by using a speech encoder of a speech recognition model to obtain speech features;
[0008] performing word embedding processing on the speech features by using a speech word embedding device of the speech recognition model to obtain word embedding features;
[0009] performing speech recognition on the to-be-recognized speech data according to the speech features and the word embedding features to obtain a speech recognition result.
[0010] A speech recognition device comprises:
[0011] The voice data to be recognized is obtained.
[0012] The voice feature is obtained by encoding the voice data to be recognized by a voice encoder of the voice recognition model.
[0013] The word embedding feature is obtained by performing word embedding processing on the voice feature by a voice word embedding device of the voice recognition model.
[0014] The voice recognition result is obtained by performing voice recognition on the voice data to be recognized according to the voice feature and the word embedding feature.
[0015] A computer device includes a memory, a processor, and computer readable instructions stored in the memory and executable on the processor, and the processor executes the computer readable instructions to implement the voice recognition method.
[0016] One or more readable storage media storing computer readable instructions, which are executed by one or more processors to cause the one or more processors to execute the voice recognition method.
[0017] The voice recognition method, device, computer device, and storage medium obtain voice data to be recognized, encode the voice data to be recognized by a voice encoder of a voice recognition model to obtain voice features, perform word embedding processing on the voice features by a voice word embedding device of the voice recognition model to obtain word embedding features, and perform voice recognition on the voice data to be recognized according to the voice features and the word embedding features to obtain a voice recognition result. The voice recognition result is obtained based on voice features and word embedding features, and the word embedding features contain hidden word information in the voice data to be recognized, so that the voice recognition result is not limited to word information in a training data set used for training of a voice recognition model, but also considers hidden word information of input data, which can improve the effect of grammar analysis and voice analysis in the voice recognition process, thereby improving the accuracy of voice recognition. The voice recognition method can be applied to intelligent diagnosis and treatment and remote consultation, so that the accuracy of voice recognition of both parties in the diagnosis can be improved, and the efficiency and effect of diagnosis can be improved. BRIEF DESCRIPTION OF DRAWINGS
[0018] To more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings needed to be used in the description of the embodiments of the present application. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.
[0019] Figure 1 is an application environment schematic diagram of a speech recognition method in an embodiment of the present application;
[0020] Figure 2 is a flow schematic diagram of a speech recognition method in an embodiment of the present application;
[0021] Figure 3 is a structure schematic diagram of a speech recognition device in an embodiment of the present application;
[0022] Figure 4 is a schematic diagram of a computer device in an embodiment of the present application. DETAILED DESCRIPTION
[0023] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are a part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all the other embodiments obtained by those of ordinary skill in the art without creative work fall within the scope of the present application.
[0024] The speech recognition method provided in the embodiment can be applied in an application environment as shown in Figure 1 , wherein a client and a server communicate. The client includes but is not limited to various personal computers, notebook computers, smart phones, tablet computers and portable wearable devices. The server can be implemented by an independent server or a server cluster composed of multiple servers.
[0025] In an embodiment, as shown in Figure 2 , a speech recognition method is provided. Taking the server in Figure 1 as an example, the method includes the following steps:
[0026] S10, obtaining speech data to be recognized.
[0027] Understandably, the speech data to be recognized refers to speech to be converted into a text.
[0028] S20, encoding the speech data to be recognized by using a speech encoder of a speech recognition model to obtain speech features.
[0029] It can be understood that the speech recognition model is a trained neural network model, and the speech recognition model includes a speech encoder. The speech encoder is configured to encode the input speech data to be recognized to convert the speech data to be recognized into encoded data. The speech recognition model is configured to extract speech features from the speech data to be recognized to obtain the speech features. Specifically, the speech encoder of the speech recognition model encodes the speech data to be recognized to convert the speech data to be recognized into encoded data. The speech features are features of the speech in the speech data to be recognized, and can be obtained from the encoded data.
[0030] In S30, the speech word embedding processor of the speech recognition model is used to perform word embedding processing on the speech features to obtain word embedding features.
[0031] It can be understood that the word embedding method includes but is not limited to neural networks, dimensionality reduction of word co-occurrence matrix, probability models, and explicit representation of the context in which the word is located. Here, the speech word embedding processor is a neural network model, which is configured to perform word embedding processing on the speech features to convert the speech features by dimensionality reduction to obtain the word embedding features. The word embedding features contain hidden word information in the speech data to be recognized. Specifically, the word embedding features can contain transliterated word information in the speech data to be recognized. For example, if the speech data to be recognized has a pronunciation of “yin”, the word embedding features can contain information of several words with the pronunciation of “yin”.
[0032] In S40, the speech recognition model is used to perform speech recognition on the speech data to be recognized according to the speech features and the word embedding features to obtain a speech recognition result.
[0033] It can be understood that the speech recognition on the speech data to be recognized refers to a process of decoding the speech features and the word embedding features into text data by a text decoder of the speech recognition model. The speech recognition result is the result of converting the speech data to be recognized into text data. The word embedding features contain hidden word information in the speech data to be recognized, and are not limited to the word information in the training data set used for training the speech recognition model, which can improve the effect of grammar analysis and speech analysis in the speech recognition process.
[0034] In steps S10-S40, the voice data to be recognized is obtained; the voice data to be recognized is encoded by a voice encoder of the voice recognition model to obtain voice features; the voice features are word embedding processed by a voice word embedding device of the voice recognition model to obtain word embedding features; and the voice data to be recognized is recognized according to the voice features and the word embedding features to obtain a voice recognition result. The voice recognition result of the embodiment is obtained based on the voice features and the word embedding features. The word embedding features contain hidden word information in the voice data to be recognized, so that the voice recognition result is not limited to the word information in the training data set used for training the voice recognition model, but also considers the hidden word information of the input data, which can improve the effect of grammar analysis and voice analysis in the voice recognition process, thereby improving the accuracy of voice recognition. The above voice recognition method can be applied to the field of digital medicine, for example, in intelligent diagnosis and treatment, remote consultation, the above voice recognition method can be used to improve the accuracy of voice recognition of the parties to the inquiry, thereby improving the efficiency and effect of the inquiry.
[0035] Optionally, in step S30, the voice features are processed by the voice word embedding device of the voice recognition model to obtain word embedding features, including:
[0036] S301, linear normalization processing is performed on the voice features to obtain normalized voice features;
[0037] S302, the normalized voice features are processed by the voice word embedding device to obtain a word embedding vector corresponding to the voice data to be recognized;
[0038] S303, the word embedding vector is decoded by a word vector decoder of the word embedding device to obtain the word embedding features.
[0039] It can be understood that linear normalization processing of voice features refers to a process of linearly transforming voice features and mapping voice features to corresponding categories to obtain normalized voice features. The voice features that are linearly transformed and normalized are processed by the word embedding device to obtain a word embedding vector, which further improves the ability of grammar analysis and voice analysis in voice recognition. Furthermore, the word embedding vector is decoded by the word vector decoder of the word embedding device to quickly obtain more accurate hidden word information, i.e., word embedding features. The word vector decoder is used to decode the word embedding vector to obtain the word embedding features. Preferably, the word vector decoder uses an attention mechanism to further strengthen learning of hidden word information in the word embedding vector, so that the word embedding features are more abundant.
[0040] Optionally, before step S20, the method further comprises:
[0041] S201, obtaining a voice data sample and a text data sample corresponding to the voice data sample;
[0042] S202, performing sample encoding processing on the voice data sample by an initial voice encoder of an initial voice recognition model to obtain a voice feature sample;
[0043] S203, performing sample word embedding processing on the voice feature sample by a voice word embedding device of the initial voice recognition model to obtain a word embedding feature sample;
[0044] S204, performing word vector conversion processing on the text data sample to obtain a word vector sample;
[0045] S205, inputting the voice feature sample, the word vector sample, and the word embedding feature sample into an initial text decoder of the initial voice recognition model for decoding processing to obtain an initial recognition result;
[0046] S206, determining a total loss value of the initial voice recognition model according to the voice data sample, the voice feature sample, the text data sample, and the initial recognition result;
[0047] S207, when the total loss value does not reach a preset convergence condition, iteratively updating initial parameters of the initial voice recognition model until the total loss value reaches the preset convergence condition, and taking the initial voice recognition model after convergence as the voice recognition model.
[0048] The voice data sample is a sample containing voice data. The text data sample is text data corresponding to the voice data sample. That is, the voice data sample and the text data sample correspond one-to-one. For example, the voice data sample is the voice data of "I am Chinese", and the text data sample is the text data of "I am Chinese". The initial voice recognition model is an untrained neural network model. The initial voice recognition model includes an initial voice encoder, an initial text decoder, and a voice word embedding. The initial voice encoder is used for sample encoding processing of the input voice data sample, converting the voice data sample into encoded data. The word vector conversion processing of the text data sample refers to the process of converting the text of the text data sample into a vector representation through a word vector conversion tool to obtain a word vector sample. The initial text decoder is used to analyze and process the voice feature sample according to the word vector sample and the word embedding feature sample, decode the voice feature sample into text information, and obtain an initial recognition result. The total loss value refers to the loss value generated by the initial voice recognition model during the training process. The training of the initial voice recognition model, that is, the joint training of the initial voice encoder and the initial text decoder.
[0049] In steps S201-S207, during the training process of the initial voice recognition model, the sample word embedding processing is performed on the voice data sample, so that the initial voice recognition result is not only limited to the word information in the text data sample used for model training, but also considers the hidden word information of the voice data sample, performs data augmentation based on the voice data, and can improve the ability of the voice recognition model to analyze grammar and voice, thereby improving the accuracy of voice recognition of the voice recognition model.
[0050] Optionally, in step S206, the total loss value of the initial voice recognition model is determined according to the voice data sample, the voice feature sample, the text data sample, and the initial recognition result, comprising:
[0051] S2061, determining a first loss value of the initial voice encoder according to the voice data sample and the voice feature sample;
[0052] S2062, determining a second loss value of the initial text decoder according to the text data sample and the initial recognition result;
[0053] S2063, determining the total loss value of the initial voice recognition model according to the first loss value and the second loss value.
[0054] The first loss value is a loss value generated by the initial speech encoder during a training process. The second loss value is a loss value generated by the initial text decoder during the training process. The loss value of the initial speech encoder can be determined according to the speech data sample and the speech feature sample, and the loss value of the initial text decoder can be determined according to the text data sample and the initial recognition result. According to the loss value of the initial speech encoder and the loss value of the initial text decoder, a total loss value of the initial speech recognition model is determined, so as to jointly train the initial speech encoder and the initial text decoder according to the total loss value.
[0055] Optionally, in step S2061, that is, determining the first loss value of the initial speech encoder according to the speech data sample and the speech feature sample, comprises:
[0056] S20611, performing time sequence alignment processing on the speech data sample and the speech feature sample through the connection time classification layer of the initial speech encoder, to obtain a first time sequence set of the speech data sample and a second time sequence set of the speech feature sample; the first time sequence set comprises a first time sequence of each sentence in the speech data sample; the second time sequence set comprises a second time sequence of speech features corresponding to each sentence in the speech feature sample;
[0057] S20612, determining the first loss value according to the sentence corresponding to the first time sequence and the speech features corresponding to the second time sequence.
[0058] Understandably, the connection time classification layer is used to perform time sequence alignment processing on the speech data sample and the speech feature sample. The time sequence alignment processing refers to aligning the first time sequence of each sentence in the speech data sample and the second time sequence of speech features corresponding to each sentence in the speech feature sample. The first time sequence refers to the time sequence of each sentence in the speech data sample. The second time sequence refers to the time sequence of speech features corresponding to each sentence in the speech feature sample. The sentence corresponding to the first time sequence and the speech features corresponding to the second time sequence are calculated through a preset loss function of the connection time classification layer, to obtain the first loss value.
[0059] Optionally, before step S10, that is, before the initial speech recognition model is obtained, comprising:
[0060] S101, obtaining initial speech data to be recognized;
[0061] S102, performing noise reduction processing on the initial speech data to be recognized, to obtain the speech data to be recognized.
[0062] It is understandable that the initial to-be-recognized speech data refers to the to-be-recognized speech data without processing. The noise reduction processing is to eliminate the noise and invalid messages in the initial to-be-recognized speech data, so as to improve the accuracy of speech recognition. Specifically, the initial to-be-recognized speech data can be filtered by a filter to eliminate the noise.
[0063] It should be understood that the size of the serial number of each step in the above embodiments does not mean the order of execution, and the execution order of each process should be determined according to its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.
[0064] In an embodiment, a speech recognition device is provided, which corresponds to the speech recognition method in the above embodiments. As shown in the figure, the speech recognition device includes a to-be-recognized speech data module 10, a speech feature module 20, a word embedding feature module 30, and a speech recognition result module 40. The functions of each module are described in detail as follows: Figure 3
[0065] The to-be-recognized speech data module 10 is configured to obtain to-be-recognized speech data;
[0066] The speech feature module 20 is configured to encode the to-be-recognized speech data by a speech encoder of a speech recognition model to obtain speech features;
[0067] The word embedding feature module 30 is configured to perform word embedding processing on the speech features by a speech word embedding device of the speech recognition model to obtain word embedding features;
[0068] The speech recognition result module 40 is configured to perform speech recognition on the to-be-recognized speech data according to the speech features and the word embedding features to obtain a speech recognition result.
[0069] Optionally, the word embedding feature module 30 includes:
[0070] The normalized speech feature unit is configured to perform linear normalization processing on the speech features to obtain normalized speech features;
[0071] The word embedding vector unit is configured to perform word embedding processing on the normalized speech features by the speech word embedding device to obtain a word embedding vector corresponding to the to-be-recognized speech data;
[0072] The word embedding feature unit is configured to perform decoding processing on the word embedding vector by a word vector decoder of the word embedding device to obtain the word embedding features.
[0073] Optionally, the speech feature module 20 includes:
[0074] A data sample unit is used to acquire voice data samples and text data samples corresponding to the voice data samples;
[0075] The speech feature sample unit is used to perform sample encoding processing on the speech data sample through the initial speech encoder of the initial speech recognition model to obtain speech feature samples;
[0076] The word embedding feature sample unit is used to perform sample word embedding processing on the speech feature sample through the speech word embedder of the initial speech recognition model to obtain the word embedding feature sample;
[0077] The word vector sample unit is used to perform word vector conversion processing on the text data sample to obtain word vector samples;
[0078] The initial recognition result unit is used to input the speech feature sample, the word vector sample, and the word embedding feature sample into the initial text decoder of the initial speech recognition model for decoding processing to obtain the initial recognition result;
[0079] The total loss value unit is used to determine the total loss value of the initial speech recognition model based on the speech data sample, the speech feature sample, the text data sample, and the initial recognition result.
[0080] The speech recognition model unit is used to iteratively update the initial parameters of the initial speech recognition model when the total loss value does not reach the preset convergence condition, until the total loss value reaches the preset convergence condition, and then use the converged initial speech recognition model as the speech recognition model.
[0081] Optionally, the total loss value unit includes:
[0082] The first loss value unit is used to determine the first loss value of the initial speech encoder based on the speech data sample and the speech feature sample.
[0083] The second loss value unit is used to determine the second loss value of the initial text decoder based on the text data sample and the initial recognition result.
[0084] The total loss value determination unit is used to determine the total loss value of the initial speech recognition model based on the first loss value and the second loss value.
[0085] Optionally, the first loss value unit includes:
[0086] a time sequence alignment unit, configured to perform time sequence alignment processing on the speech data samples and the speech feature samples by a connection time classification layer of the initial speech encoder, to obtain a first time sequence set of the speech data samples and a second time sequence set of the speech feature samples; the first time sequence set comprises a first time sequence of each sentence in the speech data samples; the second time sequence set comprises a second time sequence of speech features corresponding to the each sentence in the speech feature samples;
[0087] a first loss value determination unit, configured to determine the first loss value according to the sentence corresponding to the first time sequence and the speech feature corresponding to the second time sequence.
[0088] Optionally, before the speech data to be recognized module 10, comprising:
[0089] an initial speech data to be recognized module, configured to obtain initial speech data to be recognized;
[0090] a noise reduction module, configured to perform noise reduction processing on the initial speech data to be recognized to obtain the speech data to be recognized.
[0091] The specific limitations of the speech recognition device can be referred to the limitations of the speech recognition method in the above, which will not be repeated here. Each module in the above speech recognition device can be realized by software, hardware and their combination in whole or in part. The above modules can be embedded in or independent of the processor in the computer device in hardware form, or can be stored in the memory in the computer device in software form, so as to call and execute the operations corresponding to the above modules by the processor.
[0092] In one embodiment, a computer device is provided, which can be a server, and its internal structure diagram can be as shown in Figure 4 The computer device comprises a processor, a memory, a network interface and a database connected by a system bus. The processor of the computer device is configured to provide computing and control capabilities. The memory of the computer device comprises a readable storage medium and an internal memory. The readable storage medium stores an operating system, computer readable instructions and a database. The internal memory provides an environment for the operation of the operating system and computer readable instructions in the readable storage medium. The database of the computer device is configured to store data related to the speech recognition method. The network interface of the computer device is configured to communicate with the external terminal through the network connection. The computer readable instructions are executed by the processor to implement a speech recognition method. The readable storage medium provided in the embodiment comprises a non-volatile readable storage medium and a volatile readable storage medium.
[0093] In one embodiment, a computer device is provided, comprising a memory, a processor, and computer readable instructions stored on the memory and executable on the processor, the processor implementing the following steps when executing the computer readable instructions:
[0094] obtaining voice data to be recognized;
[0095] encoding the voice data to be recognized by a speech encoder of a speech recognition model to obtain speech features;
[0096] performing word embedding processing on the speech features by a speech word embedder of the speech recognition model to obtain word embedding features;
[0097] performing speech recognition on the voice data to be recognized according to the speech features and the word embedding features to obtain a speech recognition result.
[0098] In one embodiment, one or more computer readable storage media storing computer readable instructions are provided. The computer readable storage media provided by the present embodiment includes non-volatile readable storage media and volatile readable storage media. The computer readable instructions are stored on the readable storage media, and when executed by one or more processors, the following steps are implemented:
[0099] obtaining voice data to be recognized;
[0100] encoding the voice data to be recognized by a speech encoder of a speech recognition model to obtain speech features;
[0101] performing word embedding processing on the speech features by a speech word embedder of the speech recognition model to obtain word embedding features;
[0102] performing speech recognition on the voice data to be recognized according to the speech features and the word embedding features to obtain a speech recognition result.
[0103] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing relevant hardware through computer readable instructions, and the computer readable instructions can be stored in a non-volatile readable storage medium or a volatile readable storage medium. When the computer readable instructions are executed, the processes of the above-mentioned embodiments of the methods can be included. Any reference to memory, storage, database or other medium used in each embodiment provided by the present application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration but not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0104] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above-mentioned division of functional units and modules is exemplified, and in actual application, the above-mentioned functions can be completed by different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above.
[0105] The above-described embodiments are only used to illustrate the technical solutions of the present application, rather than limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that they can modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement for part of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should be included in the protection scope of the present application.
Claims
1. A voice recognition method, characterized by, The method comprises: obtaining to-be-recognized voice data; encoding the to-be-recognized voice data through a voice encoder of a voice recognition model to obtain voice features; performing word embedding processing on the voice features through a voice word embedding device of the voice recognition model to obtain word embedding features; performing voice recognition on the to-be-recognized voice data according to the voice features and the word embedding features to obtain a voice recognition result; before encoding the to-be-recognized voice data through a voice encoder of a voice recognition model to obtain voice features, the method comprises: obtaining voice data samples and text data samples corresponding to the voice data samples; performing sample encoding processing on the voice data samples through an initial voice encoder of an initial voice recognition model to obtain voice feature samples; performing sample word embedding processing on the voice feature samples through a voice word embedding device of the initial voice recognition model to obtain word embedding feature samples; performing word vector conversion processing on the text data samples to obtain word vector samples; inputting the voice feature samples, the word vector samples, and the word embedding feature samples into an initial text decoder of the initial voice recognition model for decoding processing to obtain an initial recognition result; determining a total loss value of the initial voice recognition model according to the voice data samples, the voice feature samples, the text data samples, and the initial recognition result; when the total loss value does not reach a preset convergence condition, iteratively updating initial parameters of the initial voice recognition model until the total loss value reaches the preset convergence condition, and taking the initial voice recognition model after convergence as the voice recognition model.
2. The voice recognition method of claim 1, wherein, The method comprises: performing linear normalization processing on the voice features to obtain normalized voice features; performing word embedding processing on the normalized voice features through the voice word embedding device to obtain a word embedding vector corresponding to the to-be-recognized voice data; performing decoding processing on the word embedding vector through a word vector decoder of the voice word embedding device to obtain the word embedding features.
3. The voice recognition method of claim 1, wherein, The method comprises: determining a first loss value of the initial voice encoder according to the voice data samples and the voice feature samples; determining a second loss value of the initial text decoder according to the text data samples and the initial recognition result; determining a total loss value of the initial voice recognition model according to the first loss value and the second loss value.
4. The voice recognition method of claim 3, wherein, The method comprises: The speech data sample and the speech feature sample are subjected to time sequence alignment processing by a connection time classification layer of the initial speech encoder, to obtain a first time sequence set of the speech data sample and a second time sequence set of the speech feature sample; the first time sequence set comprises a first time sequence of each sentence in the speech data sample; the second time sequence set comprises a second time sequence of speech features corresponding to the each sentence in the speech feature sample; The first loss value is determined according to the sentence corresponding to the first time sequence and the speech feature corresponding to the second time sequence.
5. The voice recognition method of claim 1, wherein, Before the obtaining of the to-be-recognized speech data, comprising: obtaining initial to-be-recognized speech data; performing noise reduction processing on the initial to-be-recognized speech data to obtain the to-be-recognized speech data.
6. A speech recognition apparatus characterized by comprising: Comprising: a to-be-recognized speech data module for obtaining to-be-recognized speech data; a speech feature module for performing encoding processing on the to-be-recognized speech data by a speech encoder of a speech recognition model to obtain speech features; a word embedding feature module for performing word embedding processing on the speech features by a speech word embedding device of the speech recognition model to obtain word embedding features; a speech recognition result module for performing speech recognition on the to-be-recognized speech data according to the speech features and the word embedding features to obtain a speech recognition result; The speech feature module comprises: a data sample unit for obtaining speech data samples and text data samples corresponding to the speech data samples; a speech feature sample unit for performing sample encoding processing on the speech data samples by an initial speech encoder of an initial speech recognition model to obtain speech feature samples; a word embedding feature sample unit for performing sample word embedding processing on the speech feature samples by a speech word embedding device of the initial speech recognition model to obtain word embedding feature samples; a word vector sample unit for performing word vector conversion processing on the text data samples to obtain word vector samples; an initial recognition result unit for inputting the speech feature samples, the word vector samples and the word embedding feature samples into an initial text decoder of the initial speech recognition model for decoding processing to obtain an initial recognition result; a total loss value unit for determining a total loss value of the initial speech recognition model according to the speech data samples, the speech feature samples, the text data samples and the initial recognition result; a speech recognition model unit for iteratively updating initial parameters of the initial speech recognition model when the total loss value does not reach a preset convergence condition, until the total loss value reaches the preset convergence condition, and taking the initial speech recognition model after convergence as the speech recognition model.
7. The speech recognition apparatus of claim 6, wherein The word embedding feature module comprises: a normalized speech feature unit for performing linear normalization processing on the speech features to obtain normalized speech features; a word embedding vector unit for performing word embedding processing on the normalized speech features by the speech word embedding device to obtain a word embedding vector corresponding to the to-be-recognized speech data; a word embedding feature unit configured to decode the word embedding vector by a word vector decoder of the word embedding unit to obtain a word embedding feature.
8. A computer device comprising a memory, a processor, and computer readable instructions stored in the memory and executable on the processor, wherein, The processor implements the speech recognition method in any one of claims 1-5 when executing the computer-readable instructions.
9. One or more readable storage media storing computer readable instructions, wherein, The computer-readable instructions, when executed by one or more processors, cause the one or more processors to perform the speech recognition method in any one of claims 1-5.
Citation Information
Patent Citations
Speech recognition method and related device
CN115101075A