Speech recognition method, device and equipment

By introducing a hot word coding mechanism into the speech recognition model, the voice features and hot word coding data are fused, and the problem of inaccurate hot word recognition is solved, and the accuracy and system performance of speech recognition are improved.

CN120279898APending Publication Date: 2025-07-08ALIPAY (HANGZHOU) INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510414136.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-02
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

In the prior art, in speech recognition, especially in hot word recognition, there is a problem that cannot be accurately recognized, resulting in a degradation of common word recognition performance.

Method used

By introducing an additional hot word coding mechanism into the speech recognition model, the pronunciation features are fused with the hot word coding data, and explicitly use the set hot word text data to accurately identify specific hot words.

Benefits of technology

It improves the accuracy of speech recognition and improves the recognition ability of speech recognition system in specific scenarios and specific vocabulary.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279898A_ABST
    Figure CN120279898A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a voice recognition method, device and equipment. The method comprises the following steps: receiving voice data input by a user and hot word text data which is input by the user and records a target hot word contained in the voice data; speech features in the speech data are extracted to obtain first speech features corresponding to the speech data, the first speech features are coded to obtain speech coded data corresponding to the first speech features, and each target hot word in the hot word text data is coded to obtain hot word coded data corresponding to each target hot word; performing fusion processing on the voice coding data corresponding to the first voice feature and the hot word coding data corresponding to each target hot word to obtain first fusion data, and determining voice representation corresponding to the voice data based on the first fusion data; and based on the voice representation, performing voice recognition processing on the voice data through a voice large model to obtain a voice recognition result of the voice data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This document relates to the field of computer technology, and in particular, to a speech recognition method, apparatus, and device. Background Art

[0002] With the continuous development of terminal technology and large models, on the premise of protecting users' private data, interaction between users or between humans and machines through voice has become an important interaction method. And for interaction between users or between humans and machines through voice, speech recognition is required. When applying a speech large model to actual speech recognition, it is usually encountered that hot words (or keywords, named entities, etc., referring to some special and uncommon nouns, verbs, etc. that appear in speech recognition processing, such as personal names, place names, proper nouns, etc.) cannot be accurately recognized. Therefore, it is necessary to provide a better speech recognition solution, especially for speech recognition processing that includes hot words. Summary of the Invention

[0003] The purpose of the embodiments of this specification is to provide a better speech recognition solution, especially for speech recognition processing that includes hot words.

[0004] To achieve the above technical solution, the embodiments of this specification are implemented as follows: A speech recognition method provided by the embodiments of this specification, the method includes: receiving the speech data input by the user and the hot word text data recorded by the user and containing the target hot words in the speech data; extracting the speech features in the speech data to obtain the first speech features corresponding to the speech data, and performing encoding processing on the first speech features to obtain the speech coding data corresponding to the first speech features, and performing encoding processing on each target hot word in the hot word text data to obtain the hot word coding data corresponding to each target hot word; performing fusion processing on the speech coding data corresponding to the first speech features and the hot word coding data corresponding to each target hot word to obtain the first fusion data, and based on the first fusion data, determining the speech representation corresponding to the speech data; based on the speech representation, performing speech recognition processing on the speech data through a speech large model to obtain the speech recognition result of the speech data.

[0005] A voice recognition device provided by an embodiment of this specification, the device includes: a data receiving module, which receives the voice data input by the user and the hot word text data input by the user and recording the target hot words included in the voice data; an encoding module, which extracts the voice features in the voice data to obtain the first voice features corresponding to the voice data, and performs encoding processing on the first voice features to obtain the voice encoding data corresponding to the first voice features, and performs encoding processing on each target hot word in the hot word text data to obtain the hot word encoding data corresponding to each target hot word; a fusion module, which performs fusion processing on the voice encoding data corresponding to the first voice features and the hot word encoding data corresponding to each target hot word to obtain first fusion data, and determines the voice representation corresponding to the voice data based on the first fusion data; a voice recognition module, which performs voice recognition processing on the voice data through a voice large model based on the voice representation to obtain the voice recognition result of the voice data.

[0006] A voice recognition device provided by an embodiment of this specification, the voice recognition device includes: a processor; and a memory arranged to store computer-executable instructions, the executable instructions, when executed, cause the processor to: receive the voice data input by the user and the hot word text data input by the user and recording the target hot words included in the voice data; extract the voice features in the voice data to obtain the first voice features corresponding to the voice data, and perform encoding processing on the first voice features to obtain the voice encoding data corresponding to the first voice features, and perform encoding processing on each target hot word in the hot word text data to obtain the hot word encoding data corresponding to each target hot word; perform fusion processing on the voice encoding data corresponding to the first voice features and the hot word encoding data corresponding to each target hot word to obtain first fusion data, and determine the voice representation corresponding to the voice data based on the first fusion data; perform voice recognition processing on the voice data through a voice large model based on the voice representation to obtain the voice recognition result of the voice data.

[0007] An embodiment of this specification also provides a storage medium for storing computer-executable instructions, and when the executable instructions are executed by a processor, the following process is implemented: receiving voice data input by a user and hotword text data input by the user and recording the target hotwords included in the voice data; extracting voice features in the voice data to obtain a first voice feature corresponding to the voice data, performing encoding processing on the first voice feature to obtain voice encoding data corresponding to the first voice feature, and performing encoding processing on each target hotword in the hotword text data to obtain hotword encoding data corresponding to each target hotword; performing fusion processing on the voice encoding data corresponding to the first voice feature and the hotword encoding data corresponding to each target hotword to obtain first fusion data, and determining a voice representation corresponding to the voice data based on the first fusion data; performing voice recognition processing on the voice data through a voice large model based on the voice representation to obtain a voice recognition result of the voice data.

[0008] An embodiment of this specification also provides a computer program product, including a computer program, and when the computer program is executed by a processor, the following process is implemented: receiving voice data input by a user and hotword text data input by the user and recording the target hotwords included in the voice data; extracting voice features in the voice data to obtain a first voice feature corresponding to the voice data, performing encoding processing on the first voice feature to obtain voice encoding data corresponding to the first voice feature, and performing encoding processing on each target hotword in the hotword text data to obtain hotword encoding data corresponding to each target hotword; performing fusion processing on the voice encoding data corresponding to the first voice feature and the hotword encoding data corresponding to each target hotword to obtain first fusion data, and determining a voice representation corresponding to the voice data based on the first fusion data; performing voice recognition processing on the voice data through a voice large model based on the voice representation to obtain a voice recognition result of the voice data. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] In order to more clearly illustrate the technical solutions in the embodiments of this specification or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings described below are only some embodiments recorded in this specification. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings. Figure 1 It is a schematic diagram of an embodiment of a voice recognition method in this specification; Figure 2 It is a schematic diagram of a voice recognition page in this specification; Figure 3Schematic diagram of another embodiment of the speech recognition method in this specification; Figure 4 Schematic diagram of yet another embodiment of the speech recognition method in this specification; Figure 5 Schematic diagram of the structure of a speech recognition system in this specification; Figure 6 Schematic diagram of the structure of another speech recognition system in this specification; Figure 7 Schematic diagram of the structure of yet another speech recognition system in this specification; Figure 8 Schematic diagram of a speech recognition device in this specification; Figure 9 Schematic diagram of a speech recognition device in this specification. Detailed implementation manners

[0010] The embodiments of this specification provide a speech recognition method, device and equipment.

[0011] In order to enable those skilled in the art to better understand the technical solutions in this specification, the following will clearly and completely describe the technical solutions in the embodiments of this specification with reference to the accompanying drawings in the embodiments of this specification. Obviously, the described embodiments are only a part of the embodiments of this specification, rather than all the embodiments. Based on the embodiments in this specification, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of this specification.

[0012] Embodiments of this specification provide a voice recognition mechanism. When applying a large language model to actual voice recognition, it is usually encountered that hot words (or keywords, named entities, etc., referring to some special and uncommon nouns, verbs, etc. that appear in voice recognition processing, such as personal names, place names, proper nouns, etc.) cannot be accurately recognized. Therefore, how to perform accurate voice recognition has become an important topic to be studied. Usually, a voice recognition model with a hot word module can be designed. For example, the C-LAS (Listen, Attend and Spell) model. Based on the model architecture of the C-LAS model, a complete ASR (Automatic Speech Recognition) model can be trained. Or, on the basis of an existing ASR model, a language model with a hot word enhancement function can be added for enhanced processing of voice recognition. However, the above-mentioned language model with a hot word enhancement function will cause performance loss to the original ASR model, resulting in a decline in the recognition performance of the original general words. This makes both users and model trainers have to weigh between the hot word recognition accuracy and the general recognition accuracy. For this reason, a better voice recognition solution is needed, especially the voice recognition processing including hot words. Embodiments of this specification provide an implementable technical solution. By additionally adding a hot word encoding mechanism, and then fusing the encoded data of the hot words with the encoded data of the voice data, in this way, the set hot words can be explicitly used and accurate recognition can be performed for specific hot words, thereby greatly improving the accuracy of voice recognition. The specific processing can refer to the specific content in the following embodiments.

[0013] As Figure 1 shown, embodiments of this specification provide a voice recognition method. The execution subject of this method can be a terminal device or a server, etc. The terminal device can be a mobile terminal device such as a mobile phone or a tablet computer, or a computer device such as a laptop computer or a desktop computer, or it can also be an IoT device (specifically such as a smart watch, a vehicle-mounted device, etc.). The server can be an independent server, or a server cluster composed of multiple servers. The server can be a background server in the financial field or the online shopping field, etc., or a background server of a certain application program, etc. In this embodiment, the case where the execution subject is a server is taken as an example for detailed description. For the case where the execution subject is a terminal device, it can refer to the case processing of the following server and will not be elaborated here. The method can specifically include the following steps: In step S102, receive the voice data input by the user and the hot word text data input by the user that records the target hot words included in the voice data.

[0014] Among them, the user can be any user. In this embodiment, the user can be a user who inputs voice data to request voice recognition. The voice data can include one or more different types of content. For example, the voice data can only include the data of the voice to be recognized, or the voice data can not only include the data of the voice to be recognized, but also include the user's demand information, etc. The demand information can include various types. For example, the demand information can include information related to how to process the voice to be recognized. Specifically, for example, recognize the aforementioned voice and extract the purpose expressed by the voice, etc. It can be specifically set according to the actual situation. The target hot word can be a hot word included in the voice data. The hot word can be some special and uncommon words that appear in the voice recognition process. These words can be nouns or verbs, etc. Specifically, such as personal names, place names, proper nouns, etc. In practical applications, the hot word can also be a keyword or a named entity, etc. It can be specifically set according to the actual situation.

[0015] In implementation, when the user needs to recognize a certain voice data, the user can obtain a voice recognition page through the terminal device, as Figure 2 shown. The voice recognition page can include an input trigger button for voice data, an input box for hot word text data, an output box for voice recognition results, a confirm button, a cancel button, etc. The input trigger button can implement the input of voice data in multiple different ways. For example, the user presses the input trigger button and keeps pressing it without releasing. At this time, the user can input voice data. After the input is completed, release the pressed input trigger button, or the user clicks the input trigger button. At this time, the voice recognition page can prompt the user to start inputting voice data. Then, the user can input voice data according to the prompt. After the input is completed, the user can click the input trigger button again, etc. It can be specifically set according to the actual situation. The user can input voice data through the above input trigger button and can use the above voice data input method. In addition, the user can also determine the hot word (i.e., the target hot word) included in the voice data and can input the above target hot word in the input box for hot word text data. After the input is completed, the user can click the confirm button. At this time, the user's terminal device can obtain the voice data input by the user and obtain the hot word text data input by the user recording the target hot word included in the voice data from the input box for hot word text data, and can send the voice data and the hot word text data to the server. The server can receive the voice data input by the user and the hot word text data input by the user recording the target hot word included in the voice data.

[0016] In step S104, speech features in the above speech data are extracted to obtain the first speech features corresponding to the speech data, and the first speech features are encoded to obtain speech coding data corresponding to the first speech features. Each target hot word in the hot word text data is encoded to obtain hot word coding data corresponding to each target hot word.

[0017] Among them, the first speech features can include various types. For example, the first speech features can include user voice characteristics, user pronunciation habit characteristics, etc., which can be specifically set according to the actual situation. Both the speech coding data and the hot word coding data can be represented in the form of vectors, or in the form of matrices, etc., which can be specifically set according to the actual situation.

[0018] In implementation, the speech data can be subjected to speech feature extraction in a variety of different ways. For example, a speech feature extraction algorithm or a speech feature extraction network model can be preset. The speech data can be subjected to speech feature extraction through the speech feature extraction algorithm or the speech feature extraction network model to obtain the first speech features corresponding to the speech data. Among them, the speech feature extraction algorithm can be, for example, the Mel Frequency Cepstral Coefficient algorithm, the Linear Predictive Coding algorithm, etc. The speech feature extraction network model can be a model constructed by a machine learning network. For example, a speech feature extraction network model can be constructed through a convolutional neural network, or a speech feature extraction network model can be constructed through a recurrent neural network model, or a speech feature extraction network model can be constructed based on a Transformer module, etc., which can be specifically set according to the actual situation.

[0019] It is also possible to encode the first speech features and the hot word text data separately. Specifically, an encoder or an encoding network model can be used. The first speech features can be encoded through the specified encoder or encoding network model to obtain the speech coding data corresponding to the first speech features. Similarly, each target hot word in the hot word text data can also be encoded through the specified encoder or encoding network model to obtain the hot word coding data corresponding to each target hot word. Among them, the encoder or encoding network model used for encoding the first speech features is different from the encoder or encoding network model used for encoding the hot word text data. The encoding network model can be a model constructed by a machine learning network. For example, an encoding network model can be constructed through a convolutional neural network, or an encoding network model can be constructed through a recurrent neural network model, or an encoding network model can be constructed based on BERT, etc., which can be specifically set according to the actual situation. By encoding the first speech features, the speech coding data corresponding to the first speech features can be obtained. By encoding each target hot word in the hot word text data, the hot word coding data corresponding to each target hot word can be obtained.

[0020] In step S106, the voice coding data corresponding to the first voice feature is fused with the hot word coding data corresponding to each target hot word to obtain first fusion data, and based on the first fusion data, the voice representation corresponding to the above voice data is determined.

[0021] In implementation, the voice coding data corresponding to the first voice feature can be fused with the hot word coding data corresponding to each target hot word in various ways. For example, a fusion algorithm can be preset, and the voice coding data corresponding to the first voice feature can be fused with the hot word coding data corresponding to each target hot word through this fusion algorithm. Or, a corresponding network model can be pre-constructed and trained (such as a convolutional neural network model or a recurrent neural network model, etc., which can be set according to the actual situation), and the voice coding data corresponding to the first voice feature can be fused with the hot word coding data corresponding to each target hot word through this network model, etc. In the process of fusing the voice coding data corresponding to the first voice feature with the hot word coding data corresponding to each target hot word, the voice coding data corresponding to the first voice feature and the hot word coding data corresponding to each target hot word can be fused in various fusion ways. For example, corresponding weights can be set for the voice coding data corresponding to the first voice feature and the hot word coding data corresponding to each target hot word. Among them, since the voice data contains target hot words and may also include words other than target hot words, the voice coding data corresponding to the first voice feature will be more than the hot word coding data corresponding to multiple target hot words. For the same word (i.e., the target hot word), a higher weight can be set for the hot word coding data corresponding to this target hot word, and a lower weight can be set for the voice coding data of the first voice feature corresponding to this target hot word. Or, the hot word coding data corresponding to this target hot word can be directly used to replace the voice coding data of the first voice feature corresponding to this target hot word, etc., which can be set according to the actual situation. Through the above processing, the voice coding data corresponding to the first voice feature can be fused with the hot word coding data corresponding to each target hot word to obtain first fusion data (specifically, it can be represented in the form of a vector or a matrix, etc.).

[0022] The first fusion data can be directly determined as the voice representation corresponding to the above voice data. Or, considering that the first fusion data may be different from the rules of the voice representation allowed by the voice large model, at this time, the first fusion data can also be subjected to a conversion process to convert the first fusion data into data that matches the rules of the voice representation allowed by the voice large model. The converted first fusion data can be determined as the voice representation corresponding to the above voice data. Or, other calculation processes can also be performed on the first fusion data, and finally the voice representation corresponding to the above voice data can be obtained (specifically, it can be represented in the form of a vector or a matrix, etc.), etc., which can be set according to the actual situation.

[0023] In step S108, based on the speech representation, the speech recognition processing is performed on the above-mentioned speech data through a speech large model to obtain the speech recognition result of the speech data.

[0024] In implementation, the speech representation can be input into the speech large model, and the speech large model can perform speech recognition processing on the speech data through the speech representation, and finally can output the speech recognition result of the speech data. In addition, if the speech data also includes the user's requirement information, the speech large model can also output corresponding results according to the user's requirement information, which can be specifically set according to the actual situation, and the embodiments of this specification do not limit this.

[0025] The embodiments of this specification provide a speech recognition method. By receiving the speech data input by the user and the hot word text data recording the target hot words included in the speech data, then, the speech features in the speech data can be extracted to obtain the first speech features corresponding to the speech data, and the first speech features are encoded to obtain the speech coding data corresponding to the first speech features. Each target hot word in the hot word text data is encoded to obtain the hot word coding data corresponding to each target hot word. After that, the speech coding data corresponding to the first speech features can be fused with the hot word coding data corresponding to each target hot word to obtain the first fusion data, and based on the first fusion data, the speech representation corresponding to the speech data is determined. Finally, based on the speech representation, the speech recognition processing is performed on the speech data through the speech large model to obtain the speech recognition result of the speech data. In this way, additional hot word encoding processing is added. First, the hot word text is converted into hot word coding data (such as embedding vectors, etc.) in a high-dimensional space through hot word encoding processing, and then the hot word coding data is fused with the speech coding data of the speech data, so that the set hot word text data can be explicitly used, and accurate recognition can be performed for specific hot words. Furthermore, it is beneficial to the use of the speech large model in specific scenarios and specific vocabulary, and can greatly improve the accuracy of speech recognition and the upper limit of the speech recognition system.

[0026] In practical applications, the user can input the user's requirement in text form to instruct the speech large model to input the result that meets the user's requirement, such as Figure 3 as shown, and specifically, the processing in step S110 and step S112 below can be referred to.

[0027] In step S110, the first text data recording the user's requirement is received.

[0028] Among them, user requirements can include various types. For example, user requirements can include information related to how to process the speech to be recognized. Specifically, for example, recognizing the aforementioned speech and extracting the purpose expressed by the speech, etc., which can be specifically set according to the actual situation.

[0029] In implementation, the above speech recognition page can also include an input box for user requirements. The user can trigger the above input button and can input speech data using the above speech data input method. At this time, the input speech data does not need to contain information related to user requirements. In addition, the user can also determine the hot words (i.e., target hot words) contained in the speech data and can input the above target hot words in the input box of the hot word text data. Additionally, the user can also input information about user requirements in the input box for user requirements. After the input is completed, the user can click the confirm button. At this time, the user's terminal device can obtain the speech data input by the user, obtain the hot word text data input by the user that records the target hot words contained in the speech data from the input box of the hot word text data, and obtain the information about user requirements input by the user from the input box for user requirements. The speech data, hot word text data, and information about user requirements can be sent to the server, and the server can receive the speech data input by the user, the hot word text data input by the user that records the target hot words contained in the speech data, and the information about user requirements (i.e., the first text data).

[0030] In step S112, the first text data is encoded to obtain the requirement encoding data corresponding to the first text data.

[0031] In implementation, the first text data can be encoded by a specified encoder or encoding network model to obtain the requirement encoding data corresponding to the first text data. Among them, the encoding network model can be a model constructed by a machine learning network. For example, an encoding network model can be constructed by a convolutional neural network, or an encoding network model can be constructed by a recurrent neural network model, or an encoding network model can be constructed based on BERT, etc., which can be specifically set according to the actual situation. The first text data is encoded by the above encoder or encoding network model to obtain the requirement encoding data corresponding to the first text data (specifically, it can be represented in the form of a vector or matrix, etc.).

[0032] Based on the processing of the above steps S110 and S112, the specific processing of the above step S108 can include: inputting the speech representation and requirement encoding data into a speech large model to obtain a speech recognition result that meets the user's requirements.

[0033] In implementation, the voice representation and the requirement encoding data can be mixed and input into the voice large model. Alternatively, the voice representation and the requirement encoding data can be concatenated, and the concatenated data can be input into the voice large model, etc. Specifically, it can be set according to the actual situation. The voice representation and the requirement encoding data can be input into the voice large model. In addition to performing voice recognition processing on the voice data through the voice representation and outputting the voice recognition result of the voice data, the voice large model can also, according to the information of the user's requirement, finally output the voice recognition result that meets the user's requirement. Specifically, it can be set according to the actual situation, and the embodiments of this specification do not limit this.

[0034] In practical applications, the specific processing method for determining the voice representation corresponding to the voice data based on the first fusion data in step S106 above can be various. Hereinafter, another optional processing method is provided. For example, Figure 4 as shown, it can specifically include the processing of the following step S1062 and step S1064.

[0035] In step S1062, perform conversion processing on the first fusion data to obtain first conversion data.

[0036] In implementation, the first fusion data can be converted in various ways. For example, a conversion algorithm can be preset in advance, and the first fusion data can be converted through this conversion algorithm. Alternatively, a corresponding network model can be pre-constructed and trained (specifically, such as a convolutional neural network model or a recurrent neural network model, etc., which can be specifically set according to the actual situation), and the first fusion data can be converted through this network model, etc.

[0037] In step S1064, fuse the first conversion data with the hot word encoding data corresponding to each target hot word to obtain second fusion data, and determine the second fusion data as the voice representation corresponding to the voice data.

[0038] In implementation, the first conversion data and the hot word encoding data corresponding to each target hot word can be fused in various ways. For example, a fusion algorithm can be preset in advance, and the first conversion data and the hot word encoding data corresponding to each target hot word can be fused through this fusion algorithm. Or, a corresponding network model can be pre-constructed and trained (specifically, such as a convolutional neural network model or a recurrent neural network model, etc., which can be set according to the actual situation), and the first conversion data and the hot word encoding data corresponding to each target hot word can be fused through this network model, etc. In the process of fusing the first conversion data and the hot word encoding data corresponding to each target hot word, the first conversion data and the hot word encoding data corresponding to each target hot word can be fused in various fusion ways. For example, corresponding weights can be set for the first conversion data and the hot word encoding data corresponding to each target hot word. Since the speech data contains target hot words and may also include words other than the target hot words, the first conversion data will be more than the hot word encoding data corresponding to multiple target hot words. For the same word (i.e., the target hot word), a higher weight can be set for the hot word encoding data corresponding to this target hot word, and a lower weight can be set for the data in the first conversion data corresponding to this target hot word. Or, the data in the first conversion data corresponding to this target hot word can be directly replaced with the hot word encoding data corresponding to this target hot word, etc., which can be set according to the actual situation. Through the above processing, the first conversion data and the hot word encoding data corresponding to each target hot word can be fused to obtain the second fusion data, and the second fusion data is determined as the speech representation corresponding to the speech data.

[0039] In practical applications, the above first speech feature includes FBank (Filter Bank) features. FBank features are obtained by applying a set of triangular filters at different frequency bandwidths to extract corresponding speech features from speech data. These filters are usually distributed according to the Mel Scale. The Mel Scale is a frequency scale based on human auditory perception, which simulates the sensitivity of the human ear to different frequencies. The extraction of Fbank features first requires pre-emphasis and frame segmentation. The speech data is pre-emphasized, and then the speech data is segmented into short-time frames. A window function and a fast Fourier transform (briefly speaking, a window function (such as a Hamming window) is applied to each frame) are applied to the short-time frames, and then a fast Fourier transform is performed to obtain spectral information. Through the Mel filter bank processing, a set of Mel filters is applied to the result of the Fourier transform. Each filter corresponds to a specific frequency range. Through energy calculation, the energy or power of the output of each filter is calculated to obtain the Fbank features.

[0040] In practical applications, a speech recognition system can be set, such as Figure 5As shown, the speech recognition system may include a first network model, a hotword encoding network model, a speech large model, etc. The first network model may further include a feature extraction layer, a speech encoder, a first fusion layer, a characterization sub-model, etc. Specifically, it may include the following content: The processing of step S104 may include: extracting speech features from the speech data through the feature extraction layer in the pre-trained first network model to obtain the first speech features corresponding to the speech data, and encoding the first speech features through the speech encoder in the first network model to obtain the speech encoding data corresponding to the first speech features; encoding each target hotword in the hotword text data through the pre-trained hotword encoding network model to obtain the hotword encoding data corresponding to each target hotword, and inputting the hotword encoding data corresponding to each target hotword into the first network model.

[0041] The processing of step S106 may include: fusing the speech encoding data corresponding to the first speech features with the hotword encoding data corresponding to each target hotword through the first fusion layer in the first network model to obtain the first fusion data, and based on the first fusion data, determining the speech characterization corresponding to the speech data through the characterization sub-model in the first network model.

[0042] Among them, the first network model may be a model constructed through a machine learning network. For example, the first network model may be constructed through a convolutional neural network, or may be constructed through a recurrent neural network model, or may be constructed based on a Transformer module, etc., which can be specifically set according to the actual situation. The feature extraction layer may be constructed by one or more different specified networks. For example, the feature extraction layer may be constructed through a convolutional neural network, or may be constructed through a recurrent neural network, etc., which can be specifically set according to the actual situation. The hotword encoding network model may also be a model constructed through a machine learning network. For example, the hotword encoding network model may be constructed through a convolutional neural network, or may be constructed through a recurrent neural network model, or may be constructed based on a Transformer module, etc., which can be specifically set according to the actual situation. The first fusion layer may also be constructed by one or more different specified networks. For example, the first fusion layer may be constructed through a convolutional neural network, etc., which can be specifically set according to the actual situation. The characterization sub-model may also be a sub-model for determining the speech characterization corresponding to the speech data based on the first fusion data. The characterization sub-model may also be a sub-model constructed through a machine learning network. For example, the characterization sub-model may be constructed through a convolutional neural network, or may be constructed through a recurrent neural network, or may be constructed based on a Transformer module, etc., which can be specifically set according to the actual situation.

[0043] For the specific processing of the above steps S104 and S106, reference can be made to the relevant content mentioned above, and details will not be elaborated here.

[0044] Through the processing of the above steps S104 and S106, an additional hot word encoding network model can be added based on the speech large model, which is used to encode the target hot words in the speech data into corresponding hot word encoding data. Then, the hot word encoding data of the target hot words is fused into the speech large model to accurately identify specific hot words, improving the accuracy of speech recognition and the upper limit of the speech recognition system.

[0045] In practical applications, based on Figure 5 the system structure as shown in Figure 6 it is also possible to receive the first text data recording the user's requirements input by the user. Then, the first text data can be encoded through a text encoder to obtain the requirement encoding data corresponding to the first text data.

[0046] In practical applications, the specific processing method of determining the speech representation corresponding to the speech data through the representation submodel in the first network model based on the above first fusion data can be various. Hereinafter, an optional processing method is provided, that is, the representation submodel can include a conversion layer and a second fusion layer, as shown in Figure 7 and specifically can include the processing of the following steps A2 and A4.

[0047] In step A2, the first fusion data is converted through the conversion layer in the representation submodel to obtain the first conversion data.

[0048] In step A4, the first conversion data is fused with the hot word encoding data corresponding to each target hot word through the second fusion layer in the representation submodel to obtain the second fusion data, and the second fusion data is determined as the speech representation corresponding to the speech data.

[0049] For the specific processing of the above steps A2 and A4, reference can be made to the relevant content mentioned above, and details will not be elaborated here.

[0050] In practical applications, the speech large model, the first network model, and the hot word encoding network model can be jointly trained in the following manner. Specifically, reference can be made to the processing of the following steps B02 to B12.

[0051] In step B02, a first speech data sample, the label information corresponding to the first speech data sample, and a hot word text sample recording the first hot word included in the first speech data sample are obtained.

[0052] In step B04, the speech features in the first speech data sample are extracted through the feature extraction layer in the first network model to obtain the first speech feature sample corresponding to the first speech data sample, and the first speech feature sample is encoded through the speech encoder in the first network model to obtain the speech coding data sample corresponding to the first speech feature sample.

[0053] In step B06, each first hot word in the hot word text sample is encoded through the hot word encoding network model to obtain the hot word coding data sample corresponding to each first hot word, and the hot word coding data sample corresponding to each first hot word is input into the first network model.

[0054] In step B08, the speech coding data sample corresponding to the first speech feature sample and the hot word coding data sample corresponding to each first hot word are fused through the first fusion layer in the first network model to obtain the first fusion data sample, and based on the first fusion data sample, the speech representation sample corresponding to the first speech data sample is determined through the representation sub-model in the first network model.

[0055] In step B10, based on the speech representation sample, the first speech data sample is subjected to speech recognition processing through the speech large model to obtain the speech recognition sample of the first speech data sample.

[0056] In step B12, based on the speech recognition sample and the label information corresponding to the first speech data sample, the corresponding loss information is determined through the preset first loss function, and based on this loss information, the model parameters in the speech large model, the first network model, and the hot word encoding network model are adjusted to jointly train the speech large model, the first network model, and the hot word encoding network model until the preset first loss function converges, obtaining the trained speech large model, the trained first network model, and the trained hot word encoding network model.

[0057] The specific processing of the above steps B02 to B12 can refer to the foregoing relevant content and will not be elaborated here.

[0058] In practical applications, the above first loss function includes the cross-entropy loss function.

[0059] In practical applications, before jointly training the speech large model, the first network model, and the hot word encoding network model in steps B02 to B12, a basic speech-large language model (i.e., the speech large model) can also be trained based on an existing large language model. The specific processing can refer to the processing in steps C2 and C4 below.

[0060] In step C2, a pre-trained large language model and a second speech data sample are obtained.

[0061] In step C4, the large language model is trained based on the second speech data sample and a preset second loss function to obtain a speech large model.

[0062] Among them, there can be multiple types of second loss functions, such as cross-entropy loss function or mean squared error loss function, etc., which can be specifically set according to the actual situation.

[0063] The embodiment of this specification provides a speech recognition method. By receiving the speech data input by the user and the hot word text data recorded by the user that contains the target hot words in the speech data, then, the speech features in the speech data can be extracted to obtain the first speech features corresponding to the speech data, and the first speech features are encoded to obtain the speech coding data corresponding to the first speech features. Each target hot word in the hot word text data is encoded to obtain the hot word coding data corresponding to each target hot word. After that, the speech coding data corresponding to the first speech features can be fused with the hot word coding data corresponding to each target hot word to obtain the first fusion data, and based on the first fusion data, the speech representation corresponding to the speech data is determined. Finally, based on the speech representation, the speech data is subjected to speech recognition processing through the speech large model to obtain the speech recognition result of the speech data. In this way, additional hot word encoding processing is added. First, the hot word text is converted into hot word coding data (such as embedding vectors, etc.) in a high-dimensional space through hot word encoding processing, and then the hot word coding data is fused with the speech coding data of the speech data, so that the set hot word text data can be explicitly used and accurate recognition can be performed for specific hot words. Furthermore, it is beneficial to the use of the speech large model in specific scenarios and specific vocabulary, and can greatly improve the accuracy of speech recognition and the upper limit of the speech recognition system.

[0064] The above is the speech recognition method provided by the embodiment of this specification. Based on the same idea, the embodiment of this specification also provides a speech recognition device, as Figure 8 shown.

[0065] The speech recognition device includes: a data receiving module 801, an encoding module 802, a fusion module 803, and a speech recognition module 804, where: The data receiving module 801 receives the speech data input by the user and the hot word text data recorded by the user that contains the target hot words in the speech data; The encoding module 802 extracts the speech features in the speech data to obtain the first speech features corresponding to the speech data, encodes the first speech features to obtain the speech coding data corresponding to the first speech features, and encodes each target hot word in the hot word text data to obtain the hot word coding data corresponding to each target hot word; The fusion module 803 fuses the voice coding data corresponding to the first voice feature with the hot word coding data corresponding to each target hot word to obtain first fusion data, and determines the voice representation corresponding to the voice data based on the first fusion data; The voice recognition module 804 performs voice recognition processing on the voice data through a voice large model based on the voice representation to obtain the voice recognition result of the voice data.

[0066] In the embodiments of this specification, the device further includes: The first text receiving module receives the first text data recorded with the user's needs input by the user; The text coding module performs coding processing on the first text data to obtain the demand coding data corresponding to the first text data; The voice recognition module 804 inputs the voice representation and the demand coding data into the voice large model to obtain the voice recognition result that meets the user's needs.

[0067] In the embodiments of this specification, the fusion module 803 includes: The conversion unit performs conversion processing on the first fusion data to obtain first conversion data; The fusion unit fuses the first conversion data with the hot word coding data corresponding to each target hot word to obtain second fusion data, and determines the second fusion data as the voice representation corresponding to the voice data.

[0068] In the embodiments of this specification, the first voice feature includes FBank features, The coding module 802 includes: The first coding unit extracts the voice feature of the voice data through the feature extraction layer in the pre-trained first network model to obtain the first voice feature corresponding to the voice data, and performs coding processing on the first voice feature through the voice encoder in the first network model to obtain the voice coding data corresponding to the first voice feature; The second coding unit performs coding processing on each target hot word in the hot word text data through the pre-trained hot word coding network model to obtain the hot word coding data corresponding to each target hot word, and inputs the hot word coding data corresponding to each target hot word into the first network model; The fusion module 803 fuses the voice coding data corresponding to the first voice feature with the hot word coding data corresponding to each target hot word through the first fusion layer in the first network model to obtain first fusion data, and determines the voice representation corresponding to the voice data through the representation sub-model in the first network model based on the first fusion data.

[0069] In the embodiments of the present specification, the fusion module 803 includes: A conversion unit that performs conversion processing on the first fusion data through a conversion layer in the characterization sub-model to obtain first conversion data; A fusion unit that performs fusion processing on the first conversion data and the hot word encoding data corresponding to each target hot word through a second fusion layer in the characterization sub-model to obtain second fusion data, and determines the second fusion data as the speech characterization corresponding to the speech data.

[0070] In the embodiments of the present specification, the apparatus further includes: A first sample acquisition module that acquires a first speech data sample, label information corresponding to the first speech data sample, and a hot word text sample recording the first hot words included in the first speech data sample; A first sample encoding module that extracts speech features in the first speech data sample through a feature extraction layer in the first network model to obtain a first speech feature sample corresponding to the first speech data sample, and performs encoding processing on the first speech feature sample through a speech encoder in the first network model to obtain a speech encoding data sample corresponding to the first speech feature sample; A second sample encoding module that performs encoding processing on each first hot word in the hot word text sample through the hot word encoding network model to obtain a hot word encoding data sample corresponding to each first hot word, and inputs the hot word encoding data sample corresponding to each first hot word into the first network model; A sample fusion module that performs fusion processing on the speech encoding data sample corresponding to the first speech feature sample and the hot word encoding data sample corresponding to each first hot word through a first fusion layer in the first network model to obtain a first fusion data sample, and determines a speech characterization sample corresponding to the first speech data sample through the characterization sub-model in the first network model based on the first fusion data sample; A sample speech recognition module that performs speech recognition processing on the first speech data sample through the speech large model based on the speech characterization sample to obtain a speech recognition sample of the first speech data sample; The first training module, based on the speech recognition sample and the label information corresponding to the first speech data sample, determines the corresponding loss information through a preset first loss function, and adjusts the model parameters in the speech large model, the first network model, and the hotword encoding network model based on the loss information to jointly train the speech large model, the first network model, and the hotword encoding network model until the preset first loss function converges, obtaining the trained speech large model, the trained first network model, and the trained hotword encoding network model.

[0071] In the embodiments of this specification, the first loss function includes a cross-entropy loss function.

[0072] In the embodiments of this specification, the device further includes: The second sample acquisition module, which acquires a pre-trained large language model and a second speech data sample; The second training module, which trains the large language model based on the second speech data sample and a preset second loss function to obtain the speech large model.

[0073] The embodiments of this specification provide a speech recognition device. By receiving the speech data input by the user and the hotword text data recording the target hotwords included in the speech data input by the user, then, the speech features in the speech data can be extracted to obtain the first speech features corresponding to the speech data, and the first speech features are encoded to obtain the speech encoding data corresponding to the first speech features. Each target hotword in the hotword text data is encoded to obtain the hotword encoding data corresponding to each target hotword. After that, the speech encoding data corresponding to the first speech features can be fused with the hotword encoding data corresponding to each target hotword to obtain the first fusion data, and based on the first fusion data, the speech representation corresponding to the speech data is determined. Finally, based on the speech representation, the speech data is subjected to speech recognition processing through the speech large model to obtain the speech recognition result of the speech data. In this way, additional hotword encoding processing is added. Through the hotword encoding processing, the hotword text is first converted into hotword encoding data (such as embedding vectors, etc.) in a high-dimensional space, and then the hotword encoding data is fused with the speech encoding data of the speech data, so that the set hotword text data can be explicitly used and accurate recognition can be performed for specific hotwords. Furthermore, it is beneficial to the use of the speech large model in specific scenarios and specific vocabulary, and can greatly improve the accuracy of speech recognition and the upper limit of the speech recognition system.

[0074] The above is the speech recognition device provided by the embodiments of this specification. Based on the same idea, the embodiments of this specification also provide a speech recognition device, such as Figure 9 shown.

[0075] The voice recognition device may provide a terminal device, a server, etc. for the above embodiments.

[0076] Due to different configurations or performances, the voice recognition device may have relatively large differences. It may include one or more processors 901 and a memory 902. One or more application programs or data may be stored in the memory 902. Among them, the memory 902 may be short-term storage or persistent storage. The application programs stored in the memory 902 may include one or more modules (not shown in the figure), and each module may include a series of computer-executable instructions in the voice recognition device. Further, the processor 901 may be set to communicate with the memory 902 and execute a series of computer-executable instructions in the memory 902 on the voice recognition device. The voice recognition device may also include one or more power supplies 903, one or more wired or wireless network interfaces 904, one or more input / output interfaces 905, and one or more keyboards 906.

[0077] Specifically, in this embodiment, the voice recognition device includes a memory and one or more programs. One or more of the programs are stored in the memory, and one or more of the programs may include one or more modules. Each module may include a series of computer-executable instructions in the voice recognition device and is configured to be executed by one or more processors. The one or more programs include the following computer-executable instructions: Receive the voice data input by the user and the hot word text data recorded by the user and containing the target hot words in the voice data; Extract the voice features in the voice data to obtain the first voice features corresponding to the voice data, perform encoding processing on the first voice features to obtain the voice encoding data corresponding to the first voice features, and perform encoding processing on each target hot word in the hot word text data to obtain the hot word encoding data corresponding to each target hot word; Perform fusion processing on the voice encoding data corresponding to the first voice features and the hot word encoding data corresponding to each target hot word to obtain first fusion data, and determine the voice representation corresponding to the voice data based on the first fusion data; Based on the voice representation, perform voice recognition processing on the voice data through a voice large model to obtain the voice recognition result of the voice data.

[0078] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the embodiment of the speech recognition device, since it is basically similar to the method embodiment, the description is relatively simple, and for the relevant parts, reference can be made to the description in the method embodiment.

[0079] An embodiment of this specification provides a speech recognition device. By receiving the speech data input by the user and the hot word text data input by the user and recording the target hot words included in the speech data, then, the speech features in the speech data can be extracted to obtain the first speech features corresponding to the speech data, and the first speech features are encoded to obtain the speech coding data corresponding to the first speech features. Each target hot word in the hot word text data is encoded to obtain the hot word coding data corresponding to each target hot word. After that, the speech coding data corresponding to the first speech features can be fused with the hot word coding data corresponding to each target hot word to obtain the first fusion data, and based on the first fusion data, the speech representation corresponding to the speech data is determined. Finally, based on the speech representation, the speech data is subjected to speech recognition processing through a speech large model to obtain the speech recognition result of the speech data. In this way, additional hot word encoding processing is added. First, the hot word text is converted into hot word coding data (such as embedding vectors, etc.) in a high-dimensional space through hot word encoding processing, and then the hot word coding data is fused with the speech coding data of the speech data, so that the set hot word text data can be explicitly used and accurate recognition can be performed for specific hot words, which is beneficial to the use of the speech large model in specific scenarios and specific vocabulary, and can greatly improve the accuracy of speech recognition and the upper limit of the speech recognition system.

[0080] Further, based on the above Figures 1 to 7 , one or more embodiments of this specification also provide a storage medium for storing computer-executable instruction information. In a specific embodiment, the storage medium can be a USB flash drive, an optical disc, a hard disk, etc. When the computer-executable instruction information stored in the storage medium is executed by a processor, the following process can be realized: Receive the speech data input by the user and the hot word text data input by the user and recording the target hot words included in the speech data; Extract the speech features in the speech data to obtain the first speech features corresponding to the speech data, and encode the first speech features to obtain the speech coding data corresponding to the first speech features. Each target hot word in the hot word text data is encoded to obtain the hot word coding data corresponding to each target hot word; Fuse the voice encoding data corresponding to the first voice feature with the hot word encoding data corresponding to each target hot word to obtain first fusion data, and based on the first fusion data, determine the voice representation corresponding to the voice data; Based on the voice representation, perform voice recognition processing on the voice data through a voice large model to obtain the voice recognition result of the voice data.

[0081] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. The key point of each embodiment is to illustrate the differences from other embodiments. In particular, for the above-mentioned storage medium embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.

[0082] An embodiment of this specification provides a storage medium. By receiving the voice data input by the user and the hot word text data input by the user that records the target hot words included in the voice data, then, the voice features in the voice data can be extracted to obtain the first voice feature corresponding to the voice data, and the first voice feature is encoded to obtain the voice encoding data corresponding to the first voice feature. Each target hot word in the hot word text data is encoded to obtain the hot word encoding data corresponding to each target hot word. After that, the voice encoding data corresponding to the first voice feature can be fused with the hot word encoding data corresponding to each target hot word to obtain first fusion data, and based on the first fusion data, the voice representation corresponding to the voice data is determined. Finally, based on the voice representation, voice recognition processing is performed on the voice data through a voice large model to obtain the voice recognition result of the voice data. In this way, additional hot word encoding processing is added. First, the hot word text is converted into hot word encoding data (such as embedding vectors, etc.) in a high-dimensional space through hot word encoding processing, and then the hot word encoding data is fused with the voice encoding data of the voice data, so that the set hot word text data can be explicitly used, and accurate recognition can be performed for specific hot words. Furthermore, it is beneficial to the use of the voice large model in specific scenarios and specific vocabulary, and can greatly improve the accuracy of voice recognition and the upper limit of the voice recognition system.

[0083] Furthermore, based on the above Figures 1 to 7 , one or more embodiments of this specification also provide a computer program product, including a computer program. When the computer program in this computer program product is executed by a processor, the following processes can be implemented: Receive the voice data input by the user and the hot word text data input by the user that records the target hot words included in the voice data; Extract the speech features in the speech data to obtain the first speech features corresponding to the speech data, and perform encoding processing on the first speech features to obtain the speech encoding data corresponding to the first speech features. Perform encoding processing on each target hot word in the hot word text data to obtain the hot word encoding data corresponding to each target hot word; Fuse the speech encoding data corresponding to the first speech features with the hot word encoding data corresponding to each target hot word to obtain first fusion data, and based on the first fusion data, determine the speech representation corresponding to the speech data; Based on the speech representation, perform speech recognition processing on the speech data through a speech large model to obtain the speech recognition result of the speech data.

[0084] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the above-mentioned embodiment of a computer program product, since it is basically similar to the method embodiment, the description is relatively simple. For the relevant parts, reference can be made to the partial description of the method embodiment.

[0085] An embodiment of this specification provides a computer program product. By receiving the speech data input by the user and the hot word text data recording the target hot words included in the speech data, then, the speech features in the speech data can be extracted to obtain the first speech features corresponding to the speech data, and encoding processing is performed on the first speech features to obtain the speech encoding data corresponding to the first speech features. Encoding processing is performed on each target hot word in the hot word text data to obtain the hot word encoding data corresponding to each target hot word. After that, the speech encoding data corresponding to the first speech features can be fused with the hot word encoding data corresponding to each target hot word to obtain first fusion data, and based on the first fusion data, the speech representation corresponding to the speech data is determined. Finally, based on the speech representation, speech recognition processing is performed on the speech data through a speech large model to obtain the speech recognition result of the speech data. In this way, additional hot word encoding processing is added. First, the hot word text is converted into hot word encoding data (such as embedding vectors, etc.) in a high-dimensional space through hot word encoding processing, and then the hot word encoding data is fused with the speech encoding data of the speech data, so that the set hot word text data can be explicitly used and accurate recognition can be performed for specific hot words. Furthermore, it is beneficial to the use of the speech large model in specific scenarios and specific vocabulary, and can greatly improve the accuracy of speech recognition and the upper limit of the speech recognition system.

[0086] The above description has been made of specific embodiments of the present specification. Other embodiments are within the scope of the appended claims. In some cases, the acts or steps recited in the claims may be performed in a different order than in the embodiments and still achieve the desired result. Additionally, the processes depicted in the figures do not necessarily require the particular order shown or sequential order to achieve the desired result. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0087] In the 1990s, it was quite obvious to distinguish whether an improvement to a technology was an improvement in hardware (e.g., improvement to the circuit structures such as diodes, transistors, switches, etc.) or an improvement in software (improvement to the method flow). However, with the development of technology, many improvements to the method flow today can be regarded as direct improvements to the hardware circuit structure. Almost all designers obtain the corresponding hardware circuit structure by programming the improved method flow into the hardware circuit. Therefore, it cannot be said that an improvement to a method flow cannot be implemented with a hardware entity module. For example, a Programmable Logic Device (PLD) (such as a Field Programmable Gate Array (FPGA)) is such an integrated circuit whose logical function is determined by the user's programming of the device. The designer can program by himself to "integrate" a digital system on a piece of PLD without having to ask the chip manufacturer to design and fabricate a dedicated integrated circuit chip. Moreover, nowadays, instead of manually fabricating the integrated circuit chip, this programming is mostly implemented using "logic compiler" software, which is similar to the software compiler used in program development and writing. The original code before compilation also has to be written in a specific programming language, which is called Hardware Description Language (HDL). And there is not only one type of HDL, but many types, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc. Currently, the most commonly used are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should also be aware that by simply performing a little logical programming on the method flow with the above-mentioned several hardware description languages and programming it into the integrated circuit, it is easy to obtain the hardware circuit that implements the logical method flow.

[0088] The controller can be implemented in any suitable manner. For example, the controller can take the form of, for example, a microprocessor or a processor and a computer-readable medium storing computer-readable program code (such as software or firmware) executable by the (micro)processor, logic gates, switches, an application specific integrated circuit (ASIC), a programmable logic controller, and an embedded microcontroller. Examples of the controller include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320. The memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art also know that in addition to implementing the controller in the form of pure computer-readable program code, it is entirely possible to logically program the method steps to enable the controller to be implemented in the form of logic gates, switches, application specific integrated circuits, programmable logic controllers, embedded microcontrollers, etc. to achieve the same functions. Therefore, such a controller can be considered a hardware component, and the devices included therein for implementing various functions can also be regarded as the structures within the hardware component. Or even, the devices for implementing various functions can be regarded as either software modules for implementing the method or structures within the hardware component.

[0089] The systems, devices, modules, or units illustrated in the above embodiments can be specifically implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer can be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or any combination of these devices.

[0090] For the convenience of description, when describing the above devices, the functions are divided into various units for separate description. Of course, when implementing one or more embodiments of this specification, the functions of each unit can be implemented in the same or multiple software and / or hardware.

[0091] Those skilled in the art should understand that the embodiments of this specification can be provided as a method, a system, or a computer program product. Therefore, one or more embodiments of this specification can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, one or more embodiments of this specification can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program code.

[0092] Embodiments of this specification are described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this specification. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable serial-parallel devices for fraud cases to generate a machine, such that the instructions executed by the processor of the computer or other programmable serial-parallel devices for fraud cases generate means for implementing the functions specified in the flow Figure 1 one flow or multiple flows and / or blocks Figure 1 means for implementing the functions specified in one block or multiple blocks.

[0093] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable serial-parallel devices for fraud cases to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the functions specified in the flow Figure 1 one flow or multiple flows and / or blocks Figure 1 means for implementing the functions specified in one block or multiple blocks.

[0094] These computer program instructions can also be loaded onto a computer or other programmable serial-parallel devices for fraud cases, such that a series of operational steps are executed on the computer or other programmable devices to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable devices provide steps for implementing the functions specified in the flow Figure 1 one flow or multiple flows and / or blocks Figure 1 means for implementing the functions specified in one block or multiple blocks.

[0095] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and a memory.

[0096] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM), and / or non-volatile memory, such as read-only memory (ROM) or flash memory (flash RAM). The memory is an example of computer-readable media.

[0097] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined in this article, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.

[0098] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.

[0099] Those skilled in the art will appreciate that the embodiments of this specification may be provided as methods, systems or computer program products. Therefore, one or more embodiments of this specification may be in the form of a complete hardware embodiment, a complete software embodiment or an embodiment combining software and hardware. Furthermore, one or more embodiments of this specification may be in the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0100] One or more embodiments of the present specification may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. One or more embodiments of the present specification may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules may be located in local and remote computer storage media, including storage devices.

[0101] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and for the relevant parts, reference can be made to the partial description of the method embodiment.

[0102] The above is only the embodiment of this specification and is not used to limit this document. For those skilled in the art, various changes and modifications can be made to this specification. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of this specification shall be included within the scope of the claims of this specification.

Claims

1. A speech recognition method, the method comprising: Receiving speech data input by a user and hot word text data recorded by the user and containing target hot words in the speech data; Extracting speech features in the speech data to obtain first speech features corresponding to the speech data, performing encoding processing on the first speech features to obtain speech encoding data corresponding to the first speech features, and performing encoding processing on each target hot word in the hot word text data to obtain hot word encoding data corresponding to each target hot word; Performing fusion processing on the speech encoding data corresponding to the first speech features and the hot word encoding data corresponding to each target hot word to obtain first fusion data, and determining a speech representation corresponding to the speech data based on the first fusion data; Performing speech recognition processing on the speech data through a speech large model based on the speech representation to obtain a speech recognition result of the speech data.

2. The method according to claim 1, the method further comprising: Receiving first text data recorded by the user and containing user requirements; Performing encoding processing on the first text data to obtain requirement encoding data corresponding to the first text data; The performing speech recognition processing on the speech data through a speech large model based on the speech representation to obtain a speech recognition result of the speech data includes: Inputting the speech representation and the requirement encoding data into the speech large model to obtain a speech recognition result that meets the user requirements.

3. The method according to claim 1 or 2, the determining a speech representation corresponding to the speech data based on the first fusion data includes: Performing conversion processing on the first fusion data to obtain first conversion data; Performing fusion processing on the first conversion data and the hot word encoding data corresponding to each target hot word to obtain second fusion data, and determining the second fusion data as the speech representation corresponding to the speech data.

4. The method according to claim 3, the first speech features include FBank features, The extracting speech features in the speech data to obtain first speech features corresponding to the speech data, performing encoding processing on the first speech features to obtain speech encoding data corresponding to the first speech features, and performing encoding processing on each target hot word in the hot word text data to obtain hot word encoding data corresponding to each target hot word includes: Performing speech feature extraction on the speech data through a feature extraction layer in a pre-trained first network model to obtain first speech features corresponding to the speech data, and performing encoding processing on the first speech features through a speech encoder in the first network model to obtain speech encoding data corresponding to the first speech features; Performing encoding processing on each target hot word in the hot word text data through a pre-trained hot word encoding network model to obtain hot word encoding data corresponding to each target hot word, and inputting the hot word encoding data corresponding to each target hot word into the first network model; Fusing the voice coding data corresponding to the first voice feature with the hot word coding data corresponding to each target hot word to obtain first fusion data, and determining the voice representation corresponding to the voice data based on the first fusion data includes: Fusing the voice coding data corresponding to the first voice feature with the hot word coding data corresponding to each target hot word through the first fusion layer in the first network model to obtain first fusion data, and determining the voice representation corresponding to the voice data through the representation sub-model in the first network model based on the first fusion data.

5. The method according to claim 4, wherein determining the voice representation corresponding to the voice data through the representation sub-model in the first network model based on the first fusion data includes: Performing a conversion process on the first fusion data through the conversion layer in the representation sub-model to obtain first conversion data; Fusing the first conversion data with the hot word coding data corresponding to each target hot word through the second fusion layer in the representation sub-model to obtain second fusion data, and determining the second fusion data as the voice representation corresponding to the voice data.

6. The method according to claim 4, wherein the method further includes: Obtaining a first voice data sample, label information corresponding to the first voice data sample, and a hot word text sample recording the first hot words included in the first voice data sample; Extracting the voice features in the first voice data sample through the feature extraction layer in the first network model to obtain a first voice feature sample corresponding to the first voice data sample, and performing an encoding process on the first voice feature sample through the voice encoder in the first network model to obtain a voice coding data sample corresponding to the first voice feature sample; Performing an encoding process on each first hot word in the hot word text sample through the hot word coding network model to obtain a hot word coding data sample corresponding to each first hot word, and inputting the hot word coding data sample corresponding to each first hot word into the first network model; Fusing the voice coding data sample corresponding to the first voice feature sample with the hot word coding data sample corresponding to each first hot word through the first fusion layer in the first network model to obtain a first fusion data sample, and determining a voice representation sample corresponding to the first voice data sample through the representation sub-model in the first network model based on the first fusion data sample; Performing a voice recognition process on the first voice data sample through the voice large model based on the voice representation sample to obtain a voice recognition sample of the first voice data sample; Based on the speech recognition samples and the label information corresponding to the first speech data samples, determine the corresponding loss information through a preset first loss function, and adjust the model parameters in the speech large model, the first network model, and the hotword encoding network model based on the loss information to jointly train the speech large model, the first network model, and the hotword encoding network model until the preset first loss function converges, so as to obtain the trained speech large model, the trained first network model, and the trained hotword encoding network model.

7. The method according to claim 6, wherein the first loss function comprises a cross-entropy loss function.

8. The method according to claim 7, further comprising: Obtain a pre-trained large language model and second speech data samples; Train the large language model based on the second speech data samples and a preset second loss function to obtain the speech large model.

9. A speech recognition device, the device comprising: A data receiving module, which receives the speech data input by the user and the hotword text data input by the user and recording the target hotwords included in the speech data; An encoding module, which extracts the speech features in the speech data to obtain the first speech features corresponding to the speech data, performs encoding processing on the first speech features to obtain the speech encoding data corresponding to the first speech features, and performs encoding processing on each target hotword in the hotword text data to obtain the hotword encoding data corresponding to each target hotword; A fusion module, which performs fusion processing on the speech encoding data corresponding to the first speech features and the hotword encoding data corresponding to each target hotword to obtain first fusion data, and determines the speech representation corresponding to the speech data based on the first fusion data; A speech recognition module, which performs speech recognition processing on the speech data through a speech large model based on the speech representation to obtain the speech recognition result of the speech data.

10. A speech recognition device, the speech recognition device comprising: A processor; And A memory arranged to store computer-executable instructions, which when executed cause the processor to: Receive the speech data input by the user and the hotword text data input by the user and recording the target hotwords included in the speech data; Extract the speech features in the speech data to obtain the first speech features corresponding to the speech data, perform encoding processing on the first speech features to obtain the speech encoding data corresponding to the first speech features, and perform encoding processing on each target hotword in the hotword text data to obtain the hotword encoding data corresponding to each target hotword; Perform fusion processing on the speech encoding data corresponding to the first speech features and the hotword encoding data corresponding to each target hotword to obtain first fusion data, and determine the speech representation corresponding to the speech data based on the first fusion data; Perform speech recognition processing on the speech data through a speech large model based on the speech representation to obtain the speech recognition result of the speech data.