Speech recognition method, device and equipment

By adding hot word prompt sentences to the pronunciation model, the problem of hot words being unable to accurately recognize hot words is solved, and the accuracy of speech recognition is improved, especially the recognition effect on specific vocabulary.

CN120279899APending Publication Date: 2025-07-08ALIPAY (HANGZHOU) INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510494446.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-18
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

In the prior art, in speech recognition, hot words (such as human names, place names, proper nouns, etc.) cannot be accurately recognized, resulting in a decrease in speech recognition accuracy.

Method used

By constructing prompt sentences, add hot word text to the pronunciation model, generate prompt coded data matching hot words, and perform speech recognition processing based on voice features.

Benefits of technology

It improves the accuracy of speech recognition, especially the accuracy in specific hot word recognition, and improves the overall performance of speech recognition system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279899A_ABST
    Figure CN120279899A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a voice recognition method, device and equipment. The method comprises the steps that voice data input by a user and hot word text data which are input by the user and record target hot words contained in the voice data are received; voice features in the voice data are extracted, a first voice feature corresponding to the voice data is obtained, and voice representation corresponding to the voice data is determined based on the first voice feature; on the basis of each target hot word in the hot word text data, generating a prompt statement matched with the target hot word, and coding the generated prompt statement to obtain prompt coded data corresponding to the generated prompt statement; and based on the voice representation and prompt code data corresponding to the generated prompt statement, performing voice recognition processing on the voice data through a voice large model to obtain a voice recognition result of the voice data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This document relates to the field of computer technology, and in particular, to a speech recognition method, apparatus, and device. Background Art

[0002] With the continuous development of terminal technology and large models, on the premise of ensuring the protection of users' private data, interaction between users or between humans and machines through voice has become an important interaction method. And for interaction between users or between humans and machines through voice, speech recognition is required. When applying a speech large model to actual speech recognition, it is usually encountered that hot words (or keywords, named entities, etc., referring to some special and uncommon nouns, verbs, etc. that appear in speech recognition processing, such as personal names, place names, proper nouns, etc.) cannot be accurately recognized. Therefore, a better speech recognition scheme is needed, especially the speech recognition processing including hot words. Summary of the Invention

[0003] The purpose of the embodiments of this specification is to provide a better speech recognition scheme, especially the speech recognition processing including hot words.

[0004] To achieve the above technical solution, the embodiments of this specification are implemented as follows: A speech recognition method provided by the embodiments of this specification, the method includes: receiving the speech data input by the user and the hot word text data recorded by the user and containing the target hot words in the speech data; extracting the speech features in the speech data to obtain the first speech features corresponding to the speech data, and based on the first speech features, determining the speech representation corresponding to the speech data; generating a prompt statement matching the target hot word for each target hot word in the hot word text data, and performing encoding processing on the generated prompt statement to obtain the prompt encoding data corresponding to the generated prompt statement; based on the speech representation and the prompt encoding data corresponding to the generated prompt statement, performing speech recognition processing on the speech data through a speech large model to obtain the speech recognition result of the speech data.

[0005] A voice recognition device provided by an embodiment of this specification, the device includes: a data acquisition module, which receives the voice data input by the user and the hot word text data input by the user that records the target hot words included in the voice data; a voice representation determination module, which extracts the voice features in the voice data to obtain the first voice feature corresponding to the voice data, and based on the first voice feature, determines the voice representation corresponding to the voice data; a hot word processing module, which generates a prompt statement matching the target hot word based on each target hot word in the hot word text data, and performs encoding processing on the generated prompt statement to obtain the prompt encoding data corresponding to the generated prompt statement; a voice recognition module, which performs voice recognition processing on the voice data through a voice large model based on the voice representation and the prompt encoding data corresponding to the generated prompt statement, and obtains the voice recognition result of the voice data.

[0006] A voice recognition device provided by an embodiment of this specification, the voice recognition device includes: a processor; and a memory arranged to store computer-executable instructions, the executable instructions, when executed, cause the processor to: receive the voice data input by the user and the hot word text data input by the user that records the target hot words included in the voice data; extract the voice features in the voice data to obtain the first voice feature corresponding to the voice data, and based on the first voice feature, determine the voice representation corresponding to the voice data; generate a prompt statement matching the target hot word based on each target hot word in the hot word text data, and perform encoding processing on the generated prompt statement to obtain the prompt encoding data corresponding to the generated prompt statement; perform voice recognition processing on the voice data through a voice large model based on the voice representation and the prompt encoding data corresponding to the generated prompt statement, and obtain the voice recognition result of the voice data.

[0007] An embodiment of this specification also provides a storage medium, which is used to store computer-executable instructions, and the executable instructions, when executed by a processor, implement the following processes: receive the voice data input by the user and the hot word text data input by the user that records the target hot words included in the voice data; extract the voice features in the voice data to obtain the first voice feature corresponding to the voice data, and based on the first voice feature, determine the voice representation corresponding to the voice data; generate a prompt statement matching the target hot word based on each target hot word in the hot word text data, and perform encoding processing on the generated prompt statement to obtain the prompt encoding data corresponding to the generated prompt statement; perform voice recognition processing on the voice data through a voice large model based on the voice representation and the prompt encoding data corresponding to the generated prompt statement, and obtain the voice recognition result of the voice data.

[0008] The embodiments of this specification also provide a computer program product, including a computer program, which when executed by a processor implements the following process: receiving voice data input by a user and hot word text data input by the user and recording the target hot words included in the voice data; extracting voice features in the voice data to obtain first voice features corresponding to the voice data, and determining a voice representation corresponding to the voice data based on the first voice features; generating prompt statements matching the target hot words based on each target hot word in the hot word text data, and performing encoding processing on the generated prompt statements to obtain prompt coding data corresponding to the generated prompt statements; performing voice recognition processing on the voice data through a voice large model based on the voice representation and the prompt coding data corresponding to the generated prompt statements to obtain a voice recognition result of the voice data. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] To more clearly illustrate the technical solutions in the embodiments of this specification or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments recorded in this specification. For those of ordinary skill in the art, other drawings can also be obtained based on these drawings without creative efforts. Figure 1 It is a schematic flowchart of an embodiment of a voice recognition method in this specification; Figure 2 It is a schematic diagram of a voice recognition page in this specification; Figure 3 It is a schematic structural diagram of a voice recognition system in this specification; Figure 4 It is a schematic flowchart of another embodiment of a voice recognition method in this specification; Figure 5 It is a schematic structural diagram of another voice recognition system in this specification; Figure 6 It is a schematic flowchart of yet another embodiment of a voice recognition method in this specification; Figure 7 It is a schematic structural diagram of yet another voice recognition system in this specification; Figure 8 It is a schematic flowchart of yet another embodiment of a voice recognition method in this specification; Figure 9 It is a schematic structural diagram of yet another voice recognition system in this specification; Figure 10 It is a schematic flowchart of yet another embodiment of a voice recognition method in this specification; Figure 11 It is a schematic diagram of a voice recognition device in this specification; Figure 12 This is a schematic diagram of a voice recognition device in this specification. Specific implementation manners

[0010] The embodiments of this specification provide a voice recognition method, apparatus and device.

[0011] In order to enable the personnel in the technical field to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings in the embodiments of this specification. Obviously, the described embodiments are only a part of the embodiments of this specification, rather than all the embodiments. Based on the embodiments in this specification, all other embodiments obtained by those of ordinary skill in the art without making creative efforts shall fall within the protection scope of this specification.

[0012] The embodiments of this specification provide a voice recognition mechanism. When applying a voice large model to actual voice recognition, it is usually encountered that hot words (or keywords, named entities, etc., referring to some special and uncommon nouns, verbs, etc. that appear in voice recognition processing, such as personal names, place names, proper nouns, etc.) cannot be accurately recognized. Therefore, how to perform accurate voice recognition has become an important topic to be studied. Usually, a voice recognition model with a hot word module can be designed, such as the C-LAS (Listen, Attend and Spell) model. Based on the model architecture of the C-LAS model, a complete ASR (Automatic Speech Recognition) model can be trained, or on the basis of an existing ASR model, a language model with a hot word enhancement function can be added for enhanced processing of voice recognition. However, the above language model with a hot word enhancement function will cause performance loss to the original ASR model, resulting in a decrease in the recognition performance of the original general words. This requires both users and model trainers to weigh between the hot word recognition accuracy and the general recognition accuracy. For this reason, a better voice recognition solution is needed, especially the voice recognition processing including hot words. The embodiments of this specification provide an implementable technical solution. By constructing a prompt statement, the hot word text is added to the voice large model. In this way, the set hot words can be explicitly used and accurate recognition can be performed for specific hot words, thereby greatly improving the accuracy of voice recognition. The specific processing can refer to the specific content in the following embodiments.

[0013] Such as Figure 1As shown in the figure, an embodiment of this specification provides a voice recognition method. The execution subject of this method can be a terminal device, a server, etc. The terminal device can be a mobile terminal device such as a mobile phone or a tablet computer, or a computer device such as a laptop or a desktop computer. Alternatively, it can also be an IoT device (specifically, a smart watch, a vehicle-mounted device, etc.). The server can be an independent server or a server cluster composed of multiple servers. The server can be a background server in the financial field or the online shopping field, or a background server of a certain application program. In this embodiment, the server is taken as an example of the execution subject for detailed description. For the case where the execution subject is a terminal device, refer to the following processing of the server case, which will not be elaborated here. The method can specifically include the following steps: In step S102, receive the voice data input by the user and the hot word text data input by the user and recording the target hot words included in the voice data.

[0014] Among them, the user can be any user. In this embodiment, the user can be the user who inputs voice data to request voice recognition. The voice data can include one or more different contents. For example, the voice data can only include the data of the voice to be recognized, or the voice data can not only include the data of the voice to be recognized, but also include the user's demand information, etc. The demand information can include various types. For example, the demand information can include information related to how to process the voice to be recognized. Specifically, recognize the foregoing voice and extract the purpose expressed by the voice, etc., which can be specifically set according to the actual situation. The target hot word can be a hot word included in the voice data. The hot word can be some special and uncommon words that appear in the voice recognition process. These words can be nouns or verbs, etc., specifically such as personal names, place names, proper nouns, etc. In practical applications, the hot word can also be a keyword or a named entity, etc., which can be specifically set according to the actual situation.

[0015] In implementation, when the user needs to recognize a certain voice data, the voice recognition page can be obtained through the terminal device, such as Figure 2As shown, the speech recognition page may include an input trigger button for speech data, an input box for hot word text data, an output box for speech recognition results, a confirm button, a cancel button, etc. The input trigger button can be used to input speech data in multiple different ways. For example, the user presses the input trigger button and keeps it pressed without releasing. At this time, the user can input speech data. After the input is completed, the user releases the pressed input trigger button. Or, the user clicks the input trigger button. At this time, the speech recognition page can prompt the user to start inputting speech data. Then, the user can input speech data according to the prompt. After the input is completed, the user can click the input trigger button again, etc. It can be specifically set according to the actual situation. The user can use the above input trigger button and can input speech data using the above speech data input method. In addition, the user can also determine the hot words (i.e., target hot words) included in the speech data and can input the above target hot words in the input box for hot word text data. After the input is completed, the user can click the confirm button. At this time, the user's terminal device can obtain the speech data input by the user and obtain the hot word text data recorded with the target hot words included in the speech data input by the user from the input box for hot word text data, and can send the speech data and the hot word text data to the server. The server can receive the speech data input by the user and the hot word text data recorded with the target hot words included in the speech data input by the user.

[0016] In step S104, extract the speech features in the above speech data to obtain the first speech features corresponding to the speech data, and based on the first speech features, determine the speech representation corresponding to the speech data.

[0017] Among them, the first speech features can include various types. For example, the first speech features can include user voice characteristics, user pronunciation habit characteristics, etc., and can be specifically set according to the actual situation. The speech representation can be represented in the form of a vector or in the form of a matrix, etc., and can be specifically set according to the actual situation.

[0018] In implementation, such as Figure 3As shown, speech feature extraction of speech data can be performed in a variety of different ways. For example, a speech feature extraction algorithm or a speech feature extraction network model can be preset. The speech data can be subjected to speech feature extraction through the speech feature extraction algorithm or the speech feature extraction network model to obtain the first speech feature corresponding to the speech data. Among them, the speech feature extraction algorithm can be, for example, the Mel Frequency Cepstral Coefficient algorithm, the Linear Predictive Coding algorithm, etc. The speech feature extraction network model can be a model constructed by a machine learning network. For example, a speech feature extraction network model can be constructed through a convolutional neural network, or a speech feature extraction network model can be constructed through a recurrent neural network model, or a speech feature extraction network model can be constructed based on a Transformer module, etc., which can be specifically set according to the actual situation.

[0019] The speech representation corresponding to the speech data can be determined in a variety of ways. For example, the first speech feature can be encoded. Specifically, a specified encoder or encoding network model can be pre-constructed and trained. The encoding network model can be a model constructed by a machine learning network. For example, an encoding network model can be constructed through a convolutional neural network, or an encoding network model can be constructed through a recurrent neural network model, or an encoding network model can be constructed based on BERT, etc., which can be specifically set according to the actual situation. The first speech feature can be encoded through the specified encoder or encoding network model to obtain the encoded data corresponding to the first speech feature. The obtained encoded data can be used as the speech representation corresponding to the speech data, or the obtained encoded data can be further processed (such as data conversion processing, normalization processing, scaling (or standardization) processing, dimension increase or decrease processing, etc.). The processed data can be used as the speech representation corresponding to the speech data, etc., which can be specifically set according to the actual situation, and the embodiments of this specification do not limit this.

[0020] In step S106, a prompt statement matching the target hot word is generated based on each target hot word in the hot word text data, and the generated prompt statement is encoded to obtain the prompt coding data corresponding to the generated prompt statement.

[0021] Among them, the prompt encoding data can be represented in the form of a vector or in the form of a matrix, etc., and can be specifically set according to the actual situation. The prompt statement can be a statement containing one or more target hot words, and the prompt statement can be composed of one or more sub-statements. For example, the prompt statement can be composed of one sub-statement, specifically such as "The biaswords are chamber, cyril.", and the prompt statement can also be composed of multiple sub-statements, specifically such as "The biaswords are chamber, cyril. The words are xxx…” is composed of sub-statements such as "The bias words are chamber, cyril." and "The words are xxx.", and can be specifically set according to the actual situation.

[0022] In implementation, as Figure 3 shown, the prompt statement matching the target hot word can be constructed in various ways. For example, the construction rules of the prompt statement can be preset, and based on the target hot words included in the hot word text data, the prompt statement matching the target hot word can be constructed according to the construction rules of the prompt statement. Or, a large number of hot words can be analyzed, and according to the analysis results, the common characteristics of different types of hot words can be summarized through expert experience, and then the construction rules or templates of the prompt statements for each type of hot word can be constructed. Then, each target hot word in the hot word text data can be analyzed to determine the type to which each target hot word belongs, and the construction rules or templates of the prompt statements for the hot words of this type can be used to generate the prompt statements corresponding to the hot words of this type, so that the prompt statements matching the target hot words can be generated, etc., and can be specifically set according to the actual situation.

[0023] As Figure 3 shown, the generated prompt statement can be encoded. Specifically, an encoder or an encoding network model can be pre-constructed and trained, and the generated prompt statement can be encoded through the specified encoder or encoding network model to obtain the prompt encoding data corresponding to the generated prompt statement. Among them, the encoder or encoding network model used for encoding the first voice feature is different from the encoder or encoding network model used for encoding the generated prompt statement. The encoding network model for encoding the generated prompt statement can be a model constructed by a machine learning network. For example, the encoding network model can be constructed by a convolutional neural network, or the encoding network model can be constructed by a recurrent neural network model, or the encoding network model can be constructed based on BERT, etc., and can be specifically set according to the actual situation.

[0024] In step S108, based on the speech representation and the prompt coding data corresponding to the generated prompt statement, the speech recognition process is performed on the above speech data through a speech large model to obtain the speech recognition result of the speech data.

[0025] Among them, the speech large model is a multimodal large model obtained by improving the large language model. The speech large model is a large model that fuses speech and the large language model, enabling the original large language model to have the ability to process speech.

[0026] In implementation, as Figure 3 shown, the speech representation and the prompt coding data corresponding to the generated prompt statement can be mixed and input into the speech large model, or the speech representation and the prompt coding data corresponding to the generated prompt statement can be spliced, and the spliced data can be input into the speech large model, etc., which can be specifically set according to the actual situation. The speech large model can perform speech recognition processing on the speech data through the speech representation and in combination with the prompt coding data, and finally can output the speech recognition result of the speech data. In addition, if the speech data also includes the user's requirement information, the speech large model can also output corresponding results according to the user's requirement information, which can be specifically set according to the actual situation, and the embodiments of this specification do not limit this.

[0027] The embodiments of this specification provide a speech recognition method. By receiving the speech data input by the user and the hot word text data recorded by the user that contains the target hot words in the speech data, then, the speech features in the speech data can be extracted to obtain the first speech feature corresponding to the speech data, and based on the first speech feature, the speech representation corresponding to the speech data can be determined. After that, a prompt statement matching the target hot word can be generated based on each target hot word in the hot word text data, and the generated prompt statement can be encoded to obtain the prompt coding data corresponding to the generated prompt statement. Finally, based on the speech representation and the prompt coding data corresponding to the generated prompt statement, the speech recognition process is performed on the speech data through a speech large model to obtain the speech recognition result of the speech data. In this way, the processing of additional hot word prompt statements is added. Through the processing of the hot word prompt statements, the hot word text is first converted into corresponding prompt statements, and then the hot word text is added to the speech large model through the constructed prompt statements, so that the set hot word text data can be explicitly used, and accurate recognition can be performed for specific hot words. Furthermore, it is beneficial to the use of the speech large model in specific scenarios and specific vocabulary, and can greatly improve the accuracy of speech recognition and the upper limit of the speech recognition system.

[0028] In practical applications, the user can input the user's requirements in text form to instruct the speech large model to input results that meet the user's requirements, as Figure 4As shown, specific reference can be made to the processing in step S110 and step S112 below.

[0029] In step S110, the first text data recording the user's requirements is received.

[0030] Among them, the user's requirements can include various types. For example, the user's requirements can include information related to how to process the voice to be recognized. Specifically, such as recognizing the aforementioned voice and extracting the purpose expressed by the voice, etc., which can be specifically set according to the actual situation.

[0031] In implementation, the above voice recognition page can also include an input box for the user's requirements. The user can trigger the input by the above trigger key and can input voice data using the above voice data input method. At this time, the input voice data does not need to contain information related to the user's requirements. In addition, the user can also determine the hot words (i.e., target hot words) contained in the voice data and can input the above target hot words in the input box of the hot word text data. In addition, the user can also input information about the user's requirements in the input box of the user's requirements. After the input is completed, the user can click the confirmation key. At this time, the user's terminal device can obtain the voice data input by the user, obtain the hot word text data input by the user recording the target hot words contained in the voice data from the input box of the hot word text data, and obtain the information about the user's requirements input by the user from the input box of the user's requirements. The voice data, hot word text data, and information about the user's requirements can be sent to the server, and the server can receive the voice data input by the user, the hot word text data input by the user recording the target hot words contained in the voice data, and the information about the user's requirements (i.e., the first text data).

[0032] In step S112, the first text data is encoded to obtain the requirement encoding data corresponding to the first text data.

[0033] In implementation, the first text data can be encoded through a specified encoder or encoding network model to obtain the requirement encoding data corresponding to the first text data. Among them, the encoding network model can be a model constructed through a machine learning network. For example, the encoding network model can be constructed through a convolutional neural network, or through a recurrent neural network model, or based on BERT to construct an encoding network model, etc., which can be specifically set according to the actual situation. Through the above encoder or encoding network model, the first text data is encoded to obtain the requirement encoding data corresponding to the first text data (specifically, it can be represented in the form of a vector or matrix, etc.).

[0034] Based on the processing of the above steps S110 and S112, the specific processing of the above step S108 may include: inputting the prompt coding data and the requirement coding data corresponding to the speech representation and the generated prompt sentence into the speech big model to obtain the speech recognition result that meets the user's needs.

[0035] In implementation, Figure 5 As shown, the prompt coding data and the demand coding data corresponding to the speech representation and the generated prompt sentence can be mixed and input into the speech model, or the prompt coding data and the demand coding data corresponding to the speech representation and the generated prompt sentence can be spliced, and the spliced ​​data can be input into the speech model, etc., which can be set according to the actual situation. The prompt coding data and the demand coding data corresponding to the speech representation and the generated prompt sentence can be input into the speech model. In addition to performing speech recognition processing on the speech data through speech representation and combining the prompt coding data to output the speech recognition result of the speech data, the speech model can also output the speech recognition result that meets the user's needs according to the information of the user's needs. It can be set according to the actual situation, and the embodiments of this specification are not limited to this.

[0036] In practical applications, the specific processing process of determining the speech representation corresponding to the speech data based on the first speech feature in the above step S104 can be varied. An optional processing method is provided below, such as Figure 6 As shown, the processing may specifically include the following steps S1042 and S1044.

[0037] In step S1042, the first speech feature is encoded to obtain speech encoding data corresponding to the first speech feature.

[0038] In step S1044, the speech coding data is converted to obtain first converted data, and the first converted data is determined as the speech representation corresponding to the above speech data.

[0039] In implementation, Figure 7 As shown, the speech coding data can be converted and processed in a variety of ways, for example, a conversion algorithm can be pre-set, and the speech coding data can be converted and processed by the conversion algorithm, or a corresponding network model can be pre-built and trained (specifically, a convolutional neural network model or a recurrent neural network model, etc., which can be set according to actual conditions), and the speech coding data can be converted and processed by the network model, etc. The first converted data obtained after the conversion can be determined as the speech representation corresponding to the above-mentioned speech data.

[0040] In practical applications, the above-mentioned first speech feature includes FBank (Filter Bank) features. FBank features are obtained by applying a set of triangular filters at different frequency bandwidths to extract corresponding speech features from speech data. These filters are usually distributed according to the Mel Scale, which is a frequency scale based on human auditory perception and simulates the sensitivity of the human ear to different frequencies. The extraction of FBank features first requires pre-emphasis and frame segmentation. The speech data is pre-emphasized, and then the speech data is segmented into short-time frames. A window function and the Fast Fourier Transform (briefly speaking, a window function (such as a Hamming window) is applied to each frame) are applied to the short-time frames, and then the Fast Fourier Transform is performed to obtain spectral information. Through the processing of the Mel filter bank, a set of Mel filters is applied to the result of the Fourier transform. Each filter corresponds to a specific frequency range. Through energy calculation, the energy or power of the output of each filter is calculated to obtain the FBank features.

[0041] In practical applications, the specific processing method of generating a prompt statement matching the target hot word based on each target hot word in the hot word text data in step S106 above can be various. Here is another optional processing method, as Figure 8 shown, which can specifically include the processing of the following steps S1062 and S1064.

[0042] In step S1062, based on each target hot word in the hot word text data, a prompt statement template matching the target hot word is obtained.

[0043] In implementation, as Figure 9 shown, multiple different prompt statement templates can be preset according to the actual situation. For example, as mentioned above, corresponding prompt statement templates can be set according to the category to which the hot word belongs. For different categories, different prompt statement templates can be set. Among them, the category to which the hot word belongs can be set in various ways. For example, different categories can be set according to the different parts of speech of the hot word (such as nouns, verbs, adjectives, adverbs, prepositions, conjunctions, auxiliary words, etc.), or different categories can also be set according to the different attributes of the hot word (such as personal names, animal names, plant names, etc.); or, corresponding prompt statement templates can be set according to the semantics of the hot word. For different semantics, different prompt statement templates can be set. For hot words with the same or similar semantics, the same prompt statement template can be set, etc. It can be specifically set according to the actual situation.

[0044] Each target hot word in the hot word text data can be analyzed to determine the category to which each target hot word belongs or the semantic information of each target hot word, etc. Based on the above analysis results, a prompt statement template matching the target hot word can be obtained from multiple prompt statement templates. For example, if the category to which a certain target hot word belongs is A, then a prompt statement template corresponding to A can be obtained from multiple prompt statement templates, and the obtained prompt statement template can be used as the prompt statement template matching the target hot word. Or, if the semantic information of a certain target hot word is semantic 1, then a prompt statement template corresponding to semantic 1 can be obtained from multiple prompt statement templates, and the obtained prompt statement template can be used as the prompt statement template matching the target hot word, etc.

[0045] In step S1064, one or more target hot words are respectively selected from the hot word text data, and the selected target hot words are respectively added to the corresponding prompt statement templates to obtain prompt statements matching the target hot words.

[0046] In implementation, for example, in the case of setting corresponding prompt statement templates according to the category of the hot word, one or more target hot words belonging to the same category can be selected from the hot word text data, and the selected target hot words are added to the prompt statement templates corresponding to the category to obtain corresponding prompt statements. One or more target hot words of the above category can be selected again from the hot word text data, and the selected target hot words are added to the prompt statement templates corresponding to the category to obtain corresponding prompt statements until the prompt statements of all the target hot words of this category have been constructed. In addition, for the target hot words of other categories, the above process can also be repeated until the prompt statements of all the target hot words have been constructed. Finally, prompt statements matching the target hot words can be obtained.

[0047] For the case of setting corresponding prompt statement templates according to the semantics of the hot word, one or more target hot words belonging to the same or similar semantics can be selected from the hot word text data, and the selected target hot words are added to the prompt statement templates corresponding to the semantics to obtain corresponding prompt statements. One or more target hot words of the above semantics can be selected again from the hot word text data, and the selected target hot words are added to the prompt statement templates corresponding to the semantics to obtain corresponding prompt statements until the prompt statements of all the target hot words of this semantics have been constructed. In addition, for the target hot words of other semantics, the above process can also be repeated until the prompt statements of all the target hot words have been constructed. Finally, prompt statements matching the target hot words can be obtained, etc., which can be specifically set according to the actual situation.

[0048] In practical applications, the specific processing method of generating a prompt statement matching the target hot word based on each target hot word in the hot word text data in step S106 above can be various. Here is another optional processing method. As Figure 10 shown, it may specifically include the processing of step S1066 and step S1068 below.

[0049] In step S1066, obtain the prompt statement template.

[0050] In implementation, the above processing involves multiple different prompt statement templates. In practical applications, only one prompt statement template can also be set. In this case, when a prompt statement needs to be generated, the prompt statement template can be obtained.

[0051] In step S1068, select one or more target hot words from the hot word text data respectively, and add the selected target hot words to the prompt statement template respectively to obtain one or more prompt statements matching the target hot words.

[0052] In implementation, one or more target hot words can be randomly selected from the hot word text data or selected according to a specified selection rule (such as a selection rule set according to the part of speech or attribute of the hot word, etc.). Add the selected target hot words to the prompt statement template to obtain a prompt statement. Then, one or more target hot words can be randomly selected from the target hot words other than the above-selected target hot words in the hot word text data or selected according to the specified selection rule, and the selected target hot words are added to the prompt statement template to obtain the second prompt statement, and so on, to obtain one or more prompt statements matching the target hot words.

[0053] In practical applications, the speech large model can be trained in the following manner. Specifically, refer to the processing of step A02 to step A12 below.

[0054] In step A02, obtain the first speech data sample, the label information corresponding to the first speech data sample, and the hot word text sample recording the first hot word included in the first speech data sample.

[0055] In practical applications, in addition to obtaining the above data, it is also possible to obtain a first text data sample recording user requirements. Then, the first text data sample can be encoded through a text encoder or a text encoding network model (which can be constructed through a convolutional neural network, or through a recurrent neural network model, or based on a Transformer module, etc., and can be specifically set according to actual situations). The requirement encoding data sample corresponding to the first text data sample is obtained. Subsequently, the speech representation sample, the requirement encoding data sample, and the prompt encoding data sample corresponding to the generated prompt statement sample can be input into a speech large model to perform speech recognition processing on the first speech data sample, and the speech recognition sample of the first speech data sample is obtained. For specific details, reference can be made to the relevant content mentioned above, and details will not be elaborated here.

[0056] In step A04, the speech features in the first speech data sample are extracted to obtain the first speech feature sample corresponding to the first speech data sample, and based on the first speech feature sample, the speech representation sample corresponding to the first speech data sample is determined.

[0057] In implementation, the specific processing process of extracting the speech features in the first speech data sample to obtain the first speech feature sample corresponding to the first speech data sample can be referred to the relevant content mentioned above, and details will not be elaborated here. Based on the first speech feature sample, the specific processing of determining the speech representation sample corresponding to the first speech data sample can be processed in the following manner: The first speech feature sample is encoded to obtain the speech encoding data sample corresponding to the first speech feature sample; the speech encoding data sample is converted to obtain the first conversion data sample, and the first conversion data sample is determined as the speech representation sample corresponding to the above first speech data sample. For the specific processing process, reference can be made to the relevant content mentioned above, and details will not be elaborated here.

[0058] In step A06, based on each first hot word in the hot word text sample, a prompt statement template matching the first hot word is obtained.

[0059] In step A08, one or more first hot words are respectively selected from the hot word text sample, and the selected first hot words are respectively added to the corresponding prompt statement templates to obtain a prompt statement sample matching the first hot word, and the obtained prompt statement sample is encoded to obtain the prompt encoding data sample corresponding to the prompt statement sample.

[0060] In practical applications, in addition to obtaining the prompt statement samples matching the first hot word through the processing of the above-mentioned steps A06 and A08, the prompt statement samples matching the first hot word can also be determined in the following way: Obtain the prompt statement template; select one or more first hot words from the hot word text data samples respectively, and add the selected first hot words to the prompt statement template respectively to obtain one or more prompt statement samples matching the first hot word. The specific processing process can refer to the relevant content mentioned above and will not be elaborated here.

[0061] In step A10, based on the voice characterization sample and the prompt coding data sample corresponding to the generated prompt statement sample, perform voice recognition processing on the first voice data sample through the voice large model to obtain the voice recognition sample of the first voice data sample.

[0062] In step A12, based on the voice recognition sample and the label information corresponding to the first voice data sample, determine the corresponding loss information through the preset first loss function, and adjust the model parameters in the voice large model based on the loss information to fine-tune the voice large model until the preset first loss function converges, so as to obtain the fine-tuned voice large model.

[0063] The specific processing of the above steps A02 to A12 can refer to the relevant content mentioned above and will not be elaborated here.

[0064] In practical applications, the above first loss function includes the cross-entropy loss function.

[0065] In practical applications, before training the voice large model in steps A02 to A12, a basic voice-large language model (i.e., the voice large model) can also be trained based on an existing large language model. The specific processing can refer to the processing of steps B2 and B4 below.

[0066] In step B2, obtain the pre-trained large language model and the second voice data sample.

[0067] In step B4, train the large language model based on the second voice data sample and the preset second loss function to obtain the voice large model.

[0068] Among them, the second loss function can include various types, such as the cross-entropy loss function or the mean square error loss function, etc., and can be specifically set according to the actual situation.

[0069] An embodiment of this specification provides a voice recognition method. By receiving the voice data input by the user and the hot word text data recorded by the user that contains the target hot words in the voice data, then, the voice features in the voice data can be extracted to obtain the first voice features corresponding to the voice data, and based on the first voice features, the voice representation corresponding to the voice data can be determined. After that, for each target hot word in the hot word text data, a prompt statement matching the target hot word can be generated, and the generated prompt statement can be encoded to obtain the prompt coding data corresponding to the generated prompt statement. Finally, based on the voice representation and the prompt coding data corresponding to the generated prompt statement, the voice data can be subjected to voice recognition processing through a voice large model to obtain the voice recognition result of the voice data. In this way, the processing of additional hot word prompt statements is added. First, the hot word text is converted into corresponding prompt statements through the processing of hot word prompt statements, and then the hot word text is added to the voice large model through the constructed prompt statements, so that the set hot word text data can be explicitly used, and accurate recognition can be performed for specific hot words. Furthermore, it is beneficial to the use of the voice large model in specific scenarios and specific vocabulary, and can greatly improve the accuracy of voice recognition and the upper limit of the voice recognition system.

[0070] The above is the voice recognition method provided by the embodiment of this specification. Based on the same idea, the embodiment of this specification also provides a voice recognition device, as Figure 11 shown.

[0071] The voice recognition device includes: a data acquisition module 1101, a voice representation determination module 1102, a hot word processing module 1103, and a voice recognition module 1104, where: The data acquisition module 1101 receives the voice data input by the user and the hot word text data recorded by the user that contains the target hot words in the voice data; The voice representation determination module 1102 extracts the voice features in the voice data to obtain the first voice features corresponding to the voice data, and based on the first voice features, determines the voice representation corresponding to the voice data; The hot word processing module 1103 generates a prompt statement matching the target hot word for each target hot word in the hot word text data, and encodes the generated prompt statement to obtain the prompt coding data corresponding to the generated prompt statement; The voice recognition module 1104 performs voice recognition processing on the voice data through a voice large model based on the voice representation and the prompt coding data corresponding to the generated prompt statement to obtain the voice recognition result of the voice data.

[0072] In the embodiment of this specification, the device further includes: A text acquisition module that receives first text data recording user requirements input by the user; A text encoding module that performs encoding processing on the first text data to obtain requirement encoding data corresponding to the first text data; The speech recognition module 1104 inputs the speech representation, the prompt encoding data corresponding to the generated prompt statement, and the requirement encoding data into the speech large model to obtain a speech recognition result that meets the user requirements.

[0073] In an embodiment of the present specification, the speech representation determination module 1102 includes: An encoding unit that performs encoding processing on the first speech feature to obtain speech encoding data corresponding to the first speech feature; A conversion unit that performs conversion processing on the speech encoding data to obtain first conversion data, and determines the first conversion data as the speech representation corresponding to the speech data.

[0074] In an embodiment of the present specification, the hot word processing module 1103 includes: A first template acquisition unit that acquires a prompt statement template matching the target hot word based on each target hot word in the hot word text data; A first prompt statement determination unit that respectively selects one or more target hot words from the hot word text data, and respectively adds the selected target hot words to the corresponding prompt statement template to obtain a prompt statement matching the target hot word.

[0075] In an embodiment of the present specification, the first speech feature includes FBank features, and the hot word processing module 1103 includes: A second template acquisition unit that acquires a prompt statement template; A second prompt statement determination unit that respectively selects one or more target hot words from the hot word text data, and respectively adds the selected target hot words to the prompt statement template to obtain one or more prompt statements matching the target hot word.

[0076] In an embodiment of the present specification, the apparatus further includes: A first sample acquisition module that acquires a first speech data sample, label information corresponding to the first speech data sample, and a hot word text sample recording the first hot word included in the first speech data sample; A sample representation determination module that extracts the speech feature in the first speech data sample to obtain a first speech feature sample corresponding to the first speech data sample, and determines a speech representation sample corresponding to the first speech data sample based on the first speech feature sample; A sample template acquisition module, which acquires a prompt statement template that matches each first hot word in the hot word text sample; A sample hot word processing module, which respectively selects one or more first hot words from the hot word text sample, and respectively adds the selected first hot words to the corresponding prompt statement templates to obtain prompt statement samples that match the first hot words, and performs encoding processing on the obtained prompt statement samples to obtain prompt coding data samples corresponding to the prompt statement samples; A sample recognition module, which performs speech recognition processing on the first speech data sample through the speech large model based on the speech representation sample and the prompt coding data sample corresponding to the generated prompt statement sample to obtain a speech recognition sample of the first speech data sample; A first training module, which determines corresponding loss information through a preset first loss function based on the speech recognition sample and the label information corresponding to the first speech data sample, and adjusts the model parameters in the speech large model based on the loss information to fine-tune the speech large model until the preset first loss function converges, so as to obtain a fine-tuned speech large model.

[0077] In the embodiments of this specification, the first loss function includes a cross-entropy loss function.

[0078] In the embodiments of this specification, the device further includes: A second sample acquisition module, which acquires a pre-trained large language model and a second speech data sample; A second training module, which trains the large language model based on the second speech data sample and a preset second loss function to obtain the speech large model.

[0079] An embodiment of this specification provides a voice recognition device. By receiving voice data input by a user and hot word text data recorded by the user and containing target hot words in the voice data, then, the voice features in the voice data can be extracted to obtain a first voice feature corresponding to the voice data, and based on the first voice feature, a voice representation corresponding to the voice data can be determined. After that, a prompt statement matching the target hot word can be generated for each target hot word in the hot word text data, and the generated prompt statement can be encoded to obtain prompt coding data corresponding to the generated prompt statement. Finally, based on the voice representation and the prompt coding data corresponding to the generated prompt statement, the voice data can be subjected to voice recognition processing through a voice large model to obtain a voice recognition result of the voice data. In this way, the processing of additional hot word prompt statements is added. First, the hot word text is converted into corresponding prompt statements through the processing of the hot word prompt statements, and then the hot word text is added to the voice large model through the constructed prompt statements, so that the set hot word text data can be explicitly used, and accurate recognition can be performed for specific hot words. Furthermore, it is beneficial to the use of the voice large model in specific scenarios and specific vocabulary, and can greatly improve the accuracy of voice recognition and increase the upper limit of the voice recognition system.

[0080] The above is the voice recognition device provided by the embodiment of this specification. Based on the same idea, the embodiment of this specification also provides a voice recognition device, as Figure 12 shown.

[0081] The voice recognition device may be a terminal device or a server provided in the above embodiment, etc.

[0082] The voice recognition device may vary greatly due to configuration or performance differences, and may include one or more processors 1201 and a memory 1202. One or more application programs or data may be stored in the memory 1202. Among them, the memory 1202 may be short-term storage or persistent storage. The application programs stored in the memory 1202 may include one or more modules (not shown in the figure), and each module may include a series of computer-executable instructions in the voice recognition device. Further, the processor 1201 may be set to communicate with the memory 1202 and execute a series of computer-executable instructions in the memory 1202 on the voice recognition device. The voice recognition device may also include one or more power supplies 1203, one or more wired or wireless network interfaces 1204, one or more input / output interfaces 1205, and one or more keyboards 1206.

[0083] Specifically, in this embodiment, the speech recognition device includes a memory and one or more programs. One or more of the programs are stored in the memory, and one or more of the programs may include one or more modules. Each module may include a series of computer-executable instructions in the speech recognition device and is configured to be executed by one or more processors. The one or more programs include computer-executable instructions for performing the following: Receiving the speech data input by the user and the hot word text data recorded by the user and containing the target hot words in the speech data; Extracting the speech features in the speech data to obtain the first speech features corresponding to the speech data, and determining the speech representation corresponding to the speech data based on the first speech features; Generating a prompt statement that matches each target hot word in the hot word text data, and performing encoding processing on the generated prompt statement to obtain the prompt encoding data corresponding to the generated prompt statement; Based on the speech representation and the prompt encoding data corresponding to the generated prompt statement, performing speech recognition processing on the speech data through a speech large model to obtain the speech recognition result of the speech data.

[0084] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. The key point of each embodiment is to illustrate the differences from other embodiments. In particular, for the speech recognition device embodiment, since it is basically similar to the method embodiment, the description is relatively simple. For the relevant parts, reference can be made to the description of the method embodiment.

[0085] An embodiment of this specification provides a voice recognition device. By receiving the voice data input by the user and the hot word text data input by the user that records the target hot words included in the voice data, then, the voice features in the voice data can be extracted to obtain the first voice feature corresponding to the voice data, and based on the first voice feature, the voice representation corresponding to the voice data can be determined. After that, a prompt statement matching the target hot word can be generated for each target hot word in the hot word text data, and the generated prompt statement can be encoded to obtain the prompt coding data corresponding to the generated prompt statement. Finally, based on the voice representation and the prompt coding data corresponding to the generated prompt statement, the voice data can be subjected to voice recognition processing through a voice large model to obtain the voice recognition result of the voice data. In this way, the processing of additional hot word prompt statements is added. First, the hot word text is converted into corresponding prompt statements through the processing of the hot word prompt statements, and then the hot word text is added to the voice large model through the constructed prompt statements, so that the set hot word text data can be explicitly used, and accurate recognition can be performed for specific hot words. Furthermore, it is beneficial to the use of the voice large model in specific scenarios and specific vocabulary, and can greatly improve the accuracy of voice recognition and the upper limit of the voice recognition system.

[0086] Further, based on the above Figures 1 to 10 , one or more embodiments of this specification also provide a storage medium for storing computer-executable instruction information. In a specific embodiment, the storage medium can be a USB flash drive, an optical disc, a hard disk, etc. When the computer-executable instruction information stored in the storage medium is executed by a processor, the following processes can be realized: Receive the voice data input by the user and the hot word text data input by the user that records the target hot words included in the voice data; Extract the voice features in the voice data to obtain the first voice feature corresponding to the voice data, and based on the first voice feature, determine the voice representation corresponding to the voice data; Generate a prompt statement matching the target hot word for each target hot word in the hot word text data, and encode the generated prompt statement to obtain the prompt coding data corresponding to the generated prompt statement; Based on the voice representation and the prompt coding data corresponding to the generated prompt statement, perform voice recognition processing on the voice data through a voice large model to obtain the voice recognition result of the voice data.

[0087] Each embodiment in this specification is described in a progressive manner. For the identical or similar parts among the embodiments, reference can be made to each other, and the differences between each embodiment and other embodiments are emphasized. In particular, for the above embodiment of a storage medium, since it is basically similar to the method embodiment, the description is relatively simple, and for the relevant parts, reference can be made to the partial description of the method embodiment.

[0088] An embodiment of this specification provides a storage medium. By receiving the voice data input by the user and the hot word text data input by the user that records the target hot words included in the voice data, then, the voice features in the voice data can be extracted to obtain the first voice feature corresponding to the voice data, and based on the first voice feature, the voice representation corresponding to the voice data can be determined. After that, for each target hot word in the hot word text data, a prompt statement matching the target hot word can be generated, and the generated prompt statement can be encoded to obtain the prompt coding data corresponding to the generated prompt statement. Finally, based on the voice representation and the prompt coding data corresponding to the generated prompt statement, the voice data can be subjected to voice recognition processing through a voice large model to obtain the voice recognition result of the voice data. In this way, the processing of additional hot word prompt statements is added. First, the hot word text is converted into corresponding prompt statements through the processing of the hot word prompt statements, and then the hot word text is added to the voice large model through the constructed prompt statements, so that the set hot word text data can be explicitly used, and accurate recognition can be performed for specific hot words. Furthermore, it is beneficial to the use of the voice large model in specific scenarios and specific vocabulary, and can greatly improve the accuracy of voice recognition and the upper limit of the voice recognition system.

[0089] Furthermore, based on the above Figures 1 to 10 , one or more embodiments of this specification also provide a computer program product, including a computer program. When the computer program in the computer program product is executed by a processor, the following processes can be implemented: Receive the voice data input by the user and the hot word text data input by the user that records the target hot words included in the voice data; Extract the voice features in the voice data to obtain the first voice feature corresponding to the voice data, and based on the first voice feature, determine the voice representation corresponding to the voice data; Based on each target hot word in the hot word text data, generate a prompt statement matching the target hot word, and encode the generated prompt statement to obtain the prompt coding data corresponding to the generated prompt statement; Based on the voice representation and the prompt coding data corresponding to the generated prompt statement, perform voice recognition processing on the voice data through a voice large model to obtain the voice recognition result of the voice data.

[0090] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other, and the differences between each embodiment and other embodiments are emphasized. In particular, for the above-mentioned embodiment of a computer program product, since it is basically similar to the method embodiment, the description is relatively simple, and reference can be made to the relevant part of the method embodiment for the relevant content.

[0091] An embodiment of this specification provides a computer program product. By receiving the voice data input by the user and the hot word text data recording the target hot words included in the voice data, then, the voice features in the voice data can be extracted to obtain the first voice feature corresponding to the voice data, and based on the first voice feature, the voice representation corresponding to the voice data can be determined. After that, for each target hot word in the hot word text data, a prompt statement matching the target hot word can be generated, and the generated prompt statement can be encoded to obtain the prompt coding data corresponding to the generated prompt statement. Finally, based on the voice representation and the prompt coding data corresponding to the generated prompt statement, the voice data can be processed by a voice large model for voice recognition to obtain the voice recognition result of the voice data. In this way, the processing of additional hot word prompt statements is added. First, the hot word text is converted into the corresponding prompt statement through the processing of the hot word prompt statement, and then the hot word text is added to the voice large model through the constructed prompt statement, so that the set hot word text data can be explicitly used and accurate recognition can be performed for specific hot words. Furthermore, it is beneficial to the use of the voice large model in specific scenarios and specific vocabulary, and can greatly improve the accuracy of voice recognition and the upper limit of the voice recognition system.

[0092] The above describes specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in a different order than in the embodiments and still achieve the desired result. Additionally, the processes depicted in the figures do not necessarily require the specific order or sequential order shown to achieve the desired result. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0093] In the 1990s, improvements to a technology could be clearly distinguished as either hardware improvements (e.g., improvements to circuit structures such as diodes, transistors, switches, etc.) or software improvements (improvements to method flows). However, with the development of technology, many method flow improvements today can be regarded as direct improvements to hardware circuit structures. Designers almost always obtain the corresponding hardware circuit structure by programming the improved method flow into the hardware circuit. Therefore, it cannot be said that an improvement to a method flow cannot be implemented using a hardware entity module. For example, a Programmable Logic Device (PLD) (such as a Field Programmable Gate Array (FPGA)) is an integrated circuit whose logical function is determined by the user programming the device. Designers can program themselves to "integrate" a digital system onto a single PLD, without having to ask a chip manufacturer to design and fabricate a dedicated integrated circuit chip. Moreover, nowadays, instead of manually fabricating integrated circuit chips, this programming is mostly implemented using "logic compiler" software, which is similar to the software compilers used in program development and writing. The original code before compilation also has to be written in a specific programming language, which is called a Hardware Description Language (HDL). There is not just one type of HDL, but many types, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc. The most commonly used ones currently are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should also be aware that by simply performing a little logical programming on the method flow using the above-mentioned several hardware description languages and programming it into an integrated circuit, it is easy to obtain the hardware circuit that implements the logical method flow.

[0094] The controller can be implemented in any suitable manner. For example, the controller can take the form of, for example, a microprocessor or a processor and a computer-readable medium storing computer-readable program code (such as software or firmware) executable by the (micro)processor, logic gates, switches, an application specific integrated circuit (ASIC), a programmable logic controller, and an embedded microcontroller. Examples of the controller include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320. The memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art also know that, in addition to implementing the controller in the form of pure computer-readable program code, it is entirely possible to make the controller implement the same function in the form of logic gates, switches, application specific integrated circuits, programmable logic controllers, and embedded microcontrollers by logically programming the method steps. Therefore, such a controller can be considered a hardware component, and the devices included therein for implementing various functions can also be regarded as structures within the hardware component. Or even, the devices for implementing various functions can be regarded as either software modules for implementing the method or structures within the hardware component.

[0095] The systems, devices, modules, or units illustrated in the above embodiments can be specifically implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer can be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or any combination of these devices.

[0096] For the convenience of description, the above devices are described by dividing them into various units according to their functions. Of course, when implementing one or more embodiments of this specification, the functions of each unit can be implemented in the same or multiple software and / or hardware.

[0097] Those skilled in the art should understand that the embodiments of this specification can be provided as a method, a system, or a computer program product. Therefore, one or more embodiments of this specification can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, one or more embodiments of this specification can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program code.

[0098] Embodiments of this specification are described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this specification. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processors of general-purpose computers, special-purpose computers, embedded processors, or other programmable serial and parallel devices for fraud cases to produce a machine, such that the instructions executed by the processors of the computer or other programmable serial and parallel devices for fraud cases produce means for implementing the functions specified in one Figure 1 flow or multiple flows and / or blocks Figure 1 or means for implementing the functions specified in one block or multiple blocks.

[0099] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable serial and parallel devices for fraud cases to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including instruction means that implement the functions specified in one Figure 1 flow or multiple flows and / or blocks Figure 1 or means for implementing the functions specified in one block or multiple blocks.

[0100] These computer program instructions can also be loaded onto a computer or other programmable serial and parallel devices for fraud cases, such that a series of operation steps are executed on the computer or other programmable devices to produce a computer-implemented process, so that the instructions executed on the computer or other programmable devices provide steps for implementing the functions specified in one Figure 1 flow or multiple flows and / or blocks Figure 1 or means for implementing the functions specified in one block or multiple blocks.

[0101] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.

[0102] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM), and / or non-volatile memory, such as read-only memory (ROM) or flash memory (flash RAM). The memory is an example of computer-readable media.

[0103] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined in this article, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.

[0104] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.

[0105] Those skilled in the art will appreciate that the embodiments of this specification may be provided as methods, systems or computer program products. Therefore, one or more embodiments of this specification may be in the form of a complete hardware embodiment, a complete software embodiment or an embodiment combining software and hardware. Furthermore, one or more embodiments of this specification may be in the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0106] One or more embodiments of the present specification may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. One or more embodiments of the present specification may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules may be located in local and remote computer storage media, including storage devices.

[0107] The various embodiments in this specification are described in a progressive manner. For the parts that are the same or similar among the various embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the system embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and for the relevant parts, reference can be made to the partial description of the method embodiments.

[0108] The above is only the embodiment of this specification and is not used to limit this document. For those skilled in the art, various changes and modifications can be made to this specification. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of this specification shall be included within the scope of the claims of this specification.

Claims

1. A speech recognition method, the method comprising: Receiving speech data input by a user and hot word text data input by the user and recording target hot words included in the speech data; Extracting speech features in the speech data to obtain first speech features corresponding to the speech data, and determining a speech representation corresponding to the speech data based on the first speech features; Generating a prompt statement matching each target hot word in the hot word text data, and performing encoding processing on the generated prompt statement to obtain prompt encoding data corresponding to the generated prompt statement; Performing speech recognition processing on the speech data through a speech large model based on the speech representation and the prompt encoding data corresponding to the generated prompt statement to obtain a speech recognition result of the speech data.

2. The method according to claim 1, the method further comprising: Receiving first text data input by the user and recording user requirements; Performing encoding processing on the first text data to obtain requirement encoding data corresponding to the first text data; The performing speech recognition processing on the speech data through a speech large model based on the speech representation and the prompt encoding data corresponding to the generated prompt statement to obtain a speech recognition result of the speech data includes: Inputting the speech representation, the prompt encoding data corresponding to the generated prompt statement, and the requirement encoding data into the speech large model to obtain a speech recognition result that meets the user requirements.

3. The method according to claim 1 or 2, the determining the speech representation corresponding to the speech data based on the first speech features includes: Performing encoding processing on the first speech features to obtain speech encoding data corresponding to the first speech features; Performing conversion processing on the speech encoding data to obtain first conversion data, and determining the first conversion data as the speech representation corresponding to the speech data.

4. The method according to claim 3, the generating a prompt statement matching each target hot word in the hot word text data includes: Based on each target hot word in the hot word text data, obtaining a prompt statement template matching the target hot word; Respectively selecting one or more target hot words from the hot word text data, and respectively adding the selected target hot words to the corresponding prompt statement templates to obtain prompt statements matching the target hot words.

5. The method according to claim 4, wherein the first speech features include FBank features, and the generating a prompt statement matching each target hot word in the hot word text data includes: Obtaining a prompt statement template; Respectively selecting one or more target hot words from the hot word text data, and respectively adding the selected target hot words to the prompt statement template to obtain one or more prompt statements matching the target hot words.

6. The method according to claim 4, the method further comprising: Obtain a first speech data sample, label information corresponding to the first speech data sample, and a hot word text sample recording the first hot words included in the first speech data sample; Extract the speech features in the first speech data sample to obtain a first speech feature sample corresponding to the first speech data sample, and based on the first speech feature sample, determine a speech representation sample corresponding to the first speech data sample; Based on each first hot word in the hot word text sample, obtain a prompt statement template matching the first hot word; Select one or more first hot words from the hot word text sample respectively, and add the selected first hot words to the corresponding prompt statement templates respectively to obtain prompt statement samples matching the first hot words, and perform encoding processing on the obtained prompt statement samples to obtain prompt coding data samples corresponding to the prompt statement samples; Based on the speech representation sample and the prompt coding data sample corresponding to the generated prompt statement sample, perform speech recognition processing on the first speech data sample through the speech large model to obtain a speech recognition sample of the first speech data sample; Based on the speech recognition sample and the label information corresponding to the first speech data sample, determine corresponding loss information through a preset first loss function, and adjust the model parameters in the speech large model based on the loss information to fine-tune the speech large model until the preset first loss function converges, so as to obtain a fine-tuned speech large model.

7. The method according to claim 6, wherein the first loss function comprises a cross-entropy loss function.

8. The method according to claim 7, wherein the method further comprises: Obtain a pre-trained large language model and a second speech data sample; Train the large language model based on the second speech data sample and a preset second loss function to obtain the speech large model.

9. A speech recognition device, the device comprising: A data acquisition module, receiving speech data input by a user and hot word text data recording target hot words included in the speech data input by the user; A speech representation determination module, extracting speech features in the speech data to obtain a first speech feature corresponding to the speech data, and based on the first speech feature, determining a speech representation corresponding to the speech data; A hot word processing module, generating a prompt statement matching the target hot word based on each target hot word in the hot word text data, and performing encoding processing on the generated prompt statement to obtain prompt coding data corresponding to the generated prompt statement; A speech recognition module, based on the speech representation and the prompt coding data corresponding to the generated prompt statement, performing speech recognition processing on the speech data through a speech large model to obtain a speech recognition result of the speech data.

10. A speech recognition device, the speech recognition device comprising: A processor; And A memory arranged to store computer-executable instructions, the executable instructions, when executed, causing the processor: Receive the voice data input by the user and the hot word text data input by the user and recording the target hot words included in the voice data; Extract the voice features in the voice data to obtain the first voice features corresponding to the voice data, and based on the first voice features, determine the voice representation corresponding to the voice data; Generate a prompt statement matching the target hot word based on each target hot word in the hot word text data, and perform encoding processing on the generated prompt statement to obtain the prompt coding data corresponding to the generated prompt statement; Based on the voice representation and the prompt coding data corresponding to the generated prompt statement, perform voice recognition processing on the voice data through a voice large model to obtain the voice recognition result of the voice data.