Speech recognition method and device, electronic equipment and storage medium

By introducing the fine-tuned target large language model (LLM) into the speech recognition model, the shortcomings of existing speech recognition technologies in accuracy and complex task processing are solved, and more efficient speech recognition effects are achieved.

CN119943051APending Publication Date: 2025-05-06BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510012494.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-03
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

Existing speech recognition technologies have challenges in improving accuracy, especially when processing complex tasks and low-resource annotation data, the recognition error rate is high and the model training data is strongly dependent.

Method used

By introducing the target large language model (LLM) into the speech recognition model, the LLM fine-tunes the pre-trained initial LLM through a sample data set to generate a target LLM that can decode speech representations, and unify the target decoding module and the target LLM into a speech recognition model.

Benefits of technology

It improves the accuracy of speech recognition, enhances the model's processing ability on complex tasks, reduces the dependence on high-precision labeled data, and improves the efficiency of speech recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119943051A_ABST
    Figure CN119943051A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a voice recognition method and device, electronic equipment and a storage medium. The method comprises the steps that to-be-recognized voice is acquired; converting the voice to be recognized into a target text through a voice recognition model; wherein the speech recognition model comprises a target coding module and a target decoding module, the target coding module is used for coding the speech to be recognized to obtain speech representation corresponding to the speech to be recognized, and the target decoding module comprises a target large language model LLM obtained by training; the target LLM is used for decoding the voice representation corresponding to the to-be-recognized voice to obtain a target text. According to the embodiment of the invention, the accuracy of speech recognition can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and in particular to a speech recognition method, device, electronic device and storage medium. Background Art

[0002] Automatic speech recognition is a technology that automatically transcribes speech into corresponding text. It can provide support for speech content understanding and human-computer interaction, and is widely used in various basic businesses, such as search, voice assistants, automatic subtitles, etc. At present, with a large amount of manually annotated data, such as more than 50,000 hours of speech data, the accuracy of speech recognition has reached a usable level. However, how to further improve the accuracy of speech recognition is an urgent problem to be solved. Summary of the invention

[0003] The embodiments of the present application disclose a speech recognition method, device, electronic device and storage medium, which can improve the accuracy of speech recognition and improve the efficiency of speech recognition.

[0004] The present application embodiment discloses a speech recognition method, comprising:

[0005] Get the speech to be recognized;

[0006] The speech to be recognized is converted into a target text through a speech recognition model; wherein the speech recognition model includes a target encoding module and a target decoding module, the target encoding module is used to encode the speech to be recognized to obtain a speech representation corresponding to the speech to be recognized, and the target decoding module includes a trained target large language model LLM, and the target LLM is used to decode the speech representation corresponding to the speech to be recognized to obtain the target text.

[0007] In one embodiment, the target LLM is obtained by fine-tuning the initial LLM obtained by pre-training based on the sample data set;

[0008] The sample data set includes speech representations corresponding to a plurality of sample speech and corresponding annotated texts.

[0009] In one embodiment, the sample data set also includes the task prompt information; the method further includes:

[0010] Generate a model fine-tuning instruction corresponding to the target sample speech according to the task prompt information, the speech representation corresponding to the target sample speech, and the corresponding annotated text; the target sample speech is any one of the multiple sample speech;

[0011] Decoding the speech representation corresponding to the target sample speech according to the task prompt information through the initial LLM to obtain the recognition text corresponding to the target sample speech;

[0012] Determine a target loss between a recognized text corresponding to the target sample speech and a corresponding annotated text;

[0013] According to the target losses respectively corresponding to the plurality of sample speech, the parameters of the initial LLM are adjusted until the model convergence condition is met, thereby obtaining the parameters corresponding to the target LLM.

[0014] In one embodiment, the parameters corresponding to the initial LLM include original matrix parameters; the parameters of the initial LLM are adjusted according to the target losses respectively corresponding to the multiple sample speech until the model convergence condition is met to obtain the parameters corresponding to the target LLM, including:

[0015] According to the target losses respectively corresponding to the plurality of sample speech, the target matrix parameters are adjusted; the rank of the target matrix parameters is less than the rank of the original matrix parameters;

[0016] When the model convergence condition is met, the parameters corresponding to the target LLM are obtained according to the adjusted target matrix parameters and the original matrix parameters.

[0017] In one embodiment, the method further comprises:

[0018] The target encoding module is used to encode the plurality of sample speech to obtain speech representations corresponding to the plurality of sample speech respectively.

[0019] In one embodiment, the method further comprises:

[0020] According to the target losses respectively corresponding to the multiple sample speech, the parameters of the target encoding module are adjusted to obtain an updated target encoding module.

[0021] In one embodiment, the speech recognition model further includes a connector, and the connector is used to map the speech representations corresponding to the multiple sample speech respectively to the dimensional space corresponding to the LLM;

[0022] The method further comprises:

[0023] Parameters of the connector are adjusted according to target losses respectively corresponding to the plurality of sample speech sounds.

[0024] In one embodiment, converting the speech to be recognized into target text by using a speech recognition model includes:

[0025] Encoding the speech to be recognized by the target encoding module to obtain a speech representation corresponding to the speech to be recognized;

[0026] Inputting task prompt information, context information, and a speech representation corresponding to the speech to be recognized into the target LLM;

[0027] The target LLM decodes the speech representation corresponding to the speech to be recognized according to the task prompt information and the context information to obtain the target text.

[0028] In one embodiment, the speech recognition model further includes a connector. After encoding the speech to be recognized by the target encoding module to obtain a speech representation corresponding to the speech to be recognized, the method further includes:

[0029] Mapping the speech representation corresponding to the speech to be recognized to the dimensional space corresponding to the LLM through the connector;

[0030] The input task prompt information, context information and speech representation corresponding to the speech to be recognized into the target LLM include:

[0031] The task prompt information, context information and the mapped speech representation are input into the target LLM.

[0032] The present application discloses a speech recognition device, which includes:

[0033] A speech acquisition module, used to acquire the speech to be recognized;

[0034] A speech recognition module, used for converting the speech to be recognized into a target text through a speech recognition model; wherein the speech recognition model includes a target encoding module and a target decoding module, the target encoding module is used for encoding the speech to be recognized to obtain a speech representation corresponding to the speech to be recognized, the target decoding module includes a trained target large language model LLM, the target LLM is used for decoding the speech representation corresponding to the speech to be recognized to obtain the target text.

[0035] In one embodiment, the target LLM is obtained by fine-tuning the initial LLM obtained by pre-training based on a sample data set; the sample data set includes speech representations corresponding to a plurality of sample speech and corresponding annotated texts.

[0036] In one embodiment, the sample data set also includes the task prompt information; the speech recognition device also includes a fine-tuning training module, which is used to generate a model fine-tuning instruction corresponding to the target sample speech according to the task prompt information, the speech representation corresponding to the target sample speech and the corresponding annotation text; the target sample speech is any one of the multiple sample speech; the speech representation corresponding to the target sample speech is decoded according to the task prompt information through the initial LLM to obtain the recognition text corresponding to the target sample speech; the target loss between the recognition text corresponding to the target sample speech and the corresponding annotation text is determined; according to the target losses corresponding to the multiple sample speech respectively, the parameters of the initial LLM are adjusted until the model convergence conditions are met to obtain the parameters corresponding to the target LLM.

[0037] In one embodiment, the parameters corresponding to the initial LLM include original matrix parameters; the fine-tuning training module is also used to adjust the target matrix parameters according to the target losses corresponding to the multiple sample speech respectively; the rank of the target matrix parameters is less than the rank of the original matrix parameters; when the model convergence condition is met, the parameters corresponding to the target LLM are obtained according to the adjusted target matrix parameters and the original matrix parameters.

[0038] In one embodiment, the fine-tuning training module is further used to encode the multiple sample speech through the target encoding module to obtain speech representations corresponding to the multiple sample speech respectively.

[0039] In one embodiment, the fine-tuning training module is further used to adjust parameters of the target encoding module according to the target losses respectively corresponding to the multiple sample speech sounds to obtain an updated target encoding module.

[0040] In one embodiment, the speech recognition model also includes a connector, which is used to map the speech representations corresponding to the multiple sample speech respectively to the dimensional space corresponding to the LLM; the fine-tuning training module is also used to adjust the parameters of the connector according to the target losses corresponding to the multiple sample speech respectively.

[0041] In one embodiment, the speech recognition module is also used to encode the speech to be recognized through the target encoding module to obtain the speech representation corresponding to the speech to be recognized; input task prompt information, context information and the speech representation corresponding to the speech to be recognized into the target LLM; and decode the speech representation corresponding to the speech to be recognized according to the task prompt information and the context information through the target LLM to obtain the target text.

[0042] In one embodiment, the speech recognition model also includes a connector and a speech recognition module, which is also used to map the speech representation corresponding to the speech to be recognized to the dimensional space corresponding to the LLM through the connector; input task prompt information, context information and the mapped speech representation into the target LLM.

[0043] The present application discloses an electronic device, including:

[0044] A memory storing executable program code;

[0045] a processor coupled to the memory;

[0046] The processor calls the executable program code stored in the memory to execute the method described in any one of the above embodiments.

[0047] An embodiment of the present application discloses a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, wherein when the computer program is executed by a processor, the processor executes the method described in any one of the above embodiments.

[0048] Through the speech recognition method, device, electronic device and storage medium disclosed in the embodiments of the present application, the electronic device can obtain the speech to be recognized, and convert the speech to be recognized into the target text through the speech recognition model. Among them, the speech recognition model may include a target encoding module and a target decoding module, the target encoding module can encode the speech to be recognized to obtain a speech representation, and the target decoding module may include a trained target LLM, and the target LLM can decode the speech representation to obtain a target text. By training the LLM, the trained target LLM has the ability to decode the speech representation output by the target encoding module, which solves the problem that the LLM cannot recognize the speech representation, improves the accuracy of speech recognition, and unifies the target decoding module and the target LLM as a target decoding module into a speech recognition model, thereby improving the efficiency of speech recognition. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0050] Figure 1 It is a schematic diagram of an application scenario of a speech recognition method disclosed in an embodiment of the present application;

[0051] Figure 2It is a flowchart of a speech recognition method disclosed in an embodiment of the present application;

[0052] Figure 3 It is a flow chart of a model fine-tuning method disclosed in an embodiment of the present application;

[0053] Figure 4 It is a flow chart of a model fine-tuning method disclosed in an embodiment of the present application;

[0054] Figure 5 is a flow chart of a model fine-tuning method disclosed in an embodiment of the present application;

[0055] Figure 6 It is a flowchart of a speech recognition method disclosed in an embodiment of the present application;

[0056] Figure 7 It is a modular schematic diagram of a speech recognition device disclosed in an embodiment of the present application;

[0057] Figure 8 It is a structural block diagram of an electronic device disclosed in an embodiment of the present application. DETAILED DESCRIPTION

[0058] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0059] It should be noted that the terms "including" and "having" in the embodiments of the present application and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0060] It can be understood that the terms "first", "second", etc. used in the present application can be used in this article to describe various elements, but these elements are not limited by these terms. These terms are only used to distinguish the first element from another element.

[0061] At present, speech recognition models can be divided into two types according to whether they contain a decoder. One type is without a decoder, such as mapping the input speech representation sequence to a text sequence through CTC (Connectionist Temporal Classification), which does not require direct decoding of the speech representation sequence, but has limited modeling capabilities for long sequences and complex tasks. The other type is with a decoder, such as the Transformer that contains an encoder and a decoder. The decoder can autoregressively generate a text sequence based on the speech representation sequence. Autoregression means that the output of each step of the decoder depends on the previous output when generating a sequence. The autoregressive decoder can generate more accurate text sequences with its strong context modeling capabilities. Its modeling capabilities for long sequences and complex tasks are also stronger than those of speech recognition models without a decoder. It is currently the core structure of mainstream language recognition models.

[0062] However, the recognition error rate of the speech recognition model with decoder is still not low. For example, homophone typos will occur when processing Chinese speech, and such errors account for more than 80% of all speech recognition errors. In order to improve the accuracy of speech recognition, the only way is to increase the training data of the decoder. However, the decoder in the speech recognition model usually relies on the output of the encoder for training. The decoder cannot be trained using independent text data. A large amount of high-precision speech-text alignment annotation data is required, such as the characters, words or phonemes corresponding to each frame of the speech. However, high-precision annotation data is insufficient, which limits the training effect of the decoder and the accuracy of speech recognition cannot be improved.

[0063] In order to optimize the text sequence output by the speech recognition model, the related technology also cascades the speech recognition model with the LLM (Large Language Model), and inputs the text sequence output by the speech recognition model into the LLM, that is, the speech sequence is converted into a text sequence using the large language model and then handed over to the LLM for further processing to improve the accuracy of speech recognition. However, this cascade method has the following disadvantages:

[0064] 1. Unable to handle complex tasks: Since the text sequence output by the speech recognition model does not contain other acoustic information besides semantic information, such as intonation information, voice color information, speaking speed information, etc., it is limited in its effect when processing complex tasks. For example, in scenarios where semantic overlap occurs when multiple people speak at the same time, LLM is usually unable to optimize the text sequence.

[0065] 2. The effect of LLM is limited by the effect of the speech recognition model. Since the speech recognition model may make recognition errors, if the speech recognition model recognizes the wrong keywords, it will cause a large deviation in the semantic information. LLM may further increase this deviation and still cannot improve the accuracy of speech recognition when dealing with complex tasks.

[0066] 3. High latency: Since the speech recognition model and LLM are two independent models, the response latency between the models is high.

[0067] The embodiments of the present application disclose a speech recognition method, apparatus, electronic device and storage medium. The target LLM has the ability to recognize speech representation. By using the target LLM as a target decoding module to process the speech representation sequence output by the target encoding module, the inaccurate text data input to the LLM through the traditional speech recognition model is avoided. The target LLM can obtain more acoustic information from the speech representation sequence, thereby improving the accuracy of speech recognition, and the target decoding module and the target LLM are unified into a speech recognition model, thereby improving the efficiency of speech recognition.

[0068] The following is a detailed description with reference to the accompanying drawings.

[0069] like Figure 1 As shown, Figure 1 1 is a schematic diagram of an application scenario of a speech recognition method disclosed in an embodiment of the present application. The application scenario may include an electronic device 110, which may include but is not limited to a mobile phone, a tablet computer, a wearable device, a laptop computer, a PC (Personal Computer), etc.

[0070] In one embodiment, the electronic device 110 may include a speech recognition model, and the electronic device 110 may obtain the speech to be recognized and convert the speech to be recognized into the target text through the speech recognition model. Figure 1 In the example, the speech to be recognized can be input into the electronic device 110, and the electronic device 110 can convert the speech to be recognized into text "Good morning, kids, we meet again" through the speech recognition model.

[0071] In view of the feature that the electronic device 110 can perform voice recognition, a voice recognition function can be set in some application software to facilitate users to complete specific operations through voice. For example, in the scenario where the user uses a mobile phone voice assistant, the user can send control instructions to the mobile phone through voice, and the mobile phone can detect the user's voice and convert the user's voice into target text through a voice recognition model to determine the user's control instruction and respond; in the scenario of meeting recording, the meeting voice of one or more users can be detected, converted into meeting text and recorded.

[0072] like Figure 2 As shown, Figure 2 : is a flow chart of a speech recognition method disclosed in an embodiment of the present application. The speech recognition method can be applied to the electronic device in the above embodiment. The speech recognition method may include the following steps:

[0073] Step 210: Acquire speech to be recognized.

[0074] The speech to be recognized may refer to the original speech signal or speech data to be recognized. In different scenarios, the speech to be recognized may be obtained in different ways. Optionally, in the scenario of real-time speech recognition, the speech to be recognized may be collected in real time by the electronic device through a speech collection device such as a microphone. In the scenario of offline speech processing, the speech to be recognized may be pre-stored in the electronic device or imported into the electronic device.

[0075] When an electronic device detects a speech to be recognized in real time, the electronic device may include a speech detection device and a speech processing device. The speech detection device may be used to collect sound signals outside the electronic device, such as a microphone, etc. The speech processing device may be used to perform analog-to-digital conversion on the collected sound signal to obtain the speech to be recognized, such as an analog-to-digital converter. It can be understood that the collected sound signal may be an analog signal, and the speech to be recognized may be a digital signal.

[0076] Step 220: convert the speech to be recognized into target text through the speech recognition model, and the target decoding module may include the trained target LLM.

[0077] A speech recognition model may be an artificial intelligence model that has the ability to convert speech into text.

[0078] The speech recognition model may include a target encoding module and a target decoding module. The target encoding module may be used to encode the speech to be recognized to obtain a speech representation; the target decoding module may include a trained target LLM, which may be used to decode the speech representation of the speech to be recognized according to the task prompt information to obtain the target text.

[0079] Optionally, the encoding process of the target encoding module for the speech to be recognized may include signal preprocessing, feature extraction, feature mapping, etc., and the target encoding module may also be pre-trained. Signal preprocessing may include performing noise reduction processing on the speech to be recognized to obtain a preprocessed signal, such as filtering out high-frequency signals or low-frequency signals to eliminate noise interference. Feature extraction may refer to extracting signal features of the preprocessed signal, and the signal features may include spectral features or time domain features. Feature mapping may refer to performing multi-level feature extraction and transformation on signal features through a deep neural network to obtain high-dimensional features, and then reducing the dimensionality of the high-dimensional features and mapping them to a low-dimensional vector space to obtain speech representation.

[0080] It can be understood that speech representation is an abstract expression extracted by deep neural networks. Compared with the signal characteristics that directly describe the physical or statistical properties of the speech to be recognized, speech representation can better represent the semantic information of the speech to be recognized, is more suitable for use in the field of speech recognition, and can improve the accuracy of speech recognition.

[0081] In order to enable the trained target LLM to have the ability to decode speech representations, the training process of the target LLM may include but is not limited to one or more of pre-training, self-supervised training, supervised training, fine-tuning, etc. The target LLM can perform semantic analysis on the speech representation to generate a target text.

[0082] Optionally, during the pre-training phase of the target LLM, the electronic device can use unlabeled text data for pre-training, so that the initial LLM obtained by pre-training has the ability to recognize the grammatical structure, semantic rules and contextual relationship of the text data. The electronic device can then use supervised training or fine-tuning methods to enable the target LLM obtained by supervised training to establish a mapping relationship between speech representation and text, so that the speech representation can be decoded through the target LLM to obtain the target text.

[0083] In one embodiment, the target LLM is obtained by fine-tuning the initial LLM obtained by pre-training based on a sample data set, and the sample data set includes speech representations corresponding to a plurality of sample speech and corresponding annotated texts. The speech representations contained in the sample data set correspond to the input of the decoding task, and the annotated texts contained in the sample data set correspond to the output of the decoding task. Fine-tuning refers to the process of further training the initial LLM obtained by pre-training using the sample data set, which can optimize the performance of the target LLM in the decoding task of speech recognition.

[0084] Optionally, fine-tuning of the initial LLM may be to adjust some or all parameters of the initial LLM, and the fine-tuning method may include any one of LoRA (Low-Rank Adaptation), adapter tuning, prefix tuning, etc., without limitation.

[0085] By implementing this embodiment, the target LLM for performing a decoding task of speech recognition can be obtained by fine-tuning the initial LLM obtained by pre-training. Compared with training the decoding task from the beginning, fine-tuning can utilize the knowledge learned by the model during pre-training, thereby speeding up the model training.

[0086] It should be understood that if the target LLM is not guided, the trained target LLM may not realize that the current task is to decode the speech representation when receiving the speech representation, and outputs an erroneous or irrelevant text. Therefore, the electronic device can also input task prompt information to the trained target LLM, and the task prompt information can be used to indicate the task to be performed by the target LLM. The trained target LLM can decode the speech representation according to the task prompt information to obtain the target text. The task prompt information is used to guide the trained target LLM to decode and reduce the uncertainty of the output result of the trained target LLM.

[0087] In addition to the speech recognition scenario of speech-to-text, there can also be recognition scenarios such as speech sentiment analysis and audio time detection. The initial LLM is trained through the sample data set corresponding to each recognition scenario and the preset prompt information to obtain the target LLM corresponding to each recognition scenario. The embodiment of the present application does not limit the function of the target LLM.

[0088] In an embodiment of the present application, an electronic device can obtain a speech to be recognized, and convert the speech to be recognized into a target text through a speech recognition model. Among them, the speech recognition model may include a target encoding module and a target decoding module, the target encoding module can encode the speech to be recognized to obtain a speech representation, and the target decoding module may include a trained target LLM, and the target LLM can decode the speech representation to obtain a target text. By training the LLM, the trained target LLM has the ability to decode the speech representation output by the target encoding module, which solves the problem that the LLM cannot recognize the speech representation, improves the accuracy of speech recognition, and unifies the target decoding module and the target LLM as a target decoding module into a speech recognition model, thereby improving the efficiency of speech recognition.

[0089] like Figure 3 As shown, Figure 3: is a flow chart of a model fine-tuning method disclosed in an embodiment of the present application. The model fine-tuning method can be applied to the electronic device in the above embodiment. The model fine-tuning method may include the following steps:

[0090] Step 310, generating a model fine-tuning instruction corresponding to the target sample speech according to the task prompt information, the speech representation corresponding to the target sample speech, and the corresponding annotated text.

[0091] The electronic device can train the initial LLM based on multiple sample voices. In the embodiment of the present application, the training process of the initial LLM is explained by taking the target sample voice as an example. The target sample voice is any one of the multiple sample voices.

[0092] Task prompt information refers to prompt information that guides and constrains LLM to perform decoding tasks. Optionally, task prompt information can be used to prompt one or more of the content to be recognized, the recognition result to be output, the format requirements of the recognition result, and the context information when performing the task. The embodiment of the present application does not limit the task prompt information. For example, the task identification information can be "describe this speech" or "describe the speech", so that the content that LLM needs to recognize is speech.

[0093] The speech representation corresponding to the target sample speech is generated after the target sample speech is encoded. Optionally, the encoding process corresponding to the target sample speech can be performed by a general encoder, or by a specific encoder specially set for the fine-tuning task. Among them, the general encoding module refers to the one obtained by training based on a large number of speech data sets. The general encoding module has a strong feature extraction capability, and can extract speech features such as spectral features and time domain features from the target sample speech, and generate speech representation after dimensionality reduction processing. The specific encoder refers to the one obtained after training based on a large number of speech data sets and then fine-tuning. The specific encoder can extract speech features that are more in line with the current scene. The current scene refers to the scene where the target decoding module is LLM, thereby enhancing the accuracy and robustness of the target decoding module.

[0094] The annotated text corresponding to the target sample speech may refer to standardized text data corresponding to the target sample speech, for example, may be a manually annotated speech transcription result.

[0095] The electronic device can input the model fine-tuning instruction to the initial LLM to instruct the initial LLM to perform training and parameter adjustment. Among them, the electronic device can jointly process the task prompt information, the speech representation corresponding to the target sample speech, and the corresponding annotation text to obtain the model fine-tuning instruction. Optionally, the task prompt information and the annotation text corresponding to the target sample speech can both be text sequences, and the speech representation corresponding to the target sample speech is a representation sequence. The electronic device can adjust the dimension of the speech representation corresponding to the target sample speech so that the dimension of the speech representation is the same as the dimension of the task prompt information and the annotation text, and then connect the task prompt information, the speech representation corresponding to the target sample speech, and the corresponding annotation text to obtain the model fine-tuning instruction.

[0096] Step 320 , decoding the speech representation corresponding to the target sample speech according to the task prompt information through the initial LLM to obtain the recognition text corresponding to the target sample speech.

[0097] The initial LLM can receive the model fine-tuning instruction, and in response to the model fine-tuning instruction, determine the task prompt information in the model fine-tuning instruction, the speech representation corresponding to the target sample speech, and the corresponding annotated text, so that the speech representation corresponding to the target sample speech can be decoded according to the task prompt information to obtain the recognition text corresponding to the target sample speech. Optionally, during the decoding process, the initial LLM can perform semantic analysis and context understanding on the representation vector corresponding to the target sample speech in combination with the task prompt information to obtain the recognition text corresponding to the target sample speech.

[0098] Step 330, determining a target loss between the recognized text corresponding to the target sample speech and the corresponding annotated text.

[0099] The target loss is used to characterize the difference between the recognized text and the annotated text. The target loss can be determined by the electronic device through a specific loss function. The following is an example of calculating the target loss:

[0100] 1. Target loss based on edit distance. The electronic device can compare the recognized text corresponding to the target sample speech with the annotated text corresponding to the target sample speech, calculate the edit distance corresponding to the target sample speech, and then determine the target loss based on the edit distance. The edit distance refers to the minimum number of editing operations required to convert the recognized text into the annotated text. The editing operations may include insertion, deletion, and replacement operations. The smaller the edit distance, the higher the similarity between the recognized text and the annotated text.

[0101] 2. Target loss based on character error rate or word error rate. The electronic device can compare the recognized text with the annotated text to obtain the character error rate or word error rate of the recognized text, and thus determine the target loss based on the character error rate or word error rate.

[0102] 3. Target loss based on semantic similarity. The electronic device can use a pre-trained semantic model to semantically embed the recognized text and the annotated text, obtain the embedding vector corresponding to the recognized text and the embedding vector corresponding to the annotated text, and then calculate the cosine similarity or Euclidean distance between the embedding vector corresponding to the recognized text and the embedding vector corresponding to the annotated text as the semantic similarity, thereby determining the target loss based on the semantic similarity.

[0103] The target loss calculated by the above method can quantify the performance of the initial LLM in the current task, and provide a basis for adjusting the model parameters in the subsequent steps, thereby optimizing the model performance. However, it should be understood that the above method for calculating the target loss is only an example, and the embodiment of the present application does not limit the method for calculating the target loss.

[0104] Step 340, adjusting the parameters of the initial LLM according to the target losses respectively corresponding to the plurality of sample speech until the model convergence condition is met, thereby obtaining the parameters corresponding to the target LLM.

[0105] Optionally, the electronic device can adjust the parameters of the initial LLM based on a preset fine-tuning method according to the target losses corresponding to the multiple sample voices, wherein the multiple sample voices can refer to all the sample voices in the sample data set, or can refer to part of the sample voices in the sample data set, depending on the parameter update rules in the training process.

[0106] In one example, the electronic device may adjust the parameters of the initial LLM according to the target losses corresponding to all the sample voices in the sample data set, and then retrain according to all the sample voices in the sample data set after the parameter adjustment. In another example, the electronic device may divide all the sample voices in the sample data set into multiple batches, each batch of sample voices includes multiple, and the electronic device may determine the target loss corresponding to each sample voice in each batch of sample voices, and adjust the parameters of the initial LLM, and then retrain according to another batch of sample voices after the parameter adjustment.

[0107] When the model convergence condition is met, the parameters of the initial LLM are updated to obtain the parameters corresponding to the target LLM. The model convergence condition may include one or more of the following: the number of parameter adjustments reaches a number threshold, the number of target losses calculated reaches a number threshold, the target loss is less than a loss threshold, etc.

[0108] In an embodiment of the present application, fine-tuning instructions are generated by combining task prompt information, speech representation and annotated text, and the recognized text is obtained by decoding using the initial large language model, and the target loss between the recognized text and the annotated text is further calculated. Finally, the parameters of the initial LLM are adjusted based on the target loss of multiple sample speech, thereby achieving fine-tuning of the initial LLM. The fine-tuning process can ensure that the target LLM has the ability to decode the speech representation and avoids the performance degradation of the target LLM itself, so that the target text generated according to the speech representation corresponding to the speech to be recognized is more accurate.

[0109] like Figure 4 As shown, Figure 4 : is a flow chart of a model fine-tuning method disclosed in an embodiment of the present application. The model fine-tuning method can be applied to the electronic device in the above embodiment. The model fine-tuning method may include the following steps:

[0110] Step 410: encode the plurality of sample speech sounds through a target encoding module to obtain speech representations corresponding to the plurality of sample speech sounds.

[0111] In the embodiment of the present application, the target encoding module in the speech recognition model is not a general encoder, but a target encoding module designed and optimized specifically for the current speech recognition scenario. The target encoding module is designed in collaboration with the target decoding module. The target encoding module is associated with the speech recognition task after pre-training, and can generate a speech representation suitable for LLM based on the input sample speech, which can ensure that the speech representation matches the decoding requirements, thereby improving the speech recognition accuracy of the speech recognition model.

[0112] In one embodiment, the target encoding module may be pre-trained based on a mapping matrix and a codebook set, the mapping matrix and the codebook set are generated in a random initialization manner, and the mapping matrix and the codebook set remain unchanged during the pre-training process. During the pre-training process of the target coding module, the electronic device can generate a mapping matrix and a codebook set based on a random initialization method, and the codebook set represents the corresponding relationship between the codebook vector and the index; based on the mapping matrix, each sample data in the sample data sequence is respectively subjected to vector mapping processing to obtain the mapping vector of each sample data; the target codebook vector matching each mapping vector is searched from the codebook set, and the target index corresponding to the target codebook vector matching each mapping vector is used as the reference discretized label of the sample data corresponding to each mapping vector; the masked sample data sequence is input into the target coding module to be trained for speech representation processing to obtain speech representation, and the masked sample data sequence is obtained based on masking processing of multiple positions in the sample data sequence; in the speech representation result, the representation results corresponding to each masked position are respectively subjected to discretized label prediction to obtain the predicted discretized label corresponding to each masked position; based on the difference between the predicted discretized label corresponding to each masked position and the reference discretized label of the sample data at the masked position, the parameters of the target coding module to be trained are adjusted until the preset training end condition is met to obtain the target coding module.

[0113] In implementing this embodiment, the randomly initialized mapping matrix and codebook set are fixed throughout the pre-training process, so that the same sample data in the entire pre-training process always corresponds to a uniquely determined discrete label. There is no need to discretize the sample data in advance. Efficient and stable pre-training can be performed once the sample data is prepared, thereby improving the pre-training speed, convergence and stability of the target coding module.

[0114] Discrete label prediction is performed for the representation results of the corresponding masked positions in the speech representation respectively; based on the difference between the predicted discretized labels corresponding to each masked position and the corresponding reference discretized labels, the model parameters of the speech representation model to be trained are adjusted to obtain a pre-trained speech representation model. The present disclosure improves the pre-training speed and stability of the speech representation model.

[0115] Step 420, generating a model fine-tuning instruction corresponding to the target sample speech according to the task prompt information, the speech representation corresponding to the target sample speech, and the corresponding annotated text.

[0116] Step 430 , decoding the speech representation corresponding to the target sample speech according to the task prompt information through the initial LLM to obtain the recognition text corresponding to the target sample speech.

[0117] Step 440, determining a target loss between the recognized text corresponding to the target sample speech and the corresponding annotated text.

[0118] The methods of steps 420-440 may refer to the methods of steps 310-330 in the above embodiment, and will not be described in detail here.

[0119] Step 450, adjusting the target matrix parameters according to the target losses corresponding to the plurality of sample speech respectively; the rank of the target matrix parameters is smaller than the rank of the original matrix parameters.

[0120] Since the matrix parameters of the initial LLM after pre-training are usually over-parameterized, that is, only a small part of all the parameters included in the initial LLM will be used in the decoding task, but not all the parameters will be used in the decoding task. When fine-tuning the initial LLM to adapt it to the decoding task, if the electronic device adjusts all the parameters of the initial LLM, the fine-tuning process consumes more resources and time. Therefore, the target matrix parameters can be inserted into the key network layer in the initial LLM. During the fine-tuning process, only the target matrix parameters are adjusted, while the original matrix parameters of the initial LLM are kept unchanged, thereby reducing the computational complexity, improving the efficiency of model fine-tuning, and effectively retaining the knowledge learned by the LLM during the pre-training process.

[0121] Among them, the key network layers may include the attention layer and the fully connected layer. The target matrix parameters inserted by the initial LLM can be two. For example, the original matrix parameter of the initial LLM is W0, and the parameter corresponding to the target LLM can be W. The fine-tuning process can be expressed as W = W0 + ΔW, W0∈R H×H , ΔW∈R H×H , ΔW is the parameter change in the fine-tuning process. By performing low-rank decomposition on ΔW, the two inserted target matrix parameters can include A and B, A∈R H×R , B∈R H×R , R<<H, and ΔW=A·B T That is, during the fine-tuning process, the original matrix parameter W0 can be left unchanged, but only A and B can be adjusted. In addition, since the adjusted parameters are A and B, in the forward propagation process of the multi-layer network layer included in LLM, the formula of the intermediate state h = W0·x needs to be modified to h = W0·x + A·B T x, where x refers to the input data of the current network layer. For example, the input data of the first network layer may refer to the speech representation corresponding to the target sample speech.

[0122] Step 460, when the model convergence condition is met, the parameters corresponding to the target LLM are obtained according to the adjusted target matrix parameters and the original matrix parameters.

[0123] When the model convergence condition is met, the electronic device can calculate the parameter variable ΔW of the fine-tuning process based on the adjusted target matrix parameters A and B, and then add the parameter variable ΔW to the original matrix parameter W0 to obtain the parameter corresponding to the target LLM.

[0124] Step 470 , adjusting parameters of a target coding module according to target losses corresponding to the plurality of sample speech signals, to obtain an updated target coding module.

[0125] Step 480 , adjusting parameters of the connector according to target losses corresponding to the plurality of sample speech.

[0126] Since the target encoding module and connector are the front-end modules of LLM, the target encoder is used to encode the input multiple sample speech to obtain the speech representations corresponding to the multiple sample speech respectively. The target encoder can input the speech representations corresponding to the multiple sample speech respectively into the connector. The connector can be used to map the speech representations corresponding to the multiple sample speech respectively to the dimensional space corresponding to the LLM. The form of the connector can include one of projection-based, query-based, and fusion-based connectors. During the fine-tuning process of the initial LLM by the electronic device, the parameters of the target encoding module and / or connector can also be adjusted to optimize the entire speech recognition model, realize efficient coordination between encoding, conversion and decoding, and thus improve the performance of the speech recognition model.

[0127] In order to more clearly illustrate the overall flow of the fine-tuning process, Figure 5 As shown, Figure 5 is a flow chart of a model fine-tuning method disclosed in an embodiment of the present application, wherein a sample data set includes task prompt information, a target sample voice, and annotated text corresponding to the target sample voice. The target sample voice can be first input into a target encoding module, and the target encoding module encodes the target sample voice to obtain a voice representation corresponding to the target sample voice. The voice representation corresponding to the target sample voice is then input into a connector, and the connector can reduce the dimension of the voice representation corresponding to the target sample voice, so that the voice representation E after the dimension reduction is s The task prompt information and the target sample speech corresponding to the labeled text can be input into the Tokenizer respectively. The Tokenizer can serialize the task prompt information and the labeled text corresponding to the target sample speech to obtain the serialized task prompt information E p And the serialized annotation text E t , the electronic device then converts E s 、E p and Et The model fine-tuning instructions corresponding to the target sample speech are obtained by splicing, and the model fine-tuning instructions can be input into the target LLM, and the target LLM can output the recognition text corresponding to the target sample speech according to the model fine-tuning instructions, so that the target loss can be calculated according to the recognition text corresponding to the target sample speech and the corresponding annotated text for fine-tuning.

[0128] like Figure 6 As shown, Figure 6 : is a flow chart of a speech recognition method disclosed in an embodiment of the present application. The speech recognition method can be applied to the electronic device in the above embodiment. The speech recognition method may include the following steps:

[0129] Step 610: Acquire speech to be recognized.

[0130] Step 620: Encode the speech to be recognized by a target encoding module to obtain a speech representation corresponding to the speech to be recognized.

[0131] It can be understood that the target encoding module here refers to the target encoding module updated in the above-mentioned model fine-tuning process.

[0132] Step 630 , input task prompt information, context information, and speech representation corresponding to the speech to be recognized into the target LLM.

[0133] Among them, the context information may include scene information corresponding to the speech to be recognized. Optionally, the context information may include various preset common words, such as names, place names, etc. By inputting the context information into the target LLM, the speech recognition accuracy of the target LLM can be improved.

[0134] In one embodiment, in a speech recognition scenario, the speech recognition model may further include a connector. After encoding the speech to be recognized through a target encoding module to obtain a speech representation corresponding to the speech to be recognized, the electronic device may also map the speech representation corresponding to the speech to be recognized to the dimensional space corresponding to the LLM through the connector, so that task prompt information, context information and the mapped speech representation can be input into the target LLM.

[0135] Optionally, the dimensional space corresponding to the LLM may refer to the dimensional space corresponding to the semantic representation, and the connector may refer to the connector whose parameters have also been adjusted during the fine-tuning process of the target LLM, and the connector may reduce the dimensionality of the speech representation corresponding to the speech to be recognized to the dimensional space corresponding to the semantic representation. Compared with the speech representation, the semantic representation can better represent more semantic information in the target sample speech.

[0136] By implementing this embodiment, the speech recognition accuracy of the target LLM can be improved.

[0137] Step 640 , the target LLM decodes the speech representation corresponding to the speech to be recognized according to the task prompt information and the context information to obtain the target text.

[0138] For example, the context information may include a person's name "Zhou xx". If the target LLM does not receive the context information, it may misrecognize "Zhou xx". However, if the target LLM receives the context information, the probability of recognizing "Zhou xx" can be increased.

[0139] In an embodiment of the present application, the electronic device can input task prompt information, context information, and a speech representation corresponding to the speech to be recognized into the target LLM, and can decode the speech representation corresponding to the speech to be recognized according to the task prompt information and the context information through the target LLM to obtain the target text. By providing context information to the target LLM, the accuracy of the target LLM in the speech recognition task can be improved.

[0140] like Figure 7 As shown, Figure 7 7 is a modular schematic diagram of a speech recognition device disclosed in an embodiment of the present application. The speech recognition device 700 can be applied to the electronic device in the above embodiment. The speech recognition device may include a speech acquisition module 710 and a speech recognition module 720, wherein:

[0141] The speech acquisition module 710 is used to acquire the speech to be recognized;

[0142] The speech recognition module 720 is used to convert the speech to be recognized into a target text through a speech recognition model; wherein the speech recognition model includes a target encoding module and a target decoding module, the target encoding module is used to encode the speech to be recognized to obtain a speech representation corresponding to the speech to be recognized, and the target decoding module includes a trained target large language model LLM, and the target LLM is used to decode the speech representation corresponding to the speech to be recognized to obtain a target text.

[0143] In one embodiment, the target LLM is obtained by fine-tuning the initial LLM obtained by pre-training based on a sample data set; the sample data set includes speech representations corresponding to a plurality of sample speech and corresponding annotated texts.

[0144] In one embodiment, the sample data set also includes task prompt information; the speech recognition device also includes a fine-tuning training module, which is used to generate a model fine-tuning instruction corresponding to the target sample speech according to the task prompt information, the speech representation corresponding to the target sample speech, and the corresponding annotation text; the target sample speech is any one of a plurality of sample speech; the speech representation corresponding to the target sample speech is decoded according to the task prompt information through the initial LLM to obtain the recognition text corresponding to the target sample speech; the target loss between the recognition text corresponding to the target sample speech and the corresponding annotation text is determined; according to the target losses corresponding to the plurality of sample speech respectively, the parameters of the initial LLM are adjusted until the model convergence conditions are met to obtain the parameters corresponding to the target LLM.

[0145] In one embodiment, the parameters corresponding to the initial LLM include original matrix parameters; the fine-tuning training module is also used to adjust the target matrix parameters according to the target losses corresponding to multiple sample speech respectively; the rank of the target matrix parameters is less than the rank of the original matrix parameters; when the model convergence conditions are met, the parameters corresponding to the target LLM are obtained according to the adjusted target matrix parameters and the original matrix parameters.

[0146] In one embodiment, the fine-tuning training module is further used to encode the multiple sample speech through the target encoding module to obtain speech representations corresponding to the multiple sample speech respectively.

[0147] In one embodiment, the fine-tuning training module is further used to adjust parameters of the target encoding module according to target losses corresponding to a plurality of sample speech samples, so as to obtain an updated target encoding module.

[0148] In one embodiment, the speech recognition model also includes a connector, which is used to map the speech representations corresponding to multiple sample speech samples to the dimensional space corresponding to the LLM; the fine-tuning training module is also used to adjust the parameters of the connector according to the target losses corresponding to the multiple sample speech samples.

[0149] In one embodiment, the speech recognition module 720 is also used to encode the speech to be recognized through the target encoding module to obtain the speech representation corresponding to the speech to be recognized; input the task prompt information, context information and the speech representation corresponding to the speech to be recognized into the target LLM; and decode the speech representation corresponding to the speech to be recognized according to the task prompt information and the context information through the target LLM to obtain the target text.

[0150] In one embodiment, the speech recognition model also includes a connector, a speech recognition module 720, and is also used to map the speech representation corresponding to the speech to be recognized to the dimensional space corresponding to the LLM through the connector; input task prompt information, context information and the mapped speech representation to the target LLM.

[0151] In an embodiment of the present application, an electronic device can obtain a speech to be recognized, and convert the speech to be recognized into a target text through a speech recognition model. Among them, the speech recognition model may include a target encoding module and a target decoding module, the target encoding module can encode the speech to be recognized to obtain a speech representation, and the target decoding module may include a trained target LLM, and the target LLM can decode the speech representation to obtain a target text. By training the LLM, the trained target LLM has the ability to decode the speech representation output by the target encoding module, which solves the problem that the LLM cannot recognize the speech representation, improves the accuracy of speech recognition, and unifies the target decoding module and the target LLM as a target decoding module into a speech recognition model, thereby improving the efficiency of speech recognition.

[0152] like Figure 8 As shown, in one embodiment, an electronic device is provided, which may include:

[0153] A memory 810 storing executable program code;

[0154] a processor 820 coupled to the memory 810;

[0155] The processor 820 calls the executable program code stored in the memory 810 to implement the speech recognition method provided in the above embodiments.

[0156] The memory 810 may include a random access memory (RAM) or a read-only memory (ROM). The memory 810 may be used to store instructions, programs, codes, code sets or instruction sets. The memory 810 may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for implementing at least one function (such as a touch function, a sound playback function, an image playback function, etc.), instructions for implementing the above-mentioned various method embodiments, etc. The data storage area may also store data created by the electronic device during use, etc.

[0157] The processor 820 may include one or more processing cores. The processor 820 uses various interfaces and lines to connect various parts of the entire electronic device, and executes various functions of the electronic device and processes data by running or executing instructions, programs, code sets or instruction sets stored in the memory 810, and calling data stored in the memory 810. Optionally, the processor 820 can be implemented in at least one hardware form of digital signal processing (DSP), field-programmable gate array (FPGA), and programmable logic array (PLA). The processor 820 can integrate one or a combination of a central processing unit (CPU), a graphics processing unit (GPU), and a modem. Among them, the CPU mainly processes the operating system, user interface, and application programs; the GPU is responsible for rendering and drawing display content; and the modem is used to process wireless communications. It can be understood that the above-mentioned modem may not be integrated into the processor 820, but may be implemented separately through a communication chip.

[0158] It is understandable that the electronic device may include more or fewer structural elements than those in the above structural block diagram, for example, a power module, physical buttons, a WiFi (Wireless Fidelity) module, a speaker, a Bluetooth module, a sensor, etc., and no limitation is made here.

[0159] An embodiment of the present application discloses a computer-readable storage medium storing a computer program, wherein the computer program enables a computer to execute the methods described in the above embodiments.

[0160] In addition, an embodiment of the present application further discloses a computer program product. When the computer program product is run on a computer, the computer can execute all or part of the steps in any one of the speech recognition methods described in the above embodiments.

[0161] Those skilled in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing related hardware through a program, and the program can be stored in a computer-readable storage medium, and the storage medium includes a read-only memory (ROM), a random access memory (RAM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), a one-time programmable read-only memory (OTPROM), an electronically erasable rewritable read-only memory (EEPROM), a compact disc (CD-ROM) or other optical disc storage, magnetic disk storage, magnetic tape storage, or any other computer-readable medium that can be used to carry or store data.

[0162] The above is a detailed introduction to a speech recognition method, device, electronic device and storage medium disclosed in the embodiments of the present application. Specific examples are used in this article to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea; at the same time, for general technical personnel in this field, according to the ideas of the present application, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on the present application.

Claims

1. A speech recognition method, characterized in that: include: Get the speech to be recognized; The speech to be recognized is converted into a target text through a speech recognition model; wherein the speech recognition model includes a target encoding module and a target decoding module, the target encoding module is used to encode the speech to be recognized to obtain a speech representation corresponding to the speech to be recognized, and the target decoding module includes a trained target large language model LLM, and the target LLM is used to decode the speech representation corresponding to the speech to be recognized to obtain the target text.

2. The method according to claim 1, characterized in that The target LLM is obtained by fine-tuning the initial LLM obtained by pre-training based on the sample data set; The sample data set includes speech representations corresponding to a plurality of sample speech and corresponding annotated texts.

3. The method according to claim 2, characterized in that The sample data set also includes task prompt information; the method also includes: Generate a model fine-tuning instruction corresponding to the target sample speech according to the task prompt information, the speech representation corresponding to the target sample speech, and the corresponding annotated text; the target sample speech is any one of the multiple sample speech; Decoding the speech representation corresponding to the target sample speech according to the task prompt information through the initial LLM to obtain the recognition text corresponding to the target sample speech; Determine a target loss between a recognized text corresponding to the target sample speech and a corresponding annotated text; According to the target losses respectively corresponding to the plurality of sample speech, the parameters of the initial LLM are adjusted until the model convergence condition is met, thereby obtaining the parameters corresponding to the target LLM.

4. The method according to claim 3, characterized in that The parameters corresponding to the initial LLM include original matrix parameters; the parameters of the initial LLM are adjusted according to the target losses respectively corresponding to the multiple sample speech until the model convergence condition is met to obtain the parameters corresponding to the target LLM, including: According to the target losses respectively corresponding to the plurality of sample speech, the target matrix parameters are adjusted; the rank of the target matrix parameters is less than the rank of the original matrix parameters; When the model convergence condition is met, the parameters corresponding to the target LLM are obtained according to the adjusted target matrix parameters and the original matrix parameters.

5. The method according to claim 3, characterized in that: The method further comprises: The target encoding module is used to encode the plurality of sample speech to obtain speech representations corresponding to the plurality of sample speech respectively.

6. The method according to claim 5, characterized in that The method further comprises: According to the target losses respectively corresponding to the multiple sample speech, the parameters of the target encoding module are adjusted to obtain an updated target encoding module.

7. The method according to claim 3, characterized in that The speech recognition model further includes a connector, which is used to map the speech representations corresponding to the multiple sample speech samples to the dimensional space corresponding to the LLM; The method further comprises: Parameters of the connector are adjusted according to target losses respectively corresponding to the plurality of sample speech sounds.

8. The method according to claim 1, characterized in that The converting the to-be-recognized speech into a target text by using a speech recognition model comprises: Encoding the speech to be recognized by the target encoding module to obtain a speech representation corresponding to the speech to be recognized; Inputting task prompt information, context information, and a speech representation corresponding to the speech to be recognized into the target LLM; The target LLM decodes the speech representation corresponding to the speech to be recognized according to the task prompt information and the context information to obtain the target text.

9. The method according to claim 8, characterized in that The speech recognition model further includes a connector. After encoding the speech to be recognized by the target encoding module to obtain a speech representation corresponding to the speech to be recognized, the method further includes: Mapping the speech representation corresponding to the speech to be recognized to the dimensional space corresponding to the LLM through the connector; The input task prompt information, context information and speech representation corresponding to the speech to be recognized into the target LLM include: The task prompt information, context information and the mapped speech representation are input into the target LLM.

10. A speech recognition device, characterized in that: The device comprises: A speech acquisition module, used to acquire the speech to be recognized; A speech recognition module, used for converting the speech to be recognized into a target text through a speech recognition model; wherein the speech recognition model includes a target encoding module and a target decoding module, the target encoding module is used for encoding the speech to be recognized to obtain a speech representation corresponding to the speech to be recognized, the target decoding module includes a trained target large language model LLM, the target LLM is used for decoding the speech representation corresponding to the speech to be recognized to obtain the target text.

11. An electronic device, characterized in that: include: A memory storing executable program code; a processor coupled to the memory; The processor calls the executable program code stored in the memory to execute the method according to any one of claims 1 to 9.

12. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, wherein when the computer program is executed by a processor, the processor is caused to perform the method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Speech recognition method and device, equipment and storage medium

    CN117153152A

  • Voice interaction method, server and computer readable storage medium

    CN117558277A

  • Speech recognition model training method and device, speech recognition method and device, equipment and medium

    CN117711386A

  • Large language model prompt information determination method, server and storage medium

    CN117789696A

  • Information acquisition method and device, computer equipment and storage medium

    CN118280363A