Method and device for predicting voice content based on multi-mode voice recognition model

By using a two-stage training method and a Lora adapter to fine-tune the large language model in the multimodal speech recognition model, the problem that the large language model in the existing technology cannot adapt to the actual scenario requirements is solved, and the effect of speech recognition is significantly improved.

CN120108387APending Publication Date: 2025-06-06ZUOYEBANG EDUCATION TECH (BEIJING) CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510222311.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-27
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

The large language model in the existing multimodal speech recognition system remains fixed and cannot adapt well to the actual scene requirements, resulting in poor speech recognition effect.

Method used

A method based on multimodal speech recognition model is proposed to improve the adaptability of the model through two training stages. The first stage is to train the speech-to-text big model and the modal adaptation layer, and the second stage is to adopt an optimized fine-tuning strategy for the large language model, including adding a Lora adapter to fine-tune each layer of the parameters of the large language model.

Benefits of technology

Through this method, the multimodal speech recognition model achieved the optimal word error rate on the test set in the field of education, with an average word error rate of nearly 20% better than the industry, significantly improving the effect of speech recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120108387A_ABST
    Figure CN120108387A_ABST
Patent Text Reader

Abstract

The invention discloses a method and a device for predicting voice content based on a multi-modal voice recognition model, and the multi-modal voice recognition model comprises a voice-to-text large model, a modal adaptation layer, an Embedding layer and a large language model. The training process of the multi-modal speech recognition model comprises the following steps: a first training stage: training the multi-modal speech recognition model by using automatic speech recognition data, keeping model parameters of a large language model and an Embedding layer fixed, and training a speech-to-text large model and a modal adaptation layer; and a second training stage: after the first training stage is completed, keeping model parameters of the voice-to-text large model obtained by training in the first training stage fixed, adopting an optimization fine tuning strategy for the large language model, and training a multi-modal voice recognition model by using automatic voice recognition data which is the same as that in the first stage. According to the method for predicting the voice content based on the multi-modal voice recognition model, the technical problem that parameters of a fixed language model cannot well adapt to actual scene requirements is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of large language models, and specifically provides a method and device for predicting speech content based on a multimodal speech recognition model. Background Art

[0002] Speech recognition means that the user speaks a piece of speech and the model recognizes the content of the speech. With the development of NLP large model technology, large language models based on large-scale text pre-training such as BERT, GPT, and Llama have achieved very good performance in various fields of natural language processing. In the past year, the combination of various NLP large models and speech large models for speech recognition has achieved good results.

[0003] The currently popular multimodal speech recognition systems usually consist of three parts: acoustic encoder, modal adaptation layer and large language model. The acoustic encoder is trained with large-scale speech data. The purpose of the modal adaptation layer is to transform the output of the acoustic encoder, map it to the same space as the text embedding of the language model, and then concatenate the mapping features and the text embedding features and send them to the language model to predict the speech content. In the past, the widely used multimodal speech recognition systems include seamless proposed by metaAI and qwen-audio proposed by Alibaba. Among them, qwen-audio proposed by Alibaba has the best speech recognition performance and the highest reusability.

[0004] However, in existing multimodal speech recognition systems, the large language model remains fixed, and the fixed language model cannot adapt well to actual scenario requirements.

[0005] In view of this, the present invention patent is proposed. Summary of the invention

[0006] In view of the above technical problems, the present invention proposes a method and device for predicting speech content based on a multimodal speech recognition model, aiming to improve the combination method of the multimodal speech recognition model to improve the effect of speech recognition.

[0007] Specifically, the following technical solutions are adopted:

[0008] In a first aspect, the present invention provides a method for predicting speech content based on a multimodal speech recognition model, wherein the multimodal speech recognition model includes a speech-to-text large model, a modal adaptation layer, an Embedding layer, and a large language model, and the training process of the multimodal speech recognition model includes:

[0009] First training phase: Use automatic speech recognition data to train the multimodal speech recognition model. The model parameters of the large language model and the Embedding layer remain fixed. The speech-to-text large model and the modality adaptation layer are trained.

[0010] Second training phase: After the first training phase is completed, the model parameters of the large speech-to-text model trained in the first training phase are kept fixed, an optimization fine-tuning strategy is adopted for the large language model, and the multimodal speech recognition model is trained using the same automatic speech recognition data as in the first phase.

[0011] As an optional embodiment of the present invention, in a method for predicting speech content based on a multimodal speech recognition model of the present invention, the second training stage adopts an optimization fine-tuning strategy for the large language model, including:

[0012] Add a Lora adapter to each layer parameter of the large language model to fine-tune the large language model.

[0013] As an optional embodiment of the present invention, in a method for predicting speech content based on a multimodal speech recognition model of the present invention, adding a Lora adapter to each layer parameter of the large language model to fine-tune the large language model includes:

[0014] The Lora adapter includes a first matrix W A and the second matrix W B Two parts, the first matrix W A Used to reduce the parameter dimension of a large language model, the second matrix W B The parameters for the large language model are restored to the first matrix W A Reduce the previous parameter dimension.

[0015] As an optional embodiment of the present invention, in a method for predicting speech content based on a multimodal speech recognition model of the present invention, in the second training stage, the multimodal speech recognition model is trained using the same automatic speech recognition data as in the first stage, and the first matrix W of the Lora adapter is obtained by training A and the second matrix W B ;

[0016] The first matrix W of the Lora adapter A , the second matrix W B Adding the fixed parameters in the large language model, the training is completed to obtain a multimodal speech recognition model.

[0017] As an optional embodiment of the present invention, a method for predicting speech content based on a multimodal speech recognition model of the present invention includes:

[0018] Input the speech into a multimodal speech recognition model;

[0019] The speech-to-text large model of the multimodal speech recognition model converts speech into speech-to-text, and the modal adaptation layer of the multimodal speech recognition model transforms the speech-to-text output of the speech-to-text large model, maps it to the same space as the text Embedding of the Embedding layer, and concatenates the mapping features and the text Embedding features and sends them to the large language model to predict the speech content.

[0020] As an optional embodiment of the present invention, a method for predicting speech content based on a multimodal speech recognition model of the present invention includes: performing text post-processing on the predicted speech content and limiting the maximum length of the final output predicted speech content text.

[0021] As an optional embodiment of the present invention, in a method for predicting speech content based on a multimodal speech recognition model of the present invention, the text post-processing for the predicted speech content and limiting the maximum length of the final output predicted speech content text includes:

[0022] Perform text post-processing on the predicted speech content;

[0023] According to the audio feature length of the input speech, the maximum length of the predicted speech content text finally output by the multimodal speech recognition model is proportionally limited.

[0024] As an optional embodiment of the present invention, in a method for predicting speech content based on a multimodal speech recognition model of the present invention, the text post-processing for the predicted speech content and limiting the maximum length of the final output predicted speech content text includes:

[0025] Perform text post-processing on the predicted speech content;

[0026] According to the preset length threshold of the output predicted speech content text, the maximum length of the predicted speech content text finally output by the multimodal speech recognition model is limited.

[0027] In a second aspect, the present invention provides a device for predicting speech content based on a multimodal speech recognition model, the multimodal speech recognition model comprising a speech-to-text large model, a modal adaptation layer, an Embedding layer and a large language model, the device for predicting speech content based on the multimodal speech recognition model comprising a model training module, the model training module comprising:

[0028] First training unit: Use automatic speech recognition data to train a multimodal speech recognition model. The model parameters of the large language model and the Embedding layer remain fixed. The speech-to-text large model and the modality adaptation layer are trained.

[0029] Second training unit: After the first training phase is completed, the model parameters of the large speech-to-text model trained in the first training phase are kept fixed, an optimization fine-tuning strategy is adopted for the large language model, and the multimodal speech recognition model is trained using the same automatic speech recognition data as in the first phase.

[0030] In a third aspect, the present invention provides a computer-readable recording medium storing a computer-executable program, which, when executed, implements the method for predicting speech content based on a multimodal speech recognition model.

[0031] Compared with the prior art, the present invention has the following beneficial effects:

[0032] The present invention provides a method for predicting speech content based on a multimodal speech recognition model. The training process of the multimodal speech recognition model includes two stages. The purpose of the first stage of training is to align speech features with text features, and the purpose of the second stage is to further fine-tune without changing the original language model parameters. Therefore, in order to solve the technical problem that a fixed language model cannot adapt well to the needs of actual scenarios, the present invention proposes a method for predicting speech content based on a multimodal speech recognition model. In the second training stage of training the multimodal speech recognition model, an optimization fine-tuning strategy is adopted for the large language model. Based on experiments, the present invention proposes a method based on a multimodal speech recognition model, which is trained with 40,000 hours of Chinese and English data, and has achieved the best word error rate on eight test sets related to the education field. The average word error rate is nearly 20% better than the industry.

[0033] The present invention discloses a method for predicting speech content based on a multimodal speech recognition model. In order to alleviate the repetitive hallucination caused by a large language model, text post-processing is performed on the predicted speech content to limit the maximum length of the predicted speech content text that is finally output. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] Figure 1 A flowchart of a first training phase of a multimodal speech recognition model in a method for predicting speech content based on a multimodal speech recognition model according to a first embodiment of the present invention;

[0035] Figure 2 A flowchart of a second training phase of a multimodal speech recognition model in a method for predicting speech content based on a multimodal speech recognition model according to a first embodiment of the present invention;

[0036] Figure 3 A schematic structural diagram of an electronic device according to a second embodiment of the present invention;

[0037] Figure 4 A schematic diagram of a computer-readable recording medium according to a second embodiment of the present invention. DETAILED DESCRIPTION

[0038] To make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be described clearly and completely in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them.

[0039] Therefore, the following detailed description of the embodiments of the present invention is not intended to limit the scope of the invention claimed for protection, but merely represents some embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0040] It should be noted that, in the absence of conflict, the embodiments of the present invention and the features and technical solutions in the embodiments may be combined with each other.

[0041] It should be noted that similar reference numerals and letters denote similar items in the following drawings, and therefore, once an item is defined in one drawing, further definition and explanation thereof is not required in subsequent drawings.

[0042] In the description of the present invention, it should be noted that the orientation or positional relationship indicated by the terms "upper", "lower", etc. is based on the orientation or positional relationship shown in the drawings, or the orientation or positional relationship in which the invention product is usually placed when in use, or the orientation or positional relationship commonly understood by those skilled in the art. Such terms are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as limiting the present invention. In addition, the terms "first", "second", etc. are only used to distinguish the description, and cannot be understood as indicating or implying relative importance.

[0043] Embodiment 1

[0044] In this embodiment, a method for predicting speech content based on a multimodal speech recognition model is provided. The multimodal speech recognition model includes a speech-to-text large model, a modality adaptation layer, an Embedding layer, and a large language model. The training process of the multimodal speech recognition model includes:

[0045] The first training phase, see Figure 1 As shown in the figure: the multimodal speech recognition model is trained using automatic speech recognition data. The model parameters of the large language model and the Embedding layer remain fixed, and the speech-to-text large model and the modality adaptation layer are trained.

[0046] The second training phase, see Figure 2As shown: After the first training stage is completed, the model parameters of the large speech-to-text model trained in the first training stage are kept fixed, an optimization fine-tuning strategy is adopted for the large language model, and the multimodal speech recognition model is trained using the same automatic speech recognition data as in the first stage.

[0047] The present embodiment provides a method for predicting speech content based on a multimodal speech recognition model. The training process of the multimodal speech recognition model includes two stages. The purpose of the first stage of training is to align speech features with text features, and the purpose of the second stage is to further fine-tune without changing the original language model parameters. Therefore, in view of the technical problem that a fixed language model cannot adapt well to the actual scene requirements, the present embodiment proposes a method for predicting speech content based on a multimodal speech recognition model. In the second training stage of training the multimodal speech recognition model, an optimization fine-tuning strategy is adopted for the large language model. Based on experiments, the present embodiment proposes a method based on a multimodal speech recognition model, which is trained with 40,000 hours of Chinese and English data, and has achieved the best word error rate on eight test sets related to the education field. The average word error rate is nearly 20% better than the industry.

[0048] Specifically, the automatic speech recognition data used in the training model of this embodiment includes 1.4w English open source speech recognition data and 2.4w Chinese open source and education-related speech recognition data.

[0049] See also Figure 1 As shown, in the first training phase of the training process of the multimodal speech recognition model of this embodiment, the automatic speech recognition data is used to train the entire multimodal speech recognition model, and the expected result of Next Token Prediction is the recognition text of the speech. Next Token Prediction (NTP) is an important concept in natural language processing (NLP), which is mainly used in the pre-training stage of the language model. The basic idea of ​​NTP is to predict the next token when the first n tokens of a sentence are known. This process is similar to cloze test, but unlike cloze test, NTP has no constraints on the context and candidates, which makes the prediction results more diverse. Among them, the large language model adopts Qwen-7B, the core principle of Qwen-7B, the Transformer architecture, which is a deep learning model widely used in natural language processing tasks. Transfomer captures long-distance dependencies in the input sequence through the self-attention mechanism, so as to perform well in tasks such as language modeling and text generation. Qwen-7B is trained on ultra-large-scale pre-training data, which enables it to have powerful language understanding and generation capabilities.

[0050] In the first training stage, the large language model and embedding layer parameters are fixed and not frozen. The speech-to-text large model (acoustic encoder) and modal adapter are trained. The speech-to-text large model uses Whisper, and the proposed features are sent to the C-Adapter adapter for learning.

[0051] Whisper is an end-to-end voice system with the following features:

[0052] Multilingual: English is the main language, supporting 99 languages, including Chinese.

[0053] Multi-task: mainly speech recognition, supporting VAD, language recognition, speaker log, speech translation, alignment, etc.

[0054] Data volume: 680,000 hours of speech data are used for training, and speech data in multiple languages ​​is obtained from public data sets or the Internet, far exceeding the data volume of hundreds, thousands, or at most 10,000 hours of speech recognition in the past. This will be described in detail below.

[0055] Robustness: This is mainly due to the massive amount of training data, and common enhancement operations are performed on the speech data, such as speed change, noise addition, spectral enhancement, etc.

[0056] Multiple models: Five models are provided, ranging from tiny to large, suitable for different scenarios.

[0057] See also Figure 2 As shown, in the second training phase of the training process of the multimodal speech recognition model of this embodiment, the optimization fine-tuning strategy for the large language model includes: adding a Lora adapter to each layer parameter of the large language model to fine-tune the large language model.

[0058] Furthermore, in a method for predicting speech content based on a multimodal speech recognition model of this embodiment, adding a Lora adapter to each layer parameter of the large language model to fine-tune the large language model includes:

[0059] The Lora adapter includes a first matrix W A and the second matrix W B Two parts, the first matrix W A Used to reduce the parameter dimension of a large language model, the second matrix W B The parameters for the large language model are restored to the first matrix W A Reduce the previous parameter dimension.

[0060] Specifically, in a method for predicting speech content based on a multimodal speech recognition model in this embodiment, in the second training stage, the multimodal speech recognition model is trained using the same automatic speech recognition data as in the first stage, and the first matrix W of the Lora adapter is obtained by training. A and the second matrix W B ;

[0061] The first matrix W of the Lora adapter A , the second matrix W B Adding the fixed parameters in the large language model, the training is completed to obtain a multimodal speech recognition model.

[0062] The essence of the LoRA (Low.Rank Adaptation) method is to adjust the feature representation between the source domain and the target domain based on the low-rank adaptation matrix. The LoRA adapter inserts a parallel matrix next to some matrices of the large language model. The parallel matrix consists of the first matrix W A , the second matrix W B The two matrices are composed of:

[0063] The first matrix W A : Reduce the dimensionality of the matrix, reduce the rank of the original matrix from d to r, r< <d。

[0064] The second matrix W B : Raise the dimension of the matrix, raising the rank from r to d.

[0065] When initialized, the first matrix W A Initialize with Gaussian distribution, the second matrix W B Use a matrix of all zeros.

[0066] During training, fix the parameters of the original large language model on the left and only train the first matrix W A , the second matrix W B The rank of the two matrices, the LORA adapter, is usually a very small number (1, 2, 4, 8 are all OK), so its training is very efficient and will not bring a significant increase in video memory.

[0067] After training is completed, the first matrix W can be directly A and the second matrix W B It is directly added to the original parameters, so that no additional inference delay is generated during inference. In the training of large language models, the number of neurons in the two fully connected layers may be very large. For example, if the fully connected in layer and the 0ut layer each have 10,000 neurons, then there are 10,000*10,000 parameters that need to be iterated in the full parameter training.

[0068] The inserted LORA adapter is a low-rank matrix with a very small dimension, for example, the dimension is 8*8, so the current calculation parameters are equal to 10,000*8*2+8*8, which greatly reduces the number of training parameters.

[0069] The LORA adapter is like an "intermediary" that aggregates the needs of buyers and sellers (in layer and out layer), completes information exchange within the "intermediary", and speeds up matching.

[0070] A method for predicting speech content based on a multimodal speech recognition model in this embodiment includes:

[0071] Input the speech into a multimodal speech recognition model;

[0072] The speech-to-text large model of the multimodal speech recognition model converts speech into speech-to-text, and the modal adaptation layer of the multimodal speech recognition model transforms the speech-to-text output of the speech-to-text large model, maps it to the same space as the text Embedding of the Embedding layer, and concatenates the mapping features and the text Embedding features and sends them to the large language model to predict the speech content.

[0073] In this embodiment, a method for predicting speech content based on a multimodal speech recognition model can predict speech content based on input speech by using a trained multimodal speech recognition model. Therefore, the multimodal speech recognition model of this embodiment improves the combination of a large speech-to-text model and a large language model to improve the effect of speech recognition.

[0074] A method for predicting speech content based on a multimodal speech recognition model in this embodiment includes: performing text post-processing on the predicted speech content and limiting the maximum length of the predicted speech content text that is finally output.

[0075] When inferring and predicting speech content results, in order to alleviate the repetitive hallucinations caused by large language models, the maximum length of the final output predicted speech content text is limited.

[0076] As an optional implementation of this embodiment, in a method for predicting speech content based on a multimodal speech recognition model of this embodiment, the text post-processing for the predicted speech content and limiting the maximum length of the final output predicted speech content text includes:

[0077] Perform text post-processing on the predicted speech content;

[0078] According to the audio feature length of the input speech, the maximum length of the predicted speech content text finally output by the multimodal speech recognition model is proportionally limited.

[0079] As another optional implementation of this embodiment, in a method for predicting speech content based on a multimodal speech recognition model in this embodiment, the text post-processing for the predicted speech content and limiting the maximum length of the final output predicted speech content text includes:

[0080] Perform text post-processing on the predicted speech content;

[0081] According to the preset length threshold of the output predicted speech content text, the maximum length of the predicted speech content text finally output by the multimodal speech recognition model is limited.

[0082] This embodiment also provides a device for predicting speech content based on a multimodal speech recognition model, the multimodal speech recognition model includes a speech-to-text large model, a modal adaptation layer, an Embedding layer and a large language model, the device for predicting speech content based on the multimodal speech recognition model includes a model training module, and the model training module includes:

[0083] First training unit: Use automatic speech recognition data to train a multimodal speech recognition model. The model parameters of the large language model and the Embedding layer remain fixed. The speech-to-text large model and the modality adaptation layer are trained.

[0084] Second training unit: After the first training phase is completed, the model parameters of the large speech-to-text model trained in the first training phase are kept fixed, an optimization fine-tuning strategy is adopted for the large language model, and the multimodal speech recognition model is trained using the same automatic speech recognition data as in the first phase.

[0085] In this embodiment, a device for predicting speech content based on a multimodal speech recognition model is provided. The model training module includes a first training unit and a second training unit. The purpose of the training of the first training unit is to align speech features with text features, and the purpose of the second training unit is to further fine-tune without changing the original language model parameters. Therefore, in view of the technical problem that a fixed language model cannot be well adapted to the actual scene requirements, this embodiment proposes a device for predicting speech content based on a multimodal speech recognition model. The second training unit of the model training module adopts an optimized fine-tuning strategy for the large language model. Based on experiments, the multimodal speech recognition model proposed in this embodiment has achieved the best word error rate on eight test sets related to the education field when trained with 40,000 hours of Chinese and English data, and the average word error rate is nearly 20% better than the industry.

[0086] See also Figure 1As shown, the first training unit of this embodiment uses automatic speech recognition data to train the entire multimodal speech recognition model, and expects that the result of Next Token Prediction is the recognized text of the speech. Next Token Prediction (NTP) is an important concept in natural language processing (NLP), which is mainly used in the pre-training stage of the language model. The basic idea of ​​NTP is to predict the next token when the first n tokens of a sentence are known. This process is similar to cloze test, but unlike cloze test, NTP has no constraints on the context and candidates, which makes the prediction results more diverse. Among them, the large language model adopts Qwen-7B, the core principle of Qwen-7B, the Transformer architecture, which is a deep learning model widely used in natural language processing tasks. Transfomer captures long-distance dependencies in the input sequence through the self-attention mechanism, so as to perform well in tasks such as language modeling and text generation. Qwen-7B is trained on ultra-large-scale pre-training data, which enables it to have powerful language understanding and generation capabilities.

[0087] In the training process of the first training unit, the large language model and embedding layer parameters are fixed and not frozen. The speech-to-text large model (acoustic encoder) and modal adapter are trained. The speech-to-text large model uses Whisper, and the proposed features are sent to the C-Adapter adapter for learning.

[0088] See also Figure 2 As shown, the second training unit of this embodiment adopts an optimization fine-tuning strategy for the large language model, including: adding a Lora adapter to each layer parameter of the large language model to fine-tune the large language model.

[0089] Furthermore, in this embodiment, a Lora adapter is added to each layer parameter of the large language model, and fine-tuning of the large language model includes:

[0090] The Lora adapter includes a first matrix W A and the second matrix W B Two parts, the first matrix W A Used to reduce the parameter dimension of a large language model, the second matrix W B The parameters for the large language model are restored to the first matrix W A Reduce the previous parameter dimension.

[0091] Specifically, in a method for predicting speech content based on a multimodal speech recognition model in this embodiment, in the second training stage, the multimodal speech recognition model is trained using the same automatic speech recognition data as in the first stage, and the first matrix W of the Lora adapter is obtained by training. A and the second matrix W B ;

[0092] The first matrix W of the Lora adapter A , the second matrix W B Adding the fixed parameters in the large language model, the training is completed to obtain a multimodal speech recognition model.

[0093] The essence of the LoRA (Low.Rank Adaptation) method is to adjust the feature representation between the source domain and the target domain based on the low-rank adaptation matrix. The LoRA adapter inserts a parallel matrix next to some matrices of the large language model. The parallel matrix consists of the first matrix W A , the second matrix W B The two matrices are composed of:

[0094] The first matrix W A : Reduce the dimensionality of the matrix, reduce the rank of the original matrix from d to r, r< <d。

[0095] The second matrix W B : Raise the dimension of the matrix, raising the rank from r to d.

[0096] When initialized, the first matrix W A Initialize with Gaussian distribution, the second matrix W B Use a matrix of all zeros.

[0097] During training, fix the parameters of the original large language model on the left and only train the first matrix W A , the second matrix W B The rank of the two matrices, the LORA adapter, is usually a very small number (1, 2, 4, 8 are all OK), so its training is very efficient and will not bring a significant increase in video memory.

[0098] After training is completed, the first matrix W can be directly A and the second matrix W B It is directly added to the original parameters, so that no additional inference delay is generated during inference. In the training of large language models, the number of neurons in the two fully connected layers may be very large. For example, if the fully connected in layer and the 0ut layer each have 10,000 neurons, then there are 10,000*10,000 parameters that need to be iterated in the full parameter training.

[0099] The inserted LORA adapter is a low-rank matrix with a very small dimension, for example, the dimension is 8*8, so the current calculation parameters are equal to 10,000*8*2+8*8, which greatly reduces the number of training parameters.

[0100] A device for predicting speech content based on a multimodal speech recognition model in this embodiment includes:

[0101] Input the speech into a multimodal speech recognition model;

[0102] The speech-to-text large model of the multimodal speech recognition model converts speech into speech-to-text, and the modal adaptation layer of the multimodal speech recognition model transforms the speech-to-text output of the speech-to-text large model, maps it to the same space as the text Embedding of the Embedding layer, and concatenates the mapping features and the text Embedding features and sends them to the large language model to predict the speech content.

[0103] The present embodiment is a device for predicting speech content based on a multimodal speech recognition model. By using a trained multimodal speech recognition model, the speech content can be predicted according to the input speech. Therefore, the multimodal speech recognition model of the present embodiment improves the combination of the speech-to-text large model and the large language model to improve the effect of speech recognition.

[0104] A device for predicting speech content based on a multimodal speech recognition model in this embodiment includes a text post-processing module: performing text post-processing on the predicted speech content to limit the maximum length of the predicted speech content text that is finally output.

[0105] When inferring and predicting speech content results, in order to alleviate the repetitive hallucinations caused by large language models, the text post-processing module limits the maximum length of the final output predicted speech content text.

[0106] As an optional implementation of this embodiment, in a device for predicting speech content based on a multimodal speech recognition model of this embodiment, the text post-processing module performs text post-processing on the predicted speech content, and limits the maximum length of the final output predicted speech content text, including:

[0107] Perform text post-processing on the predicted speech content;

[0108] According to the audio feature length of the input speech, the maximum length of the predicted speech content text finally output by the multimodal speech recognition model is proportionally limited.

[0109] As another optional implementation of this embodiment, in a device for predicting speech content based on a multimodal speech recognition model of this embodiment, the text post-processing module performs text post-processing on the predicted speech content, and limits the maximum length of the final output predicted speech content text, including:

[0110] Perform text post-processing on the predicted speech content;

[0111] According to the preset length threshold of the output predicted speech content text, the maximum length of the predicted speech content text finally output by the multimodal speech recognition model is limited.

[0112] Embodiment 2

[0113] The following describes an electronic device embodiment of the present invention, which can be regarded as a specific physical implementation of the method and device embodiments of the present invention. The details described in the electronic device embodiment of the present invention should be regarded as a supplement to the above method or device embodiments; details not disclosed in the electronic device embodiment of the present invention can be implemented with reference to the above method or device embodiments.

[0114] Figure 3 It is a structural schematic diagram of an electronic device of an embodiment of the present invention, the electronic device includes a processor and a memory, the memory is used to store a computer executable program, when the computer program is executed by the processor, the processor executes a method for predicting speech content based on a multimodal speech recognition model of embodiment one.

[0115] like Figure 3 As shown, the electronic device is presented in the form of a general computing device. The processor may be one or more and work in coordination. The present invention does not exclude distributed processing, that is, the processor may be dispersed in different physical devices. The electronic device of the present invention is not limited to a single entity, but may also be the sum of multiple physical devices.

[0116] The memory stores a computer executable program, which is usually a machine-readable code. The computer-readable program can be executed by the processor to enable the electronic device to perform the method of the present invention, or at least part of the steps in the method.

[0117] The memory includes a volatile memory, such as a random access memory unit (RAM) and / or a cache memory unit, and may also be a non-volatile memory, such as a read-only memory unit (ROM).

[0118] Optionally, in this embodiment, the electronic device further includes an I / O interface, which is used for the electronic device to exchange data with an external device. The I / O interface can represent one or more of several types of bus structures, including a storage unit bus or a storage unit controller, a peripheral bus, a graphics acceleration port, a processing unit, or a local bus using any of a variety of bus structures.

[0119] It should be understood that Figure 3 The electronic device shown is only an example of the present invention, and the electronic device of the present invention may also include elements or components not shown in the above examples. For example, some electronic devices also include display units such as display screens, and some electronic devices also include human-computer interaction elements such as buttons, keyboards, etc. As long as the electronic device can execute the computer-readable program in the memory to implement the method of the present invention or at least part of the steps of the method, it can be considered as an electronic device covered by the present invention.

[0120] Figure 4 Schematic diagram of a computer readable recording medium according to an embodiment of the present invention. Figure 4 As shown, a computer executable program is stored in a computer-readable recording medium, and when the computer executable program is executed, a method for predicting speech content based on a multimodal speech recognition model in the first embodiment of the present invention is implemented. The computer-readable recording medium may include a data signal propagated in a baseband or as part of a carrier, which carries a readable program code. This propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The readable recording medium may also be any readable medium other than a readable recording medium, which may send, propagate, or transmit a program for use by or in combination with an instruction execution system, device, or device. The program code contained on the readable recording medium may be transmitted using any appropriate medium, including but not limited to wireless, wired, optical cable, RF, etc., or any suitable combination of the above.

[0121] Program code for performing the operations of the present invention may be written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Java, C++, etc., and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may be executed entirely on the user computing device, partially on the user device, as a separate software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server. In cases involving a remote computing device, the remote computing device may be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device (e.g., through the Internet using an Internet service provider).

[0122] Through the above description of the implementation mode, it is easy for those skilled in the art to understand that the present invention can be implemented by hardware capable of executing a specific computer program, such as the system of the present invention, and the electronic processing unit, server, client, mobile phone, control unit, processor, etc. contained in the system. The present invention can also be implemented by computer software that executes the method of the present invention, such as control software executed by a microprocessor, an electronic control unit, a client, a server, etc. However, it should be noted that the computer software that executes the method of the present invention is not limited to being executed by one or a specific hardware entity, and it can also be implemented in a distributed manner by unspecified specific hardware. For computer software, the software product can be stored in a computer-readable recording medium (which can be a CD-ROM, a USB flash drive, a mobile disk, etc.), and can also be distributed and stored on the network, as long as it enables the electronic device to execute the method according to the present invention.

[0123] The above embodiments are only used to illustrate the present invention and are not intended to limit the technical solutions described in the present invention. Although the present invention has been described in detail with reference to the above embodiments, the present invention is not limited to the above specific implementation methods. Therefore, any modification or equivalent replacement of the present invention; and all technical solutions and improvements thereof that do not depart from the spirit and scope of the invention are included in the scope of the claims of the present invention.

Claims

1. A method for predicting speech content based on a multimodal speech recognition model, characterized in that: The multimodal speech recognition model includes a speech-to-text model, a modality adaptation layer, an embedding layer, and a large language model. The training process of the multimodal speech recognition model includes: First training phase: Use automatic speech recognition data to train the multimodal speech recognition model. The model parameters of the large language model and the Embedding layer remain fixed. The speech-to-text large model and the modality adaptation layer are trained. Second training phase: After the first training phase is completed, the model parameters of the large speech-to-text model trained in the first training phase are kept fixed, an optimization fine-tuning strategy is adopted for the large language model, and the multimodal speech recognition model is trained using the same automatic speech recognition data as in the first phase.

2. The method for predicting speech content based on a multimodal speech recognition model according to claim 1, characterized in that: The second training stage adopts an optimization fine-tuning strategy for the large language model including: Add a Lora adapter to each layer parameter of the large language model to fine-tune the large language model.

3. The method for predicting speech content based on a multimodal speech recognition model according to claim 2, characterized in that: The adding of a Lora adapter to each layer parameter of the large language model and fine-tuning of the large language model includes: The Lora adapter includes a first matrix W A and the second matrix W B Two parts, the first matrix W A Used to reduce the parameter dimension of a large language model, the second matrix W B The parameters for the large language model are restored to the first matrix W A Reduce the previous parameter dimension.

4. The method for predicting speech content based on a multimodal speech recognition model according to claim 3, characterized in that: In the second training stage, the multimodal speech recognition model is trained using the same automatic speech recognition data as in the first stage to obtain the first matrix W of the Lora adapter. A and the second matrix W B ; The first matrix W of the Lora adapter A , the second matrix W B Adding the fixed parameters in the large language model, the training is completed to obtain a multimodal speech recognition model.

5. The method for predicting speech content based on a multimodal speech recognition model according to claim 4, characterized in that: include: Input the speech into a multimodal speech recognition model; The speech-to-text large model of the multimodal speech recognition model converts speech into speech-to-text, and the modal adaptation layer of the multimodal speech recognition model transforms the speech-to-text output of the speech-to-text large model, maps it to the same space as the text embedding of the embedding layer, and concatenates the mapping features and the text embedding features and sends them to the large language model to predict the speech content.

6. The method for predicting speech content based on a multimodal speech recognition model according to claim 5, characterized in that: include: Perform text post-processing on the predicted speech content to limit the maximum length of the final output predicted speech content text.

7. The method for predicting speech content based on a multimodal speech recognition model according to claim 6, characterized in that: The text post-processing for the predicted speech content and limiting the maximum length of the final output predicted speech content text includes: Perform text post-processing on the predicted speech content; According to the audio feature length of the input speech, the maximum length of the predicted speech content text finally output by the multimodal speech recognition model is proportionally limited.

8. The method for predicting speech content based on a multimodal speech recognition model according to claim 6, characterized in that: The text post-processing for the predicted speech content to limit the maximum length of the final output predicted speech content text includes: Perform text post-processing on the predicted speech content; According to the preset length threshold of the output predicted speech content text, the maximum length of the predicted speech content text finally output by the multimodal speech recognition model is limited.

9. A device for predicting speech content based on a multimodal speech recognition model, characterized in that: The multimodal speech recognition model includes a speech-to-text large model, a modal adaptation layer, an Embedding layer, and a large language model. The device for predicting speech content based on the multimodal speech recognition model includes a model training module, and the model training module includes: First training unit: Use automatic speech recognition data to train a multimodal speech recognition model. The model parameters of the large language model and the Embedding layer remain fixed. The speech-to-text large model and the modality adaptation layer are trained. Second training unit: After the first training phase is completed, the model parameters of the large speech-to-text model trained in the first training phase are kept fixed, an optimization fine-tuning strategy is adopted for the large language model, and the multimodal speech recognition model is trained using the same automatic speech recognition data as in the first phase.

10. A computer-readable recording medium storing a computer-executable program, characterized in that: When the computer executable program is executed, a method for predicting speech content based on a multimodal speech recognition model as described in any one of claims 1 to 8 is implemented.

Citation Information

Cited By

  • Cross-modal context speech recognition method and system, and storage medium

    CN121662047A