Multi-modal speech recognition method based on large model, storage medium, electronic equipment and product
Through the multimodal speech recognition method of the large model, combined with preprocessing, feature extraction and encoder-decoder structure, the problem of insufficient speech recognition accuracy of the HMM-GMM model in variable scenarios is solved, and the speech recognition effect with high accuracy and wide adaptability is achieved.
Patent Information
- Application Number
- CN202510749396.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-06
- Publication Date
- 2025-08-05
AI Technical Summary
When the existing HMM-GMM model faces variable scenarios such as different languages, speakers and background environments, the speech recognition accuracy is low, and the robustness and adaptability are insufficient.
The multimodal speech recognition method based on large models is adopted, and the original speech signal is preprocessed, feature extraction and spliced text vectors are used to convert speech signals to text using pre-trained encoder and decoder layers. Combined with large language models and feature extraction technology, high-quality text data is output.
It improves the accuracy of speech recognition, has a wide range of adaptability, and can maintain high-precision speech recognition effect in different languages, speakers and background environments, reducing the need for complex computing resources and training data.
Smart Images

Figure CN120431933A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of speech recognition technology, and in particular to a large-model-based multimodal speech recognition method, storage medium, electronic device, and product. Background Art
[0002] With the increasing popularity and rapid development of intelligent voice customer service, it is gradually being applied across various industries. This has led to issues with speech recognition accuracy. Currently, speech recognition technology in practical applications primarily relies on several common models, such as a combination of Hidden Markov Models (HMMs) and Gaussian Mixture Models (GMMs). The basic principle is to preprocess and extract features from the input speech signal to obtain a series of feature vectors. These feature vectors are then modeled using the HMM, and each HMM state is parameterized using the GMM to describe the probability distribution of the speech signal. However, due to the fixed model parameters of the HMM-GMM model, it exhibits poor robustness and adaptability in diverse scenarios, such as different languages, speakers, and background environments, resulting in reduced speech recognition accuracy.
[0003] Therefore, how to provide a technical solution for a method of speech recognition with higher accuracy has become a technical problem that needs to be solved urgently. Summary of the Invention
[0004] The purpose of some embodiments of the present application is to provide a multimodal speech recognition method, storage medium, electronic device and product based on a large model. The technical solutions of the embodiments of the present application can improve the accuracy of speech recognition, and are not affected by various scenarios such as language, speaker, background environment, etc. The speech recognition effect is good and the applicability is wide.
[0005] In a first aspect, some embodiments of the present application provide a method for multimodal speech recognition based on a large model, comprising: preprocessing the user's original speech signal to obtain a processed speech signal, wherein the preprocessing includes: filtering and gain adjustment; inputting the speech coding data and historical conversation data corresponding to the processed speech signal into a large language model to obtain a text vector corresponding to the processed speech signal; performing feature extraction on the processed speech signal to obtain a speech feature vector; using a pre-trained speech recognition module to process a target vector sequence obtained by concatenating the speech feature vector and the text vector to obtain a text sequence; wherein the speech recognition module includes multiple pre-trained encoder layers and multiple decoder layers; cleaning and formatting the text sequence to obtain text data corresponding to the original speech signal.
[0006] Some embodiments of the present application process the original speech signal using a large language model and feature extraction, concatenating the resulting text vector and speech feature vector into a speech recognition module, outputting a text sequence, and finally formatting the text sequence to obtain text data. Some embodiments of the present application can achieve precise conversion of speech signals into text, improving the accuracy of speech recognition, and are not affected by various scenarios such as language, speaker, and background environment. The speech recognition effect is good and has wide applicability.
[0007] In some embodiments, the preprocessing of the user's original voice signal to obtain a processed voice signal includes: dividing the original voice signal into multiple frames; filtering each frame of the multiple frames to obtain a filtered frame; and adjusting the gain of each frame in the filtered frame based on a preset target energy range to obtain the processed voice signal.
[0008] Some embodiments of the present application obtain a processed speech signal by filtering and adjusting the gain after framing the original speech signal, so as to obtain a high-quality processed speech signal to provide support for subsequent accurate speech recognition.
[0009] In some embodiments, the feature extraction of the processed speech signal to obtain a speech feature vector includes: dividing the processed speech signal according to a preset length to obtain multiple signal segments; adding a window function and performing Fourier transform on each signal segment in the multiple signal segments to obtain speech frequency domain features; and performing feature processing on the speech frequency domain features to obtain the speech feature vector.
[0010] Some embodiments of the present application obtain speech frequency domain features by segmenting and converting the processed speech signal, and then perform feature processing on the signal to obtain speech feature vectors, thereby achieving accurate extraction and representation of speech features.
[0011] In some embodiments, the pre-trained speech recognition module is used to process the target vector sequence obtained by concatenating the speech feature vector and the text vector to obtain a text sequence, including: using the multiple encoder layers to perform feature processing on the target vector sequence to obtain a hidden layer speech state sequence; and using the multiple decoder layers to perform feature processing on the hidden layer speech state sequence to generate the text sequence.
[0012] Some embodiments of the present application process a target vector sequence through multiple encoder layers and multiple decoder layers to obtain a text sequence, which can achieve accurate and efficient conversion from speech to text.
[0013] In some embodiments, the use of the multiple encoder layers to perform feature processing on the target vector sequence to obtain a hidden layer speech state sequence includes: each encoder layer in the multiple encoder layers performs the following operations until the last encoder layer outputs the hidden layer speech state sequence: performing high-dimensional space embedding processing on the target vector sequence and then adding position coding to obtain a time series vector; performing a linear transformation on the time series vector to obtain a first linear transformation matrix; and generating an intermediate state sequence based on the first linear transformation matrix and a feedforward neural network.
[0014] Some embodiments of the present application perform a series of operations on the target vector sequence through multiple encoder layers until a hidden layer speech state sequence output by the last encoder layer is obtained, thereby achieving efficient processing of the speech signal.
[0015] In some embodiments, the use of the multiple decoder layers to perform feature processing on the hidden layer speech state sequence to generate the text sequence includes: each decoder layer in the multiple decoder layers performs the following operations until the last decoder layer outputs the text sequence: embedding the hidden layer speech state sequence into a high-dimensional space and adding position encoding to obtain a coding sequence; performing a linear transformation on the coding sequence to obtain a second linear transformation matrix; and generating a text intermediate sequence based on the second linear transformation matrix and a feedforward fully connected network.
[0016] Some embodiments of the present application perform a series of operations on the hidden layer speech state sequence through multiple decoder layers until a text sequence output by the last decoder layer is obtained, thereby achieving efficient conversion of the speech signal.
[0017] In some embodiments, the cleaning and formatting of the text sequence to obtain text data corresponding to the original speech signal includes: cleaning spaces and spelling errors in the text sequence to obtain a cleaned text sequence; and formatting the cleaned text sequence using a regular expression to obtain the text data.
[0018] Some embodiments of the present application can obtain standard and accurate text data by cleaning and formatting text sequences, thereby achieving efficient and accurate recognition of voice signals.
[0019] In a second aspect, some embodiments of the present application provide a device for multimodal speech recognition based on a large model, comprising: a preprocessing module for preprocessing the user's original speech signal to obtain a processed speech signal, wherein the preprocessing includes: filtering and gain adjustment; a large model processing module for inputting the speech coding data and historical conversation data corresponding to the processed speech signal into a large language model to obtain a text vector corresponding to the processed speech signal; an extraction module for performing feature extraction on the processed speech signal to obtain a speech feature vector; a recognition module for using a pre-trained speech recognition module to process a target vector sequence obtained by splicing the speech feature vector and the text vector to obtain a text sequence; wherein the speech recognition module includes multiple pre-trained encoder layers and multiple decoder layers; a text processing module for cleaning and formatting the text sequence to obtain text data corresponding to the original speech signal.
[0020] In a third aspect, some embodiments of the present application provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, can implement the method described in any embodiment of the first aspect.
[0021] In a fourth aspect, some embodiments of the present application provide an electronic device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the processor can implement a method as described in any embodiment of the first aspect when executing the program.
[0022] In a fifth aspect, some embodiments of the present application provide a computer program product, comprising a computer program, wherein the computer program, when executed by a processor, can implement the method described in any embodiment of the first aspect. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] In order to more clearly illustrate the technical solutions of some embodiments of the present application, the following is a brief introduction to the drawings required for use in some embodiments of the present application. It should be understood that the following drawings only show certain embodiments of the present application and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other relevant drawings can be obtained based on these drawings without paying any creative work.
[0024] Figure 1 A system diagram of large-model-based multimodal speech recognition provided for some embodiments of the present application; Figure 2 One of the flow charts of the method for multimodal speech recognition based on a large model provided in some embodiments of the present application; Figure 3A second flowchart of a method for multimodal speech recognition based on a large model provided in some embodiments of the present application; Figure 4 A block diagram of a device for large-model-based multimodal speech recognition provided in some embodiments of the present application; Figure 5 A schematic diagram of an electronic device is provided for some embodiments of the present application. DETAILED DESCRIPTION
[0025] The technical solutions in some embodiments of the present application will be described below in conjunction with the drawings in some embodiments of the present application.
[0026] It should be noted that similar reference numerals and letters represent similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined or explained in subsequent drawings. At the same time, in the description of this application, the terms "first", "second", etc. are only used to distinguish the description and should not be understood as indicating or implying relative importance.
[0027] In related technologies, a trained HMM-GMM model is usually used for speech recognition; however, it still has the following problems: the speech recognition accuracy in continuous speech and natural environments is still insufficient, and the presence of different languages, accent variability, and background noise all pose great challenges to the HMM-GMM model; the preprocessing and feature extraction of speech signals require complex and large-scale computing resources, and the model's excessive reliance on features makes it difficult to adapt to changing practical application scenarios; and a large amount of labeled speech data is required during model training, and many unlabeled speech data are difficult to use, resulting in huge data requirements and high training costs.
[0028] In view of this, some embodiments of the present application provide a method for multimodal speech recognition based on a large model, in which a text vector and a speech feature vector are obtained by preprocessing the original speech signal and combining it with a large language model and feature extraction technology; the two are spliced and input into the speech recognition module, and a text sequence is output. Finally, the text sequence is integrated and processed to obtain text data. The speech recognition module mainly includes two core components: an encoder and a decoder. The encoder is responsible for extracting high-level features of speech signals and text information, and the decoder converts speech features into text sequences. Compared with traditional methods, the encoder-decoder architecture has significant advantages in processing complex sequential data and heterogeneous data. At the same time, combined with multimodal information, it further improves the speech-to-text effect and the accuracy of speech recognition.
[0029] The following is combined with Figure 1The overall composition structure of the system for multimodal speech recognition based on a large model provided by some embodiments of the present application is exemplified.
[0030] like Figure 1 As shown, some embodiments of the present application provide a system diagram of multimodal speech recognition based on a large model. The system of multimodal speech recognition based on a large model may include: a mobile terminal 100 and an intelligent voice server 200. The user can communicate with the intelligent voice server 200 through the mobile terminal 100, and the mobile terminal 100 can capture the user's original voice signal through a microphone or other recording device. After receiving the user's original voice signal, the intelligent voice server 200 can pre-process it to obtain a processed voice signal; then, the large language model deployed inside the intelligent voice server 200 can analyze the processed voice signal and output the corresponding text vector; at the same time, feature processing is performed on the processed voice signal to obtain a voice feature vector; then, the text vector and the voice feature vector are spliced and input into the deployed voice recognition module to obtain a text sequence; finally, the text sequence is integrated to obtain text data corresponding to the original voice signal. This system can achieve accurate and efficient conversion of voice signals to text data, and has wide adaptability.
[0031] In some embodiments of the present application, the mobile terminal 100 may be a smart phone, a landline phone, a tablet computer with communication functions, etc., which is not specifically limited in the embodiments of the present application.
[0032] The following is combined with Figure 2 The implementation process of large-model-based multimodal speech recognition performed by the intelligent speech server 200 provided in some embodiments of the present application is exemplified.
[0033] Please see the attached Figure 2 , Figure 2 A flowchart of a method for multimodal speech recognition based on a large model is provided for some embodiments of the present application. The method for multimodal speech recognition based on a large model may include: S210 , preprocessing the original voice signal of the user to obtain a processed voice signal, wherein the preprocessing includes filtering and gain adjustment.
[0034] For example, in some embodiments of the present application, a high-quality processed speech signal may be obtained by filtering and gain-adjusting the original speech signal.
[0035] In some embodiments of the present application, S210 may include: dividing the original speech signal into multiple frames; filtering each frame in the multiple frames to obtain a filtered frame; adjusting the gain of each frame in the filtered frame based on a preset target energy range to obtain the processed speech signal.
[0036] For example, in some embodiments of the present application, background noise in the original speech signal is eliminated by a filter or other method to improve the signal quality. For example, an adaptive filter is used, and the filter parameters are first initialized, and the initial weight and learning rate of the filter are set. In this process, the filter parameters can be adjusted first, and the collected historical speech signal is used as a sample. Based on the error between the output of the historical speech signal after filtering and the expected signal, the filter parameters are updated using the LMS algorithm to set a more appropriate initial weight and learning rate.
[0037] During filtering, the continuous original speech signal is divided into multiple small frames (e.g., 25ms per frame), each containing a certain number of speech signal samples. An adaptive filter is applied to each frame, automatically adjusting its parameters based on changes in the input speech signal to reduce background noise, resulting in a filtered frame. The signal corresponding to the filtered frame is then amplified or compressed to ensure it remains within a preset amplitude range (as a specific example of a target energy range). Specifically, the energy or amplitude of the signal in each filtered frame is first calculated. Based on the preset amplitude range, the gain of each frame is dynamically adjusted to maintain its amplitude within this range, thereby achieving signal gain adjustment and obtaining the processed speech signal.
[0038] S220: Input the speech coding data and historical conversation data corresponding to the processed speech signal into a large language model to obtain a text vector corresponding to the processed speech signal.
[0039] For example, in some embodiments of the present application, for different business scenarios to which the speech belongs (e.g., insurance, banking telemarketing), by writing appropriate prompts and using the speech coding data and historical conversation data corresponding to the processed speech signal as input to the large language model, the large language model outputs the corresponding text vector in combination with the contextual conversation (only the text and punctuation corresponding to the speech is output, without additional text), thereby enhancing the recognition effect by providing contextual text semantic information. The speech coding data corresponding to the processed speech signal can be obtained through processing by a text input module, which can use transformer encoding or other encoding methods.
[0040] For example, in a car model collection business scenario, the salesperson's line of speech might be: "Hello, which brand of car are you interested in?" The user's line of speech is: "Geek." If the speech "jike" is recognized alone, it would be difficult to identify it as "geek." However, when combined with the context, the large language model recognizes that the user's "jike" refers to a specific car brand, thus increasing the probability of recognizing "geek."
[0041] S230: Perform feature extraction on the processed speech signal to obtain a speech feature vector.
[0042] For example, in some embodiments of the present application, the features in the processed speech signal are extracted by a speech feature extraction algorithm or Mel-frequency cepstral coefficients to obtain a speech feature vector.
[0043] In some embodiments of the present application, S230 may include: dividing the processed speech signal according to a preset length to obtain multiple signal segments; adding a window function and performing Fourier transform on each signal segment in the multiple signal segments to obtain speech frequency domain features; performing feature processing on the speech frequency domain features to obtain the speech feature vector.
[0044] For example, in some embodiments of the present application, the processed speech signal is divided into multiple segments of a fixed length (as a specific example of a preset length). A window function is applied to each segment to reduce boundary effects. Each segment is then Fourier transformed to obtain a speech frequency domain feature representation. The spectrum of the speech frequency domain features is then mapped to a set scale (e.g., the Mel scale) and filtered using a triangular filter. The logarithm of the power spectrum is then taken to obtain the logarithmic power. Finally, the logarithmic power is transformed, and the first N data points are taken as the speech feature vector.
[0045] S240, using a pre-trained speech recognition module to process the target vector sequence obtained by concatenating the speech feature vector and the text vector to obtain a text sequence; wherein the speech recognition module includes a plurality of pre-trained encoder layers and a plurality of decoder layers.
[0046] For example, in some embodiments of the present application, a speech recognition module composed of a trained encoder-decoder processes a target vector sequence spliced from the results of a large language model and feature extraction to obtain a text sequence. The splicing method is concat. The encoder can be an encoder based on a Gate Recurrent Unit (GRU): in certain tasks, GRU can simplify the structure of an LSTM (Long Short-Term Memory); the decoder can be a bidirectional LSTM / GRU decoder: bidirectional LSTM / GRU can simultaneously process the preceding and following relationships of the sequence during the decoding process, which helps to generate a more coherent text sequence. It should be understood that the types of encoders and decoders can be flexibly selected according to the actual application scenario, and the embodiments of the present application are not specifically limited here.
[0047] In some embodiments of the present application, S240 may include: S241: Perform feature processing on the target vector sequence using the multiple encoder layers to obtain a hidden layer speech state sequence.
[0048] For example, in some embodiments of the present application, the encoder is responsible for extracting high-level representations of features within the input target vector sequence and generating hidden layer speech state sequences that can better capture the semantic information and structural features in the original speech signal.
[0049] Specifically, S241 may include: each encoder layer in the multiple encoder layers performs the following operations until the last encoder layer outputs the hidden layer speech state sequence: performing high-dimensional space embedding processing on the target vector sequence and adding position coding to obtain a time series vector; performing a linear transformation on the time series vector to obtain a first linear transformation matrix; based on the first linear transformation matrix and the feedforward neural network, generating an intermediate state sequence.
[0050] For example, in some embodiments of the present application, since the speech recognition module contains multiple encoder layers, any encoder layer is taken as an example below to illustrate the sequence processing process of the encoder layer.
[0051] For example, the target vector sequence is input into a linear layer in the encoder and embedded into a high-dimensional space. Positional encoding is then added to the input target vector sequence to generate a time series vector, preserving the sequential information of the time series. A multi-head self-attention mechanism is used to perform a linear transformation on the embedded time series vector, resulting in query, key, and value matrices (as a specific example of the first linear transformation matrix). The dot product of the query and key matrices is taken and divided by a scaling factor to calculate the attention score, which is then normalized using a softmax function to obtain the attention weights. The attention weights are then weighted and summed over the value matrix to obtain the attention output. The attention output is passed through a feedforward fully connected network (i.e., a feedforward neural network consisting of two linear layers and a ReLU activation function), followed by residual connections and layer normalization, to obtain the intermediate state sequence output by any encoder layer. By repeating these steps, the hidden layer speech state sequence output by the last encoder layer is obtained.
[0052] S242: Perform feature processing on the hidden layer speech state sequence using the multiple decoder layers to generate the text sequence.
[0053] For example, in some embodiments of the present application, the decoder layer adopts a structure with an attention mechanism; the introduction of the attention mechanism allows the decoder to focus on different parts of the intermediate sequence of the input when generating each character or word, thereby more accurately capturing the alignment relationship between the input and output and obtaining a more accurate text sequence.
[0054] Specifically, S242 may include: each decoder layer in the multiple decoder layers performs the following operations until the last decoder layer outputs the text sequence: embedding the hidden layer speech state sequence into a high-dimensional space and adding position encoding to obtain a coding sequence; performing a linear transformation on the coding sequence to obtain a second linear transformation matrix; based on the second linear transformation matrix and the feedforward fully connected network, generating a text intermediate sequence.
[0055] For example, in some embodiments of the present application, since the speech recognition module contains multiple decoder layers, any decoder layer is taken as an example below to illustrate the sequence processing process of the decoder layer.
[0056] Specifically, the hidden layer speech state sequence is embedded into a high-dimensional space as input, and positional encoding is added to obtain an encoded sequence. The embedded encoded sequence is linearly transformed using masked multi-head self-attention (MMSA) to obtain query, key, and value matrices (as a specific example of the second linear transformation matrix). The dot product of the query and key matrices is taken and divided by the scaling factor to calculate the attention score, which is then normalized using the softmax function to obtain the attention weight. The attention weights are weighted and summed over the value matrix to obtain the attention output. Multi-head attention is calculated using the decoder's query and the encoder's key and value. The attention output is passed through a feed-forward fully connected network (e.g., comprising two linear layers and a ReLU activation function), followed by residual connections and layer normalization, to obtain the intermediate text sequence output by any decoder layer. By repeating these steps, the text sequence output by the last decoder layer is obtained.
[0057] S250: Clean and format the text sequence to obtain text data corresponding to the original speech signal.
[0058] For example, in some embodiments of the present application, relatively standard text data can be obtained by cleaning and standardizing the text sequence.
[0059] In some embodiments of the present application, S250 may include: cleaning spaces and spelling errors in the text sequence to obtain a cleaned text sequence; and formatting the cleaned text sequence using a regular expression to obtain the text data.
[0060] For example, in some embodiments of the present application, the text sequence output by the decoder layer is cleaned up to remove extra spaces and correct spelling errors to obtain a cleaned text sequence. The cleaned text sequence is formatted using regular expressions to add punctuation and paragraph separators to obtain text data with good readability and clear formatting.
[0061] In some embodiments of the present application, a speech recognition module composed of an encoder and a decoder replaces the model in the prior art, which can have a faster training speed and can also effectively extract high-level hidden features.
[0062] The following is combined with Figure 3 The specific process of large-model-based multimodal speech recognition provided by some embodiments of the present application is exemplified.
[0063] Please see the attached Figure 3 , Figure 3 A flowchart of a method for multimodal speech recognition based on a large model is provided for some embodiments of the present application.
[0064] The above process is explained below as an example.
[0065] S310 , dividing the original speech signal into multiple frames, and filtering each frame to obtain a filtered frame.
[0066] S320 , adjusting the gain of each frame in the filtered frames based on a preset target energy range to obtain a processed speech signal.
[0067] S330: Input the speech coding data and historical conversation data corresponding to the processed speech signal into the large language model to obtain a text vector corresponding to the processed speech signal.
[0068] S340: Extract features from the processed speech signal to obtain a speech feature vector.
[0069] S350: Concatenate the speech feature vector and the text vector to obtain a target vector sequence.
[0070] S360: Perform feature processing on the target vector sequence using multiple encoder layers to obtain a hidden layer speech state sequence.
[0071] S370: Utilize multiple decoder layers to perform feature processing on the hidden layer speech state sequence to generate a text sequence.
[0072] S380: Clean and format the text sequence to obtain text data corresponding to the original speech signal.
[0073] It should be noted that the specific implementation process of S310 to S380 can refer to the method embodiment provided above. To avoid repetition, detailed description is appropriately omitted here.
[0074] Through some of the above-mentioned embodiments of the present application, it can be seen that the embodiments of the present application can better utilize the recognition capabilities of the large speech model by adjusting the appropriate prompt. At the same time, the speech recognition module composed of the encoder-decoder structure can more effectively capture the high-level feature information in the speech signal. After integrating the attention mechanism, it can more finely process the alignment relationship between input and output, thereby significantly improving the recognition accuracy. Combined with the versatility of the large language model, the present application shows stronger adaptability and good generalization performance in different languages, accents, speakers and background noise. The present application uses an end-to-end framework, and feature extraction, encoding, decoding and text generation can be integrated, reducing the reliance on complex feature engineering and simplifying the processing flow. Since the speech recognition module composed of the encoder-decoder structure also performs well on small-scale unlabeled data, it reduces the need for large amounts of labeled data, reduces training costs, and improves feasibility in practical applications; and it can better handle speaker changes, accent deviations and different background noise environments, improving the durability and stability of the system in real-world scenarios.
[0075] Please refer to Figure 4 , Figure 4 The following is a block diagram illustrating the components of a large-model-based multimodal speech recognition apparatus provided in some embodiments of the present application. It should be understood that the large-model-based multimodal speech recognition apparatus corresponds to the aforementioned method embodiment and is capable of executing each of the steps involved in the aforementioned method embodiment. The specific functions of the large-model-based multimodal speech recognition apparatus can be found in the description above, and a detailed description is omitted here to avoid repetition.
[0076] Figure 4The device for multimodal speech recognition based on a large model includes at least one software functional module that can be stored in a memory in the form of software or firmware or solidified in the speech recognition device. The device for multimodal speech recognition based on a large model includes: a preprocessing module 410, which is used to preprocess the user's original speech signal to obtain a processed speech signal, wherein the preprocessing includes: filtering and gain adjustment; a large model processing module 420, which is used to input the speech coding data and historical conversation data corresponding to the processed speech signal into a large language model to obtain a text vector corresponding to the processed speech signal; an extraction module 430, which is used to extract features from the processed speech signal to obtain a speech feature vector; a recognition module 440, which is used to use a pre-trained speech recognition module to process a target vector sequence obtained by splicing the speech feature vector and the text vector to obtain a text sequence; wherein the speech recognition module includes a plurality of pre-trained encoder layers and a plurality of decoder layers; a text processing module 450, which is used to clean and format the text sequence to obtain text data corresponding to the original speech signal.
[0077] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working process of the device described above can refer to the corresponding process in the aforementioned method, and will not be described in detail here.
[0078] Some embodiments of the present application further provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, can implement the operations corresponding to any of the above methods provided in the above embodiments.
[0079] Some embodiments of the present application further provide a computer program product, which includes a computer program, wherein when the computer program is executed by a processor, it can implement the operations corresponding to any of the above methods provided in the above embodiments.
[0080] like Figure 5 As shown, some embodiments of the present application provide an electronic device 500, which includes: a memory 510, a processor 520, and a computer program stored in the memory 510 and executable on the processor 520, wherein the processor 520 can implement a method as described in any of the above embodiments when reading the program from the memory 510 through the bus 530 and executing the program.
[0081] Processor 520 can process digital signals and can include various computing architectures, such as a complex instruction set computer architecture, a reduced instruction set computer architecture, or an architecture that implements a combination of multiple instruction sets. In some examples, processor 520 can be a microprocessor.
[0082] The memory 510 can be used to store instructions executed by the processor 520 or data related to the execution of instructions. These instructions and / or data may include code for implementing some or all functions of one or more modules described in the embodiments of this application. The processor 520 of the embodiment of the present disclosure can be used to execute the instructions in the memory 510 to implement the method shown above. The memory 510 includes dynamic random access memory, static random access memory, flash memory, optical memory, or other memory known to those skilled in the art.
[0083] The foregoing is merely an embodiment of the present application and is not intended to limit the scope of protection of the present application. Various modifications and variations are possible for those skilled in the art. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present application shall be included within the scope of protection of the present application. It should be noted that similar reference numerals and letters represent similar items in the following figures. Therefore, once an item is defined in one figure, it does not need to be further defined or explained in subsequent figures.
[0084] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.
[0085] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply the existence of any such actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or device comprising the element.
Claims
1. A method for multimodal speech recognition based on a large model, characterized in that: include: Preprocessing the user's original voice signal to obtain a processed voice signal, wherein the preprocessing includes: filtering and gain adjustment; Inputting speech coding data and historical conversation data corresponding to the processed speech signal into a large language model to obtain a text vector corresponding to the processed speech signal; Performing feature extraction on the processed speech signal to obtain a speech feature vector; Using a pre-trained speech recognition module to process the target vector sequence obtained by concatenating the speech feature vector and the text vector to obtain a text sequence; wherein the speech recognition module includes a plurality of pre-trained encoder layers and a plurality of decoder layers; The text sequence is cleaned and formatted to obtain text data corresponding to the original speech signal.
2. The method according to claim 1, wherein The preprocessing of the user's original voice signal to obtain a processed voice signal includes: Dividing the original speech signal into multiple frames; filtering each of the plurality of frames to obtain a filtered frame; Based on a preset target energy range, the gain of each frame in the filtered frames is adjusted to obtain the processed speech signal.
3. The method according to claim 1 or 2, wherein: The step of extracting features from the processed speech signal to obtain a speech feature vector comprises: Dividing the processed speech signal according to a preset length to obtain multiple segment signals; Adding a window function to each of the multiple signal segments and performing Fourier transform to obtain a speech domain feature; Feature processing is performed on the audio-domain features of the speech to obtain the speech feature vector.
4. The method according to claim 1 or 2, wherein: The method of using a pre-trained speech recognition module to process a target vector sequence obtained by concatenating the speech feature vector and the text vector to obtain a text sequence includes: Performing feature processing on the target vector sequence using the multiple encoder layers to obtain a hidden layer speech state sequence; The plurality of decoder layers are used to perform feature processing on the hidden layer speech state sequence to generate the text sequence.
5. The method according to claim 4, wherein The step of performing feature processing on the target vector sequence by using the multiple encoder layers to obtain a hidden layer speech state sequence includes: Each encoder layer in the plurality of encoder layers performs the following operations until the last encoder layer outputs the hidden layer speech state sequence: Performing high-dimensional spatial embedding processing on the target vector sequence and adding position coding to obtain a time series vector; Performing a linear transformation on the time series vector to obtain a first linear transformation matrix; Based on the first linear transformation matrix and the feedforward neural network, an intermediate state sequence is generated.
6. The method according to claim 4, wherein The step of performing feature processing on the hidden layer speech state sequence by using the multiple decoder layers to generate the text sequence comprises: Each decoder layer in the plurality of decoder layers performs the following operations until the last decoder layer outputs the text sequence: Embedding the hidden layer speech state sequence into a high-dimensional space and adding position coding to obtain a coding sequence; Performing a linear transformation on the coding sequence to obtain a second linear transformation matrix; Based on the second linear transformation matrix and the feedforward fully connected network, an intermediate text sequence is generated.
7. The method according to claim 1 or 2, wherein: The cleaning and formatting of the text sequence to obtain text data corresponding to the original speech signal includes: Cleaning spaces and spelling errors in the text sequence to obtain a cleaned text sequence; The cleaned text sequence is formatted using a regular expression to obtain the text data.
8. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, wherein the computer program is executed by a processor to perform the method according to any one of claims 1 to 7.
9. An electronic device, characterized in that: The method comprises a memory, a processor, and a computer program stored in the memory and running on the processor, wherein the computer program executes the method according to any one of claims 1 to 7 when run by the processor.
10. A computer program product, characterized in that The computer program product comprises a computer program, wherein the computer program is executed by a processor to perform the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Speech recognition method and computer storage medium
CN114220424A
Speech recognition method and related device
CN115101075A
Speech recognition method and device, equipment and storage medium
CN115512695A
Long context end-to-end speech recognition system
CN116324974A
Speech recognition method, device and equipment and computer readable medium
CN116994591A
Cited By
Cross-modal context speech recognition method and system, and storage medium
CN121662047A