Speech recognition method and device, electronic equipment and storage medium
Through the cross-modal feature alignment and loss function optimization methods, combined with speech encoder and domain model, the problem of poor recognition performance of ASR systems in different fields is solved, and efficient understanding of professional terms and recognition accuracy is improved.
Patent Information
- Application Number
- CN202510529488.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-25
- Publication Date
- 2025-07-08
AI Technical Summary
The existing ASR systems have poor recognition performance in different fields and scenarios, especially under professional term understanding and noise interference, and it is difficult to make full use of text prior knowledge in single-modal speech training.
By acquiring the audio feature matrix and transcription text feature matrix of the speech dataset, cross-modal feature alignment and comprehensive loss function optimization are performed, combined with speech encoder and domain big models, match the sequence length of audio and text features, and complete feature mapping through a learnable linear projection layer, combining CTC loss and cosine similarity loss to train the target speech recognition model.
It significantly improves the recognition performance of ASR systems in different fields, enhances the understanding of professional terms and cross-modal consistency, and improves the recognition accuracy and stability.
Smart Images

Figure CN120279918A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of speech recognition, and in particular, to a speech recognition method, apparatus, electronic device, and storage medium. Background Art
[0002] Speech recognition technology, also known as Automatic Speech Recognition (ASR), aims to convert the lexical content in human speech into computer-readable input; this automatic speech recognition technology has important significance in applications in different industrial fields, such as professional document entry, meeting record generation, intelligent assistant systems, etc. However, due to the complexity of professional terms, noise interference in speech signals, and the diversity of expression methods, the recognition accuracy of general ASR systems in specific field environments is relatively low. Therefore, how to improve the understanding ability of ASR systems for domain terms has become the focus and difficulty of current research.
[0003] Currently, supervised learning often relies on large-scale domain speech datasets, but due to the limitations of data privacy protection and annotation costs, there are challenges in obtaining high-quality, large-scale training data. In addition, single-modal speech training is difficult to fully utilize the prior knowledge in text, resulting in insufficient recognition ability of the model for rare terms or expressions.
[0004] Therefore, there is an urgent need for an optimization method that can effectively integrate speech and text modal knowledge to comprehensively improve the recognition performance of ASR systems in different field scenarios. Summary of the Invention
[0005] In view of this, embodiments of the present invention provide a speech recognition method, apparatus, electronic device, and storage medium to solve the problem of poor recognition performance of ASR systems in different field scenarios in the prior art.
[0006] To achieve the above object, embodiments of the present invention provide the following technical solutions:
[0007] A first aspect of an embodiment of the present invention shows a speech recognition method, and the method includes:
[0008] Obtain a speech dataset, and extract an audio feature matrix of the speech dataset, where the speech dataset is speech data in a field;
[0009] Perform feature extraction using a transcription text corresponding to the field corresponding to the speech dataset to obtain a corresponding text feature matrix;
[0010] For the audio feature matrix and the text feature matrix in the same field, perform cross-modal feature alignment on the features in the audio feature matrix and the text feature matrix to obtain aligned features;
[0011] Process based on the alignment features to determine a comprehensive loss function;
[0012] Optimize the initial speech recognition model using the comprehensive loss function to obtain a target speech recognition model, so as to process the speech to be recognized input by the user based on the target speech recognition model to obtain a transcription text corresponding to the speech to be recognized, wherein the initial speech recognition model is pre-trained according to the speech data and text feature matrix in each field.
[0013] Optionally, extracting the audio feature matrix of the speech dataset includes:
[0014] Perform audio noise reduction on the speech dataset to obtain a speech dataset after noise reduction processing;
[0015] Remove the abnormal part in the speech dataset after noise reduction processing to obtain a processed speech dataset;
[0016] Randomly split the processed speech dataset to obtain multiple audios;
[0017] Extract the audio features corresponding to the audio and combine them to obtain an audio feature matrix.
[0018] Optionally, perform feature extraction using the transcription text corresponding to the field of the speech dataset to obtain a corresponding text feature matrix, including:
[0019] Obtain the transcription text of the corresponding field based on the field of the speech dataset;
[0020] Perform feature extraction on the transcription text to obtain text features and construct a text feature matrix corresponding to the text features.
[0021] Optionally, perform cross-modal feature alignment on the features in the audio feature matrix and the text feature matrix to obtain alignment features, including:
[0022] Adjust the lengths of the audio features and text features based on the sequence lengths of the features in the audio feature matrix and the text feature matrix;
[0023] Perform cross-modal feature alignment on the audio features and text features after length adjustment to obtain alignment features.
[0024] Optionally, adjusting the lengths of the audio features and text features based on the sequence lengths of the features in the audio feature matrix and the text feature matrix includes:
[0025] When the sequence length of the audio features in the audio feature matrix is less than the sequence length of the corresponding text features in the text feature matrix, adjust the sequence length of the audio features based on the sequence length of the text features;
[0026] When the sequence length of the audio features in the audio feature matrix is greater than the sequence length of the corresponding text features in the text feature matrix, adjust the sequence length of the text features based on the sequence length of the audio features.
[0027] Optionally, process based on the alignment features to determine a comprehensive loss function, including:
[0028] Calculate the cosine similarity loss between the audio features and the text features in the alignment features;
[0029] Perform cross-modal knowledge distillation based on the cosine similarity loss and the CTC loss to obtain a comprehensive loss function.
[0030] Optionally, use the comprehensive loss function to optimize an initial speech recognition model to obtain a target speech recognition model, including:
[0031] Optimize the mapping from the speech dataset to the transcribed text in the corresponding domain based on the comprehensive loss function;
[0032] Adjust the initial speech recognition model based on the optimized mapping from the speech dataset to the transcribed text in the corresponding domain to obtain a target speech recognition model.
[0033] A second aspect of the embodiments of the present invention shows a speech recognition device, the device includes:
[0034] A speech processing module, configured to obtain a speech dataset and extract an audio feature matrix of the speech dataset, where the speech dataset is speech data in a domain;
[0035] A text processing module, configured to perform feature extraction using the transcribed text corresponding to the domain of the speech dataset to obtain a corresponding text feature matrix;
[0036] A feature alignment module, configured to perform cross-modal feature alignment on the features in the audio feature matrix and the text feature matrix for the same domain to obtain alignment features;
[0037] A determination module, configured to process based on the alignment features to determine a comprehensive loss function;
[0038] An optimization module, configured to optimize an initial speech recognition model by using the comprehensive loss function to obtain a target speech recognition model, so as to process the speech to be recognized input by a user based on the target speech recognition model to obtain a transcription text corresponding to the speech to be recognized, where the initial speech recognition model is pre-trained according to speech data and text feature matrices in each field.
[0039] A third aspect of an embodiment of the present invention shows an electronic device, where the electronic device includes a processor and a memory. The memory is configured to store program codes and data for steering wheel control of a storage medium, and the processor is configured to call program instructions in the memory to execute any one of the speech recognition methods shown in the first aspect of the embodiment of the present invention.
[0040] A fourth aspect of an embodiment of the present invention shows a storage medium, characterized in that the storage medium includes a stored program, where when the program runs, it controls a device where the storage medium is located to execute any one of the speech recognition methods shown in the first aspect of the embodiment of the present invention.
[0041] Based on the speech recognition method, device, electronic device, and storage medium provided in the above embodiments of the present invention, the method includes: obtaining a speech data set, and extracting an audio feature matrix of the speech data set, where the speech data set is speech data in one field; performing feature extraction by using a transcription text corresponding to the field corresponding to the speech data set to obtain a corresponding text feature matrix; for the audio feature matrix and the text feature matrix in the same field, performing cross-modal feature alignment on features in the audio feature matrix and the text feature matrix to obtain aligned features; processing based on the aligned features to determine a comprehensive loss function; optimizing an initial speech recognition model by using the comprehensive loss function to obtain a target speech recognition model, so as to process the speech to be recognized input by a user based on the target speech recognition model to obtain a transcription text corresponding to the speech to be recognized, where the initial speech recognition model is pre-trained according to speech data and text feature matrices in each field. In the embodiment of the present invention, a feature aggregation mechanism is adopted to match the sequence lengths of audio and text features and perform feature mapping, thereby effectively enhancing the cross-modal consistency between audio and text features; and significantly improving the recognition ability of domain-specific terms. Thereby improving the recognition performance of the ASR system in different domain scenarios. Description of the Drawings
[0042] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained according to the provided drawings.
[0043] Figure 1 It is a schematic diagram of a speech recognition method shown in an embodiment of the present invention;
[0044] Figure 2 It is a schematic diagram of the implementation process of speech recognition shown in an embodiment of the present invention;
[0045] Figure 3 It is a schematic structural diagram of a speech recognition device shown in an embodiment of the present invention. Detailed implementation manners
[0046] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0047] The terms "first", "second", "third", "fourth", etc. (if any) in the specification, claims and drawings of this application are used to distinguish similar objects and do not necessarily need to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments described here can be implemented in an order other than that illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0048] It should be noted that in the present invention, the descriptions involving "first", "second", etc. are only for descriptive purposes and should not be construed as indicating or implying their relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of such features. Additionally, the technical solutions between various embodiments may be combined with each other, but it must be based on the ability of those of ordinary skill in the art to implement. When the combination of technical solutions results in contradictions or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection required by the present invention.
[0049] In this application, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the phrase "comprising a..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.
[0050] The present invention improves the understanding and transcription ability of a speech recognition system for domain-specific terms through the deep integration of a speech encoder and a domain text model. The present invention includes steps such as speech data preprocessing, audio feature extraction, text feature extraction, feature alignment, cross-modal knowledge distillation, model optimization, and inference, and finally realizes more accurate industry speech transcription. The specific steps are as follows:
[0051] See Figure 1 , which shows a schematic diagram of a speech recognition method according to an embodiment of the present invention. The method includes:
[0052] Step S101: Obtain a speech data set and extract the audio feature matrix of the speech data set.
[0053] Wherein, the speech data set is a speech data set of a domain;
[0054] It should be noted that the process of specifically implementing step S101 includes the following steps:
[0055] Step S11: Perform audio noise reduction on the speech data set to obtain a speech data set after noise reduction processing.
[0056] In the process of specifically implementing step S11, for each piece of speech in the speech dataset, perform multi-speech signal recognition on the speech to determine speech signals such as background noise, reverberation, device noise, the speech of other speakers, and the target speech of the target speaker; then, remove the speech signals corresponding to background noise, reverberation, the speech of other speakers, and / or device noise to obtain a denoised speech dataset.
[0057] Optionally, the self-attention mechanism and time-frequency domain feature enhancement ability in the speech separation model can be used to perform multi-speech signal recognition on the speech, that is, to distinguish speech signals such as background noise, reverberation, device noise, and target speech.
[0058] It should be noted that the speech separation model is pre-trained according to the speech signals of multi-speech.
[0059] Step S12: Remove the abnormal parts in the denoised speech dataset to obtain a processed speech dataset.
[0060] In the process of specifically implementing step S12, for each piece of speech, first detect the start position and end position of the speech, and remove the silence, that is, the blank part, within this speech segment based on the start position and end position of the speech; then, based on the speech data key table, remove the parts different from the speech data key table within the speech segment, that is, the irrelevant parts, to obtain a processed speech dataset.
[0061] It should be noted that the speech data key table is composed of relevant statements, words, etc. recorded in the corresponding field.
[0062] Optionally, the voice activity detection (VAD) technology can be used to detect the start position and end position of the speech.
[0063] Step S13: Randomly split the processed speech dataset to obtain multiple audio segments.
[0064] In the process of specifically implementing step S13, randomly split each piece of speech in the processed speech dataset to ensure that the audio length of each segment is within a preset range, and use the segments within the preset range as audio segments.
[0065] It should be noted that the preset range is set by technicians in advance based on multiple experiments. Generally, it can be set between 5 and 20 seconds to make the data more uniform and adapt to speech inputs with different speech rates and expression styles.
[0066] Step S14: Extract the audio features corresponding to the audio segments and combine them to obtain an audio feature matrix.
[0067] In the process of specifically implementing step S14, the pre-trained speech encoder whisper-large-v3 is used to encode the input domain speech data, that is, the above-mentioned multiple audio files, to extract high-dimensional audio feature representations, that is, to capture phonemes, prosody, and context information in the speech; that is, the multiple audio files are used as input and processed through multiple deep learning layers in the speech encoder to extract the hidden state output by the last layer as the audio feature matrix.
[0068] It should be noted that the audio feature matrix retains the temporal information of the speech and the high-dimensional representation of the speech signal, providing a basis for subsequent cross-modal alignment.
[0069] In the embodiment of the present invention, for the original speech data set of each domain, the original domain speech data is processed such as noise reduction, segmentation, and voice detection. By optimizing the quality and consistency of the domain speech data, high-quality input is provided for subsequent feature extraction, model training, and speech recognition, improving the adaptability and accuracy of the ASR system to the domain scenario.
[0070] Step S102: Use the transcription text corresponding to the domain corresponding to the speech data set to extract features, and obtain the corresponding text feature matrix.
[0071] It should be noted that the process of specifically implementing step S102 includes the following steps:
[0072] Step S21: Obtain the transcription text of the corresponding domain based on the domain of the speech data set.
[0073] In the process of specifically implementing step S21, for the speech data set of each domain in advance, each speech in the speech data set is recognized based on the initial speech recognition model to obtain the transcription text corresponding to the speech data set; the number of the transcription texts is multiple; then the corresponding relationship between the speech data sets of each domain and the transcription texts is stored, so as to subsequently find the transcription text corresponding to the domain of the speech data set based on the corresponding relationship between the speech data sets of each domain and the transcription texts. Here, the transcription text is a set of transcription texts.
[0074] It should be noted that each transcription text has a corresponding speech data;
[0075] Step S22: Extract features from the transcription text to obtain text features, and construct a text feature matrix corresponding to the text features;
[0076] In the process of specifically implementing step S22, a domain large model (such as a medical large model can be used in the medical field) is used to encode the corresponding transcribed text. That is, taking the transcribed text as input, it is input into multiple in-depth learning layers of the domain large model for processing, and the hidden state output by the last layer is obtained as the text feature; and all the text features are used as a text feature matrix, which contains rich domain semantic information to optimize the text understanding ability of the ASR system.
[0077] It should be noted that the domain large model is an encoder.
[0078] It should be noted that since each transcribed text has a corresponding speech data, there are corresponding audio features for the text features in the text feature matrix.
[0079] Step S103: For the audio feature matrix and the text feature matrix in the same domain, perform cross-modal feature alignment on the features in the audio feature matrix and the text feature matrix to obtain aligned features;
[0080] It should be noted that in the process of specifically implementing step S103, the following steps are included.
[0081] Step S31: Adjust the lengths of the audio features and the text features based on the sequence lengths of the features in the audio feature matrix and the text feature matrix.
[0082] It should be noted that the present invention proposes a temporal feature aggregation mechanism for matching the sequence lengths of audio features and text features. Specifically, in the process of specifically implementing step S31, the following steps are included.
[0083] Step S41: Determine whether the sequence length of the audio features in the audio feature matrix is less than the sequence length of the corresponding text features in the text feature matrix. If so, execute step S42. If not, when the sequence length of the audio features in the audio feature matrix is greater than the sequence length of the corresponding text features in the text feature matrix, execute step S43. If the sequence length of the audio features in the audio feature matrix is equal to the sequence length of the corresponding text features in the text feature matrix, directly execute step S32.
[0084] It should be noted that since the length corresponding to each audio feature is within a preset range, it can be considered that the sequence lengths of each audio feature are the same.
[0085] In the process of specifically implementing step S41, for each audio feature in the audio feature matrix and the corresponding text feature in the text feature matrix, for the audio feature, it is determined whether the sequence length of the audio feature is less than the sequence length of the corresponding text feature; if the sequence length of the corresponding text feature is less than the sequence length of the audio feature, step S42 is executed; if the sequence length of the corresponding text feature is greater than the sequence length of the audio feature, steps S33 to S43 are executed; if the sequence length of the audio feature in the audio feature matrix is equal to the sequence length of the corresponding text feature in the text feature matrix, step S32 is directly executed.
[0086] Step S42: Adjust the sequence length of the audio feature based on the sequence length of the text feature;
[0087] In the process of specifically implementing step S42, first, divide the sequence length of the audio feature according to the sequence length of the text feature to obtain multiple segments of audio features; that is to say, divide the sequence length of the audio feature into segments corresponding to the sequence length of the text feature to obtain multiple segments of audio features.
[0088] Next, based on the average value of the sequence lengths of the multiple segments of audio features, that is, take the average value of each segment to aggregate it, to obtain the aggregated audio feature, that is, determine the length of the aggregated audio feature to ensure the alignment of audio and text features in the time dimension.
[0089] For example: if the sequence length of the audio feature sequence is T2 and the sequence length of the text feature sequence is T1, divide T2 into T1 segments to obtain T1 segments of audio features; then, take the average value of each T1 segment of audio features for aggregation to determine the length of the aggregated audio feature.
[0090] Step S43: Adjust the sequence length of the text feature based on the sequence length of the audio feature.
[0091] It should be noted that the specific process of specifically implementing step S43 is the same as the specific process of the above step S32, and reference can be made to each other.
[0092] Step S32: Perform cross-modal feature alignment on the audio feature and text feature after length adjustment to obtain the aligned feature.
[0093] In the process of specifically implementing step S32, use a learnable linear projection layer to map the audio feature after length adjustment into the semantic space of the corresponding text feature to obtain the aligned feature, so as to complete the cross-modal feature mapping from audio to text, make the audio feature more conform to the distribution of the text feature, improve the cross-modal consistency of speech recognition, and reduce the modal difference.
[0094] It should be noted that the alignment features include audio features and text features in the same semantic space as the audio features.
[0095] The sequence length of the audio feature refers to the length of a series of frames into which the audio signal is segmented during processing.
[0096] The sequence length of the text feature refers to the text length of each recorded text.
[0097] Step S104: Process based on the alignment features to determine the comprehensive loss function.
[0098] In the specific process of implementing step S104, the following steps are included.
[0099] Step S51: Calculate the cosine similarity loss between the audio feature and the text feature in the alignment features;
[0100] In the specific process of implementing step S51, the vector representation H of the audio feature in the alignment features SM , and the vector representation H of the text feature LM are substituted into formula (1) for processing to obtain the corresponding cosine similarity loss, enabling the audio feature to gradually approach the text feature in direction, thereby optimizing the cross-modal alignment ability.
[0101] Formula (1):
[0102]
[0103] Where is the corresponding cosine similarity loss.
[0104] Step S52: Perform cross-modal knowledge distillation based on the cosine similarity loss and the CTC loss to obtain the comprehensive loss function.
[0105] In the specific process of implementing step S52, the cosine similarity loss , and (Connectionist Temporal Classification Loss, CTC loss) are subjected to cross-modal knowledge distillation to optimize the mapping from speech to text, that is, it is substituted into formula (2) for calculation to obtain the comprehensive loss function.
[0106] Formula (2):
[0107]
[0108] Where is the comprehensive loss function, and λ is a hyperparameter used to balance the CTC loss of the ASR task and the cross-modal knowledge distillation loss;
[0109] It should be noted that the cross-modal knowledge distillation technology is a technology for transferring knowledge between different modalities. It allows knowledge to be distilled from an audio recognition task to another task, text classification. This method can improve the performance of cross-modal tasks, enable the sharing of knowledge between different tasks, and enhance the efficiency of cross-modal learning.
[0110] The CTC loss estimates several possible cases based on the speech dataset and the corresponding transcription text in the corresponding field, and then calculates the function corresponding to the one with the highest probability according to the probabilities of these possible cases.
[0111] The embodiment of the present invention uses the cosine similarity loss to optimize the feature mapping from speech to text, and at the same time combines the CTC loss to improve the accuracy and stability of industrial speech recognition.
[0112] Step S105: Optimize the initial speech recognition model using the comprehensive loss function to obtain a target speech recognition model, so as to process the speech to be recognized input by the user based on the target speech recognition model to obtain a transcription text corresponding to the speech to be recognized.
[0113] Among them, the initial speech recognition model is pre-trained according to the speech data and text feature matrix in each field.
[0114] It should be noted that in the process of specifically implementing step S105, the following steps are included.
[0115] Step S61: Optimize the mapping from the speech dataset to the transcription text in the corresponding field based on the comprehensive loss function;
[0116] In the process of specifically implementing step S61, for each audio feature of the speech dataset, the comprehensive loss function is used to determine the text feature in the transcription text in the corresponding field that best matches the audio feature, and a mapping from the audio feature in the speech dataset to the text feature in the transcription text in the corresponding field is established. That is to say, the comprehensive loss function can minimize the loss of the speech dataset in model training.
[0117] Step S62: Adjust the initial speech recognition model based on the optimized mapping from the speech dataset to the transcription text in the corresponding field to obtain a target speech recognition model.
[0118] In the process of specifically implementing step S62, the initial speech recognition model is retrained using the mapping of the optimized speech dataset to the transcription text in the corresponding field until the output transcription text is consistent with the transcription text mapped to the speech data. At this time, the initial speech recognition model is determined as the target speech recognition model. If they are inconsistent, the initial speech recognition model is retrained using the mapping of the optimized speech dataset to the transcription text in the corresponding field.
[0119] In one embodiment, the process of processing the speech to be recognized input by the user based on the target speech recognition model to obtain the transcription text corresponding to the speech to be recognized includes:
[0120] Obtain the speech to be recognized input by the user; first, process the speech to be recognized according to the methods shown in steps S11 to S13 to obtain the corresponding audio features; and input them into the target speech recognition model so that the target speech recognition model recognizes the speech to be recognized and obtains the transcription text corresponding to the speech to be recognized.
[0121] In another embodiment, the process of processing the speech to be recognized input by the user based on the target speech recognition model to obtain the transcription text corresponding to the speech to be recognized includes:
[0122] Obtain the speech to be recognized input by the user; and directly input it into the target speech recognition model so that the target speech recognition model recognizes the speech to be recognized and obtains the transcription text corresponding to the speech to be recognized.
[0123] Optionally, based on the specific implementation process of the above steps S101 to S105, it can also be illustrated by Figure 2 description.
[0124] In the embodiments of the present invention, by aligning the audio features generated by the speech encoder with the text features generated by the domain large model; adopting a feature aggregation mechanism to match the sequence lengths of the audio and text features, and completing the feature mapping through a learnable linear projection layer, thereby effectively enhancing the cross-modal consistency of the audio and text features. Training the ASR model by combining the CTC loss and the cosine similarity loss enables its recognition ability for domain-specific terms to be significantly improved. This method does not require modifying the ASR model architecture, only adding loss calculations, is compatible with existing speech recognition systems, and improves the automatic speech recognition ability in domain-specific customized scenarios.
[0125] Based on the speech recognition method shown in the above embodiments of the present invention, correspondingly, the embodiments of the present invention also correspondingly show a structural schematic diagram of a speech recognition device, as Figure 3 shown, the device includes:
[0126] A voice processing module 301 is configured to obtain a voice data set and extract an audio feature matrix of the voice data set, where the voice data set is voice data in a field;
[0127] A text processing module 302 is configured to perform feature extraction by using a transcription text corresponding to the field corresponding to the voice data set to obtain a corresponding text feature matrix;
[0128] A feature alignment module 303 is configured to perform cross-modal feature alignment on the features in the audio feature matrix and the text feature matrix for the audio feature matrix and the text feature matrix in the same field to obtain aligned features;
[0129] A determination module 304 is configured to process based on the aligned features to determine a comprehensive loss function;
[0130] An optimization module 305 is configured to optimize an initial speech recognition model by using the comprehensive loss function to obtain a target speech recognition model, so as to process a to-be-recognized voice input by a user based on the target speech recognition model to obtain a transcription text corresponding to the to-be-recognized voice, where the initial speech recognition model is pre-trained according to voice data and text feature matrices in each field.
[0131] For the specific principles and execution processes of each unit in the speech recognition device disclosed in the embodiments of the present invention above, they are the same as the corresponding content in the speech recognition method provided in the embodiments of the present invention above. Refer to the corresponding parts in the speech recognition method disclosed in the embodiments of the present invention above, and details are not described herein again.
[0132] In the embodiments of the present invention, by aligning the audio features generated by a voice encoder with the text features generated by a domain large model; adopting a feature aggregation mechanism to match the sequence lengths of the audio and text features, and completing feature mapping through a learnable linear projection layer, thereby effectively enhancing the cross-modal consistency between the audio and text features. Combining the CTC loss and the cosine similarity loss to train the initial ASR model, so that its recognition ability for domain-specific terms is significantly improved. This method does not require modifying the architecture of the initial ASR model, and improves the automatic speech recognition ability in domain customization scenarios.
[0133] Optionally, based on the speech recognition device shown in the embodiments of the present invention above, the voice processing module 301 that extracts the audio feature matrix of the voice data set is specifically configured to:
[0134] Perform audio noise reduction on the voice data set to obtain a voice data set after noise reduction processing;
[0135] Remove abnormal parts in the voice data set after noise reduction processing to obtain a processed voice data set;
[0136] Randomly split the processed speech data set to obtain multiple audio files;
[0137] Extract the audio features corresponding to the audio files and combine them to obtain an audio feature matrix.
[0138] Optionally, based on the speech recognition device shown in the embodiments of the present invention above, the text processing module 302 is specifically configured to:
[0139] Obtain a transcription text corresponding to the field based on the field of the speech data set;
[0140] Extract text features from the transcription text and construct a text feature matrix corresponding to the text features.
[0141] Optionally, based on the speech recognition device shown in the embodiments of the present invention above, the feature alignment module 303 for cross-modal feature alignment of the features in the audio feature matrix and the text feature matrix is specifically configured to:
[0142] Adjust the lengths of the audio features and the text features based on the sequence lengths of the features in the audio feature matrix and the text feature matrix;
[0143] Perform cross-modal feature alignment on the audio features and the text features after length adjustment to obtain aligned features.
[0144] Among them, adjusting the lengths of the audio features and the text features based on the sequence lengths of the features in the audio feature matrix and the text feature matrix includes:
[0145] If the sequence length of the audio features in the audio feature matrix is less than the sequence length of the corresponding text features in the text feature matrix, adjust the sequence length of the audio features based on the sequence length of the text features;
[0146] If the sequence length of the audio features in the audio feature matrix is greater than the sequence length of the corresponding text features in the text feature matrix, adjust the sequence length of the text features based on the sequence length of the audio features.
[0147] Optionally, based on the speech recognition device shown in the embodiments of the present invention above, the determination module 304 is specifically configured to:
[0148] Calculate the cosine similarity loss between the audio features and the text features in the aligned features;
[0149] Perform cross-modal knowledge distillation based on the cosine similarity loss and the CTC loss to obtain a comprehensive loss function.
[0150] Optionally, based on the voice recognition device shown in the embodiments of the present invention above, an optimization module 305 for obtaining a target voice recognition model by optimizing an initial voice recognition model using the comprehensive loss function is specifically configured to:
[0151] Optimize the mapping of the voice data set to the transcription text in the corresponding field based on the comprehensive loss function;
[0152] Adjust the initial voice recognition model based on the optimized mapping of the voice data set to the transcription text in the corresponding field to obtain a target voice recognition model.
[0153] An embodiment of the present application provides an electronic device, which includes a processor and a memory. The memory is used to store voice recognition program codes and data, and the processor is used to call the program instructions in the memory to execute the steps shown in the voice recognition method in the above embodiments.
[0154] An embodiment of the present invention provides a storage medium, which includes the electronic device provided in the embodiment of the present application above, and this electronic device is used to execute the voice recognition method disclosed in the embodiment of the present application.
[0155] Each embodiment in this specification is described in a progressive manner. For the same or similar parts between the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for a system or system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and for the relevant parts, reference can be made to the partial description of the method embodiment. The systems and system embodiments described above are only illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement without creative work.
[0156] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed in this article can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.
[0157] The foregoing description of the disclosed embodiments enables those skilled in the art to practice or use the present invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Thus, the present invention is not intended to be limited to the embodiments shown herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A speech recognition method, characterized in that, The method includes: Obtaining a speech dataset and extracting an audio feature matrix of the speech dataset, where the speech dataset is speech data in a certain domain; Performing feature extraction using a transcription text corresponding to the domain of the speech dataset to obtain a corresponding text feature matrix; For the audio feature matrix and the text feature matrix in the same domain, performing cross-modal feature alignment on the features in the audio feature matrix and the text feature matrix to obtain aligned features; Processing based on the aligned features to determine a comprehensive loss function; Optimizing an initial speech recognition model using the comprehensive loss function to obtain a target speech recognition model, so as to process the speech to be recognized input by the user based on the target speech recognition model to obtain a transcription text corresponding to the speech to be recognized, where the initial speech recognition model is pre-trained according to the speech data and the text feature matrix in each domain.
2. The method according to claim 1, characterized in that Extracting the audio feature matrix of the speech dataset includes: Performing audio noise reduction on the speech dataset to obtain a speech dataset after noise reduction processing; Removing abnormal parts in the speech dataset after noise reduction processing to obtain a processed speech dataset; Randomly splitting the processed speech dataset to obtain multiple audios; Extracting the audio features corresponding to the audios and combining them to obtain an audio feature matrix.
3. The method according to claim 1, characterized in that, Performing feature extraction using a transcription text corresponding to the domain of the speech dataset to obtain a corresponding text feature matrix, including: Obtaining a transcription text of the corresponding domain based on the domain of the speech dataset; Performing feature extraction on the transcription text to obtain text features and constructing a text feature matrix corresponding to the text features.
4. The method according to claim 1, characterized in that, Performing cross-modal feature alignment on the features in the audio feature matrix and the text feature matrix to obtain aligned features, including: Adjusting the lengths of the audio features and the text features based on the sequence lengths of the features in the audio feature matrix and the text feature matrix; Performing cross-modal feature alignment on the audio features and the text features after length adjustment to obtain aligned features.
5. The method according to claim 4, characterized in that, Adjusting the lengths of the audio features and the text features based on the sequence lengths of the features in the audio feature matrix and the text feature matrix includes: When the sequence length of the audio features in the audio feature matrix is less than the sequence length of the corresponding text features in the text feature matrix, adjusting the sequence length of the audio features based on the sequence length of the text features; When the sequence length of the audio features in the audio feature matrix is greater than the sequence length of the corresponding text features in the text feature matrix, adjusting the sequence length of the text features based on the sequence length of the audio features.
6. The method according to claim 1, wherein Processing based on the aligned features to determine a comprehensive loss function, including: Calculating the cosine similarity loss between the audio features and the text features in the aligned features; Performing cross-modal knowledge distillation based on the cosine similarity loss and the CTC loss to obtain a comprehensive loss function.
7. The method according to claim 1, characterized in that Optimizing an initial speech recognition model using the comprehensive loss function to obtain a target speech recognition model, including: Optimize the mapping of the speech dataset to the transcribed text in the corresponding domain based on the comprehensive loss function; Adjust the initial speech recognition model based on the optimized mapping of the speech dataset to the transcribed text in the corresponding domain to obtain a target speech recognition model.
8. A voice recognition device, characterized in that, The device includes: A speech processing module, configured to obtain a speech dataset and extract an audio feature matrix of the speech dataset, where the speech dataset is speech data in a domain; A text processing module, configured to perform feature extraction using the transcribed text corresponding to the domain of the speech dataset to obtain a corresponding text feature matrix; A feature alignment module, configured to perform cross-modal feature alignment on the features in the audio feature matrix and the text feature matrix for the same domain to obtain aligned features; A determination module, configured to process based on the aligned features to determine a comprehensive loss function; An optimization module, configured to optimize an initial speech recognition model using the comprehensive loss function to obtain a target speech recognition model, so as to process the speech to be recognized input by a user based on the target speech recognition model to obtain a transcribed text corresponding to the speech to be recognized, where the initial speech recognition model is pre-trained according to the speech data and the text feature matrix in each domain.
9. An electronic device, characterized in that, The electronic device includes a processor and a memory, where the memory is used to store program codes and data for steering wheel control of the storage medium, and the processor is used to call the program instructions in the memory to execute the speech recognition method according to any one of claims 1-7.
10. A storage medium, characterized in that, The storage medium includes a stored program, where when the program runs, it controls the device where the storage medium is located to execute the speech recognition method according to any one of claims 1-7.