Speech-to-Text Translation Method and System Based on ASR Engine

Through the integration of feature extraction, enhancement and decoding of large pre-trained models and small hot word capture pre-trained models, the problem of low accuracy of banking ASR engine in professional vocabulary translation is solved, and the translation effect is improved.

CN116052670BActive Publication Date: 2025-07-29IND BANK CO +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202211455010.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-21
Publication Date
2025-07-29
Estimated Expiration
2042-11-21

AI Technical Summary

Technical Problem

In the existing bank intelligent customer service system, the general ASR engine has low translation accuracy when processing professional vocabulary and special expression methods, which cannot meet the translation needs of special scenarios in the banking industry.

Method used

Large pre-trained models are used for feature extraction, combined with small hot word capture pre-trained models for feature enhancement, and decoding and fusion are finally performed to improve translation accuracy.

Benefits of technology

It improves the accuracy of translation of special vocabulary for ASR engine in designated fields and improves the effect of voice and text translation in special scenarios in the banking industry.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116052670B_ABST
    Figure CN116052670B_ABST
Patent Text Reader

Abstract

The present invention provides a method and system for speech-to-text translation based on an ASR engine, including: after receiving the speech information of a customer, using a large pre-trained model for feature extraction, transmitting the extracted feature information to a small hot-word capture pre-trained model, and finally splicing the features output by the large pre-trained model and the small hot-word capture pre-trained model, and translating the spliced features to obtain the final result. The present invention improves the translation accuracy of the ASR engine for domain-specific vocabulary.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology. Specifically, it relates to a method and system for speech-to-text translation based on an ASR engine. Background Art

[0002] In recent years, with the increasing complexity and diversification of banking services, the telephone banking service provided by banks to customers is undergoing a process of gradually transforming from a simple manual customer service to an efficient combination of manual customer service + intelligent voice customer service. Among them, the performance of the ASR engine for speech-to-text translation in the intelligent voice customer service system is particularly important. In the context of the increasing diversification, personalization, and customization of customer needs, as well as the continuous increase in targeted customer service scenarios, using a more efficient ASR engine to provide higher-quality speech-to-text translation services can effectively improve the service quality of intelligent customer service, and can also effectively reduce the occurrence of situations such as inaccurate recognition of accurate intentions and failure to provide corresponding services accurately due to inaccurate translation, further enhancing the user experience.

[0003] Patent document CN115039171A (application number: CN202180011577.0) discloses a method (600) that includes obtaining a plurality of training data sets (202), each training data set being associated with a corresponding native language and including a plurality of corresponding training data samples (204). For each corresponding training data sample (204) of each of the training data sets (202) for the corresponding native language: the method includes translating the corresponding transcription of the corresponding native text into a corresponding translated text (121) of the corresponding native language representing the target text and associating the corresponding translated text of the target text with the corresponding audio (210) of the corresponding native language to generate a corresponding standardized training data sample (240). The method further includes using the standardized training data sample to train a multilingual E2E ASR model (300) to predict a speech recognition result (120) in the target text for a corresponding speech utterance (106) spoken in any of the native languages associated with the plurality of training data sets.

[0004] The general ASR translation engine currently used by bank intelligent customer service contains a large pre-trained model based on daily communication corpus. After long-term training with a large amount of corpus for daily communication and general expressions, the overall structure of the model is mature and stable, and it has accurate and efficient performance in most general scenarios. However, in the special targeted scenarios of the banking industry, there are the following disadvantages: There are many business channels and product categories in the banking industry, and professional terms or special expressions are generally used. In the current large pre-trained model included in the ASR engine, since the proportion of training corpus containing professional terms and special expressions is too small, the required features cannot be accurately captured through the corresponding voice feature information, resulting in a high error rate during the translation process and unable to meet the translation requirements of this part of the scenarios. Summary of the Invention

[0005] In view of the deficiencies in the prior art, the purpose of the present invention is to provide a speech-to-text translation method and system based on an ASR engine.

[0006] According to the speech-to-text translation method based on an ASR engine provided by the present invention, it includes:

[0007] Step 1: Obtain an initial training data set for targeted hot words;

[0008] Step 2: Based on the initial training data set, perform feature extraction through a preset large pre-trained model to obtain the output features of the large pre-trained model;

[0009] Step 3: Based on the output features of the large pre-trained model, perform feature enhancement through a preset small hot word capture pre-trained model to obtain vector information with enhanced hot word features;

[0010] Step 4: Perform weighted connection on the features output by the large pre-trained model and the small hot word capture pre-trained model;

[0011] Step 5: Input the features after weighted connection into the large pre-trained model, and then perform decoding fusion through the large pre-trained model and the small hot word capture pre-trained model to obtain the final result.

[0012] Preferably, the said Step 1 includes: According to the specified hot word list, find the data related to the hot words in the hot word list from the set of audio data where errors occur during the translation process of the existing unoptimized ASR engine, and give the corresponding hot word labels and corresponding text information to construct the initial training data set. The hot word labels included in the initial training data set are used for data set classification, and the text information corresponding to the included audio files is used to verify the recognition results.

[0013] Preferably, step 2 includes: using a large pre-trained model in the ASR engine to extract features from the speech data in the obtained initial training dataset, and converting the data information contained in the corresponding audio into corresponding pinyin feature vector information through a preset acoustic model and outputting it to obtain the output features of the large pre-trained model;

[0014] Step 3 includes: inputting the output features of the large pre-trained model into a small hotword capture pre-trained model for further feature extraction, randomly replacing the feature vectors corresponding to at least one word in the input feature vectors with corresponding hotword feature vector information, and then inputting the data to be filled in into the small hotword pre-trained model for semantic recognition to obtain a prediction vector, thereby completing the feature enhancement for the hotword and obtaining the vector information of the hotword feature strengthened output by the small hotword pre-trained model.

[0015] Preferably, step 4 includes: according to the specific business scenario requirements, performing weighted connection on the output features of the large pre-trained model and the output features of the small hotword capture pre-trained model, setting weights, and splicing the prediction vectors of the large pre-trained model and the small hotword capture pre-trained model to obtain the composite enhanced feature vector information. The composite feature information includes the speech text translation features in the general scenario extracted by the large pre-trained model, the speech text translation features for the hotword extracted and enhanced by the small hotword capture pre-trained model, and the speech text translation features of the hotword itself.

[0016] Preferably, step 5 includes: re-inputting the composite enhanced feature vector information into the large pre-trained model, and at the same time decoding through the large pre-trained model and the small hotword capture pre-trained model. The decoded result of the small hotword capture pre-trained model and the decoded result of the large pre-trained model are further fused in such a way that the hotword part of the decoded result of the small hotword capture pre-trained model has a high weight and the rest has a low weight, and the hotword part of the decoded result of the large pre-trained model has a low weight and the rest has a high weight. The output result obtained after fusion is output as the final result.

[0017] According to the speech text translation system based on the ASR engine provided by the present invention, it includes:

[0018] Module M1: Obtain an initial training dataset for targeted hotwords;

[0019] Module M2: Based on the initial training dataset, perform feature extraction through a preset large pre-trained model to obtain the output features of the large pre-trained model;

[0020] Module M3: Based on the output features of the large pre-trained model, perform feature enhancement through a preset small hotword capture pre-trained model to obtain the vector information of the hotword feature strengthened;

[0021] Module M4: Perform weighted connection on the features output by the large pre-trained model and the small hot-word capture pre-trained model;

[0022] Module M5: Input the weighted-connected features into the large pre-trained model, and then perform decoding fusion through the large pre-trained model and the small hot-word capture pre-trained model to obtain the final result.

[0023] Preferably, the module M1 includes: According to the specified hot-word list, find the data related to the hot words in the hot-word list from the set of audio data with errors in the translation process of the existing unoptimized ASR engine, and give the corresponding hot-word tags and corresponding text information to construct the initial training dataset. The hot-word tags included in the initial training dataset are used for dataset classification, and the text information corresponding to the audio files included is used to verify the recognition results.

[0024] Preferably, the module M2 includes: Use the large pre-trained model in the ASR engine to extract features from the speech data in the obtained initial training dataset, and convert the data information contained in the corresponding audio into the corresponding pinyin feature vector information through a preset acoustic model and output it to obtain the output features of the large pre-trained model.

[0025] The module M3 includes: Input the output features of the large pre-trained model into the small hot-word capture pre-trained model for further feature extraction, randomly replace the feature vectors corresponding to at least one word in the input feature vector with the corresponding hot-word feature vector information, and then input the data to be filled in the blank into the small hot-word pre-trained model for semantic recognition to obtain the prediction vector, so as to complete the feature enhancement for the hot words and obtain the vector information of the hot-word feature strengthened output by the small hot-word pre-trained model.

[0026] Preferably, the module M4 includes: According to the specific business scenario requirements, perform weighted connection on the output features of the large pre-trained model and the small hot-word capture pre-trained model, set the weights, splice the prediction vectors of the large pre-trained model and the small hot-word capture pre-trained model to obtain the composite enhanced feature vector information. The composite feature information includes the speech text translation features in the general scenario extracted by the large pre-trained model, the speech text translation features for hot words extracted and enhanced by the small hot-word capture pre-trained model, and the speech text translation features of the hot words themselves.

[0027] Preferably, the module M5 includes: re - inputting the compounded and enhanced feature vector information into the large - scale pre - trained model, and at the same time performing decoding through the large - scale pre - trained model and the small - scale hot - word capture pre - trained model. Further fuse the results decoded by the small - scale hot - word capture pre - trained model and the results decoded by the large - scale pre - trained model in such a way that the hot - word part of the decoding result of the small - scale hot - word capture pre - trained model has high weight and the rest has low weight, while the hot - word part of the decoding result of the large - scale pre - trained model has low weight and the rest has high weight. Then output the fused output result as the final result.

[0028] Compared with the prior art, the present invention has the following beneficial effects:

[0029] The present invention extracts features by using a large - scale pre - trained model, constructs a small - scale hot - word capture pre - trained model for feature enhancement, then obtains enhanced features through feature splicing, and finally uses the enhanced features for output fusion to obtain the final translation result, thus solving the problem of low translation accuracy of the ASR engine for domain - specific vocabulary. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] By reading the following detailed description of non - restrictive embodiments with reference to the accompanying drawings, other features, objects, and advantages of the present invention will become more apparent:

[0031] Figure 1 It is a flowchart of the overall optimization method;

[0032] Figure 2 It is a schematic diagram of the small - model feature enhancement method;

[0033] Figure 3 Schematic diagram of the feature weighted connection method;

[0034] Figure 4 Schematic diagram of the result fusion output method. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0035] The present invention will be described in detail below with reference to specific embodiments. The following embodiments will help those skilled in the art to further understand the present invention, but do not limit the present invention in any form. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several changes and improvements can still be made. These all belong to the protection scope of the present invention.

[0036] Embodiment 1:

[0037] The present invention provides an optimization method for improving the translation recognition accuracy by information transmission between a large pre-trained model based on speech-to-text translation and a small hot-word capture pre-trained model. After receiving the customer's speech information, the large pre-trained model is used for feature extraction, and the extracted feature information is transmitted to the small hot-word capture pre-trained model. Finally, the features output by the large pre-trained model and the small hot-word capture pre-trained model are concatenated, and the concatenated features are translated to obtain the final result, improving the translation accuracy of the ASR engine for domain-specific vocabulary.

[0038] Such as Figure 1 , the present invention specifically includes the following steps:

[0039] Step 1: Obtain the initial training dataset for targeted hot words

[0040] According to the specified hot-word list, from the audio data set where errors occur during the translation process of the existing unoptimized ASR engine, find the data related to the hot words in the hot-word list, and give the corresponding hot-word labels and corresponding text information for constructing the initial training dataset. Among them, the hot-word labels included in the initial training dataset are used for dataset classification, and the text information corresponding to the audio files included is used to verify the recognition results.

[0041] Step 2: Feature extraction of the large pre-trained model

[0042] Use the large pre-trained model in the ASR engine to extract features from the speech data in the initial training dataset obtained in Step 1. After the acoustic model converts the data information contained in the corresponding audio into the corresponding pinyin feature vector information, skip the feature decoding stage of using the language model for pinyin-to-character conversion, and directly output the extracted pinyin feature vector information to obtain the large-model output features, which will be used as the input information for the small hot-word capture pre-trained model.

[0043] Step 3: Feature enhancement of the small hot-word capture pre-trained model

[0044] Such as Figure 2 , input the large-model output features into the small hot-word capture pre-trained model for the next step of feature extraction. Use the corresponding hot-word feature vector information to randomly replace at least one word-corresponding feature vector in the input feature vector, and then input the data to be filled in into the small hot-word pre-trained model for semantic recognition to obtain the prediction vector, completing the feature enhancement for the hot words and obtaining the vector information of the hot-word feature strengthened output by the small model.

[0045] Step 4: Weighted connection of the features of the large and small models

[0046] Such as Figure 3, according to the requirements of specific business scenarios, the output features of the large model and the output features of the small model are weighted and connected. A reasonable weight is set for the prediction vector obtained from the language model part in the large model and the prediction vector output by the small model, and the two prediction vectors are concatenated (it can be done by weighting all components, or by sampling some components in the distribution for weighting in the hot word part) to obtain the enhanced composite feature vector information. The composite feature information includes the speech text translation features in the general scenario extracted by the large model, the speech text translation features for hot words extracted and enhanced by the small model, and the speech text translation features of the hot words themselves. Among them, α + β = 1, γ + δ = 1, and β < α, γ > δ, where β ≥ 0, δ ≥ 0.

[0047] Step 5: Decoding of enhanced features and fusion of decoded outputs

[0048] Such as Figure 4 , the enhanced composite feature vector information is re-input into the large model, and at the same time, decoding is performed through the large model and the small model. The result decoded by the small model and the result decoded by the large model are further fused in such a way that the hot word part of the result decoded by the small model has a high weight and the rest has a low weight, and the hot word part of the result decoded by the large model has a low weight and the rest has a high weight. The output result obtained after fusion is used as the final result for output. Among them, α + β = 1, γ + δ = 1, and β < α, γ > δ, where β ≥ 0, δ ≥ 0.

[0049] Embodiment 2:

[0050] The present invention also provides a speech text translation system based on an ASR engine. The speech text translation system based on an ASR engine can be implemented by executing the process steps of the speech text translation method based on an ASR engine. That is, those skilled in the art can understand the speech text translation method based on an ASR engine as a preferred implementation manner of the speech text translation system based on an ASR engine.

[0051] The speech text translation system based on an ASR engine includes: Module M1: Obtain an initial training data set for targeted hot words; Module M2: Based on the initial training data set, perform feature extraction through a preset large pre-trained model to obtain the output features of the large pre-trained model; Module M3: Based on the output features of the large pre-trained model, perform feature enhancement through a preset small hot word capture pre-trained model to obtain vector information with enhanced hot word features; Module M4: Perform weighted connection on the features output by the large pre-trained model and the small hot word capture pre-trained model; Module M5: Input the weighted connected features into the large pre-trained model, and then perform decoding and fusion through the large pre-trained model and the small hot word capture pre-trained model to obtain the final result.

[0052] The module M1 includes: according to the specified hot word list, find the data related to the hot words in the hot word list from the set of audio data with errors in the translation process of the existing unoptimized ASR engine, and give the corresponding hot word tags and corresponding text information to construct an initial training data set. The hot word tags included in the initial training data set are used for data set classification, and the text information corresponding to the included audio files is used to verify the recognition results.

[0053] The module M2 includes: using the large pre-trained model in the ASR engine, extracting features from the speech data in the obtained initial training data set, and converting the data information contained in the corresponding audio into the corresponding pinyin feature vector information through a preset acoustic model and outputting it to obtain the output features of the large pre-trained model.

[0054] The module M3 includes: inputting the output features of the large pre-trained model into a small hot word capture pre-trained model for further feature extraction, randomly replacing the feature vector corresponding to at least one word in the input feature vector with the corresponding hot word feature vector information, and then inputting the data to be filled in the blank into the small hot word pre-trained model for semantic recognition to obtain a prediction vector, so as to complete the feature enhancement for the hot word and obtain the vector information of the hot word feature strengthened output by the small hot word pre-trained model.

[0055] The module M4 includes: according to the specific business scenario requirements, performing weighted connection on the output features of the large pre-trained model and the output features of the small hot word capture pre-trained model, setting weights, and splicing the prediction vectors of the large pre-trained model and the small hot word capture pre-trained model to obtain the composite enhanced feature vector information. The composite feature information includes the speech text translation features in the general scenario extracted by the large pre-trained model, the speech text translation features for the hot word extracted and enhanced by the small hot word capture pre-trained model, and the speech text translation features of the hot word itself.

[0056] The module M5 includes: re-inputting the composite enhanced feature vector information into the large pre-trained model, and at the same time decoding through the large pre-trained model and the small hot word capture pre-trained model. The decoded result of the small hot word capture pre-trained model and the decoded result of the large pre-trained model are further fused in such a way that the hot word part of the decoded result of the small hot word capture pre-trained model has a high weight and the rest has a low weight, and the hot word part of the decoded result of the large pre-trained model has a low weight and the rest has a high weight. The output result obtained after fusion is output as the final result.

[0057] Those skilled in the art know that, in addition to implementing the systems, devices and their respective modules provided by the present invention in the form of pure computer-readable program codes, the method steps can be logically programmed to enable the systems, devices and their respective modules provided by the present invention to be implemented in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, embedded microcontrollers, etc. Therefore, the systems, devices and their respective modules provided by the present invention can be regarded as a kind of hardware components, and the modules included therein for implementing various programs can also be regarded as the structures within the hardware components; the modules for implementing various functions can also be regarded as either software programs for implementing the methods or the structures within the hardware components.

[0058] The specific embodiments of the present invention have been described above. It should be understood that the present invention is not limited to the above specific embodiments, and those skilled in the art can make various changes or modifications within the scope of the claims, which does not affect the essence of the present invention. Without conflict, the embodiments of the present application and the features in the embodiments can be combined with each other arbitrarily.

Claims

1. A speech-to-text translation method based on an ASR engine, characterized in that, Including: Step 1: Obtain the initial training dataset for targeted hot words; Step 2: Based on the initial training dataset, perform feature extraction through a preset large pre-trained model to obtain the output features of the large pre-trained model; Step 3: Based on the output features of the large pre-trained model, perform feature enhancement through a preset small hot word capture pre-trained model to obtain vector information with enhanced hot word features; Step 4: Perform weighted connection on the features output by the large pre-trained model and the small hot word capture pre-trained model; Step 5: Input the features after weighted connection into the large pre-trained model, and then perform decoding fusion through the large pre-trained model and the small hot word capture pre-trained model to obtain the final result; The said Step 3 includes: Input the output features of the large pre-trained model into the small hot word capture pre-trained model for the next step of feature extraction, randomly replace the feature vectors corresponding to at least one word in the input feature vectors with the corresponding hot word feature vector information, and then input the data to be filled in into the small hot word pre-trained model for semantic recognition to obtain the predicted vectors, so as to complete the feature enhancement for the hot words and obtain the vector information with enhanced hot word features output by the small hot word pre-trained model; The said Step 4 includes: According to the requirements of the specific business scenario, perform weighted connection on the output features of the large pre-trained model and the output features of the small hot word capture pre-trained model, set the weights, splice the predicted vectors of the large pre-trained model and the small hot word capture pre-trained model to obtain the composite enhanced feature vector information, and the composite feature information includes the speech text translation features in the general scenario extracted by the large pre-trained model, the speech text translation features for the hot words extracted and enhanced by the small hot word capture pre-trained model, and the speech text translation features of the hot words themselves; The said Step 5 includes: Re-input the composite enhanced feature vector information into the large pre-trained model, and at the same time perform decoding through the large pre-trained model and the small hot word capture pre-trained model, and further fuse the results decoded by the small hot word capture pre-trained model and the results decoded by the large pre-trained model in such a way that the hot word part of the decoding result of the small hot word capture pre-trained model has a high weight and the rest has a low weight, and the hot word part of the decoding result of the large pre-trained model has a low weight and the rest has a high weight, and take the output result after fusion as the final result for output.

2. The speech-to-text translation method based on the ASR engine according to claim 1, wherein The said Step 1 includes: According to the specified hot word list, find the data related to the hot words in the hot word list from the set of audio data with errors that occur during the translation process of the existing unoptimized ASR engine, and give the corresponding hot word labels and corresponding text information for constructing the initial training dataset. The hot word labels included in the initial training dataset are used for dataset classification, and the text information corresponding to the included audio files is used to verify the recognition results.

3. The method for speech-to-text translation based on an ASR engine according to claim 1, characterized in that, Step 2 includes: using a large pre-trained model in the ASR engine to extract features from the speech data in the obtained initial training dataset, and converting the data information contained in the corresponding audio into corresponding pinyin feature vector information through a preset acoustic model and outputting it to obtain the output features of the large pre-trained model.

4. A speech-to-text translation system based on an ASR engine, characterized in that, including: Module M1: Obtain an initial training dataset for targeted hotwords; Module M2: Based on the initial training dataset, perform feature extraction through a preset large pre-trained model to obtain the output features of the large pre-trained model; Module M3: Based on the output features of the large pre-trained model, perform feature enhancement through a preset small hotword capture pre-trained model to obtain vector information with enhanced hotword features; Module M4: Perform weighted connection on the features output by the large pre-trained model and the small hotword capture pre-trained model; Module M5: Input the weighted-connected features into the large pre-trained model, and then perform decoding fusion through the large pre-trained model and the small hotword capture pre-trained model to obtain the final result; The said Module M3 includes: inputting the output features of the large pre-trained model into the small hotword capture pre-trained model for the next step of feature extraction, randomly replacing the feature vectors corresponding to at least one word in the input feature vectors with the corresponding hotword feature vector information, and then inputting the data to be filled in into the small hotword pre-trained model for semantic recognition to obtain a prediction vector, so as to complete the feature enhancement for the hotwords and obtain the vector information with enhanced hotword features output by the small hotword pre-trained model; The said Module M4 includes: according to the requirements of the specific business scenario, performing weighted connection on the output features of the large pre-trained model and the output features of the small hotword capture pre-trained model, setting weights, and splicing the prediction vectors of the large pre-trained model and the small hotword capture pre-trained model to obtain the composite enhanced feature vector information. The composite feature information includes the speech text translation features in the general scenario extracted by the large pre-trained model, the speech text translation features for hotwords extracted and enhanced by the small hotword capture pre-trained model, and the speech text translation features of the hotwords themselves; The said Module M5 includes: re-inputting the composite enhanced feature vector information into the large pre-trained model, and at the same time performing decoding through the large pre-trained model and the small hotword capture pre-trained model. The results decoded by the small hotword capture pre-trained model and the results decoded by the large pre-trained model are further fused in such a way that the hotword part of the decoding result of the small hotword capture pre-trained model has a high weight and the rest has a low weight, and the hotword part of the decoding result of the large pre-trained model has a low weight and the rest has a high weight. The output result obtained after fusion is output as the final result.

5. The speech-to-text translation system based on the ASR engine according to claim 4, characterized in that The module M1 includes: according to the specified hot word list, find the data related to the hot words in the hot word list from the set of audio data with errors that occur during the translation process of the existing unoptimized ASR engine, and give the corresponding hot word tags and corresponding text information for constructing the initial training data set. The hot word tags included in the initial training data set are used for data set classification, and the text information corresponding to the included audio files is used to verify the recognition results.

6. The speech-to-text translation system based on the ASR engine according to claim 4, characterized in that, The module M2 includes: using the large pre-trained model in the ASR engine, extract the features of the speech data in the obtained initial training data set, and convert the data information contained in the corresponding audio into the corresponding pinyin feature vector information through a preset acoustic model and output it to obtain the output features of the large pre-trained model.

Citation Information

Patent Citations

  • Language independent multilingual modeling using efficient text standardization

    CN115039171A

  • Time domain feature extraction method and system for automatic speech recognition

    CN110660382A

  • End-to-end speech recognition model processing method, speech recognition method and related device

    CN114299930A