Speech Recognition Evaluation Method, Device, Electronic Device and Storage Medium

By acquiring and translating speech recognition text and tags, and aligning them, inputting the evaluation model to evaluate speech recognition effect, the problem of evaluation misjudgment in the prior art is solved and the accuracy of speech recognition evaluation is improved.

CN119785831BActive Publication Date: 2025-06-27IFLYTEK CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510279973.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-11
Publication Date
2025-06-27
Estimated Expiration
2045-03-11

AI Technical Summary

Technical Problem

In the prior art, word accuracy evaluation methods cannot accurately evaluate the true quality of speech recognition effects, resulting in misjudgment in actual applications.

Method used

By obtaining the speech recognition text and labels of the target speech, translating and aligning, a pre-trained evaluation model is entered to obtain the evaluation results of the speech recognition text.

Benefits of technology

It improves the accuracy of speech recognition evaluation, can more comprehensively evaluate the semantic correctness of speech recognition text, and reduces evaluation misjudgment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119785831B_ABST
    Figure CN119785831B_ABST
Patent Text Reader

Abstract

The present invention provides a method, apparatus, electronic device, and storage medium for speech recognition evaluation, belonging to the technical field of natural language processing. The method includes: obtaining a speech recognition text and a speech recognition text label of a target speech; translating the speech recognition text and the speech recognition text label to obtain a first translation of the speech recognition text and a second translation of the speech recognition text label; aligning the speech recognition text with the first translation to obtain first alignment information, and aligning the speech recognition text label with the second translation to obtain second alignment information; inputting the first alignment information and the second alignment information into an evaluation model to obtain an evaluation result of the speech recognition text; the evaluation model is trained based on first sample alignment information corresponding to a sample speech recognition text, second sample alignment information corresponding to a sample speech recognition text label, and an evaluation result label of the sample speech recognition text. The present invention can improve the accuracy of speech recognition evaluation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of natural language processing, and in particular, to a method, device, electronic device, and storage medium for speech recognition evaluation. Background Art

[0002] With the rapid development of computer and artificial intelligence technologies, the application of speech recognition technology is becoming more and more extensive. In some key fields, the accuracy of speech recognition is crucial. How to accurately evaluate the speech recognition effect is a key problem to be solved urgently. The traditional manual evaluation method has high cost and low efficiency. The traditional character accuracy evaluation method cannot accurately evaluate the true quality of the speech recognition effect, resulting in misjudgment in actual applications. Summary of the Invention

[0003] The present invention provides a method, device, electronic device, and storage medium for speech recognition evaluation to solve the defect that the character accuracy evaluation method in the prior art cannot accurately evaluate the true quality of the speech recognition effect, resulting in misjudgment in actual applications.

[0004] The present invention provides a method for speech recognition evaluation, including:

[0005] Obtaining the speech recognition text of the target speech and determining the speech recognition text label of the target speech;

[0006] Translating the speech recognition text and the speech recognition text label to obtain a first translation of the speech recognition text and a second translation of the speech recognition text label;

[0007] Aligning the speech recognition text with the first translation to obtain first alignment information, and aligning the speech recognition text label with the second translation to obtain second alignment information;

[0008] Inputting the first alignment information and the second alignment information into a pre-trained evaluation model to obtain the evaluation result of the speech recognition text output by the evaluation model;

[0009] Wherein, the evaluation model is trained based on the first sample alignment information between the sample speech recognition text and the corresponding first sample translation, the second sample alignment information between the sample speech recognition text label and the corresponding second sample translation, and the evaluation result label of the sample speech recognition text.

[0010] In some embodiments, the step of inputting the first alignment information and the second alignment information into a pre-trained evaluation model to obtain the evaluation result of the speech recognition text output by the evaluation model includes:

[0011] Based on the evaluation model, feature extraction is performed on the first alignment information to obtain the semantic features of the first translation, and feature extraction is performed on the second alignment information to obtain the semantic features of the second translation. Based on the semantic features of the first translation and the semantic features of the second translation, the semantic similarity between the first translation and the second translation is calculated, and based on the semantic similarity, the intelligibility evaluation result of the speech recognition text is obtained.

[0012] In some embodiments, the calculating the semantic similarity between the first translation and the second translation based on the semantic features of the first translation and the semantic features of the second translation, and obtaining the intelligibility evaluation result of the speech recognition text based on the semantic similarity includes:

[0013] Align each word in each sentence of the first translation with each word in each sentence of the second translation to obtain a plurality of word pairs;

[0014] Based on the semantic features of the first translation and the semantic features of the second translation, determine the semantic features of the two words in each word pair;

[0015] Based on the semantic features of the two words in each word pair, calculate the semantic similarity between the two words in each word pair;

[0016] Determine the weight of each word pair, and based on the semantic similarity between the two words in each word pair and the weight of each word pair, calculate the intelligibility of the speech recognition text.

[0017] In some embodiments, the aligning the speech recognition text with the first translation to obtain the first alignment information includes:

[0018] Input the speech recognition text and the first translation of the speech recognition text into a pre-trained first alignment model to obtain the first alignment information output by the first alignment model;

[0019] Wherein, the first alignment model is trained based on a sample speech recognition text and a first sample translation of the sample speech recognition text, and a first alignment information label of the sample speech recognition text and the first sample translation.

[0020] In some embodiments, the aligning the speech recognition text label with the second translation to obtain the second alignment information includes:

[0021] Input the speech recognition text label and the second translation of the speech recognition text label into a pre-trained second alignment model to obtain the second alignment information output by the second alignment model;

[0022] Among them, the second alignment model is trained based on the sample speech recognition text label, the second sample translation of the sample speech recognition text label, and the second alignment information label of the sample speech recognition text and the second sample translation.

[0023] In some embodiments, translating the speech recognition text and the speech recognition text label to obtain a first translation of the speech recognition text and a second translation of the speech recognition text label includes:

[0024] Using a pre-trained machine translation model to translate the speech recognition text and the speech recognition text label to obtain a first translation of the speech recognition text and a second translation of the speech recognition text label;

[0025] Among them, the machine translation model is trained based on the sample speech recognition text and the sample speech recognition text label, as well as the first translation label of the sample speech recognition text and the second translation label of the sample speech recognition text label.

[0026] In some embodiments, the training process of the evaluation model includes:

[0027] Obtaining the sample speech recognition text of the sample speech, determining the sample speech recognition text label of the sample speech, and determining the evaluation result label of the sample speech recognition text;

[0028] Translating the sample speech recognition text and the sample speech recognition text label to obtain a first sample translation of the sample speech recognition text and a second sample translation of the sample speech recognition text label;

[0029] Aligning the sample speech recognition text and the first sample translation to obtain first sample alignment information, and aligning the sample speech recognition text label and the second sample translation to obtain second sample alignment information;

[0030] Using the first sample alignment information and the second sample alignment information as training samples, and using the evaluation result label of the sample speech recognition text as a sample label to train an initial evaluation model. After the training is completed, the evaluation model is obtained.

[0031] The present invention also provides a speech recognition evaluation device, including:

[0032] An acquisition unit for acquiring the speech recognition text of the target speech and determining the speech recognition text label of the target speech;

[0033] A translation unit, configured to translate the speech recognition text and the speech recognition text label to obtain a first translation of the speech recognition text and a second translation of the speech recognition text label;

[0034] An alignment unit, configured to align the speech recognition text with the first translation to obtain first alignment information, and align the speech recognition text label with the second translation to obtain second alignment information;

[0035] An evaluation unit, configured to input the first alignment information and the second alignment information into a pre-trained evaluation model to obtain an evaluation result of the speech recognition text output by the evaluation model;

[0036] Wherein, the evaluation model is trained based on first sample alignment information between a sample speech recognition text and a corresponding first sample translation, second sample alignment information between a sample speech recognition text label and a corresponding second sample translation, and an evaluation result label of the sample speech recognition text.

[0037] The present invention further provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the speech recognition evaluation method as described in any one of the above is implemented.

[0038] The present invention further provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the speech recognition evaluation method as described in any one of the above is implemented.

[0039] The speech recognition evaluation method, device, electronic device, and storage medium provided by the present invention can improve the accuracy of speech recognition evaluation by obtaining a speech recognition text and a speech recognition text label of a target voice; translating the speech recognition text and the speech recognition text label to obtain a first translation of the speech recognition text and a second translation of the speech recognition text label; aligning the speech recognition text with the first translation to obtain first alignment information, and aligning the speech recognition text label with the second translation to obtain second alignment information; and inputting the first alignment information and the second alignment information into a pre-trained evaluation model to obtain an evaluation result of the speech recognition text. Description of the Drawings

[0040] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0041] Figure 1It is one of the schematic flowcharts of the speech recognition evaluation method provided by an embodiment of the present invention.

[0042] Figure 2 It is a schematic flowchart of translating the speech recognition text of a target speech provided by an embodiment of the present invention.

[0043] Figure 3 It is an example of the first alignment information and the second alignment information provided by an embodiment of the present invention.

[0044] Figure 4 It is the second of the schematic flowcharts of the speech recognition evaluation method provided by an embodiment of the present invention.

[0045] Figure 5 It is the third of the schematic flowcharts of the speech recognition evaluation method provided by an embodiment of the present invention.

[0046] Figure 6 It is a schematic flowchart of the training process of the evaluation model provided by an embodiment of the present invention.

[0047] Figure 7 It is a schematic structural diagram of the speech recognition evaluation device provided by an embodiment of the present invention.

[0048] Figure 8 It is a schematic structural diagram of the electronic device provided by an embodiment of the present invention. Detailed implementation manners

[0049] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below with reference to the accompanying drawings in the present invention. Apparently, the described embodiments are some but not all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without creative efforts shall fall within the protection scope of the present invention.

[0050] The terms "first", "second", etc. in the present invention are used to distinguish similar objects, rather than to describe a specific order or sequence. It should be understood that such terms can be interchanged under appropriate circumstances so that the embodiments of the present invention can be implemented in an order different from those illustrated or described herein, and the objects distinguished by "first" and "second" are usually of the same category, and the number of objects is not limited. For example, the first object can be one or more.

[0051] The existing speech recognition effect evaluation solutions mainly focus on evaluating the accuracy of speech recognition systems through traditional quantitative indicators. These evaluation methods are generally based on the following aspects:

[0052] (1) Word Error Rate (WER):

[0053] The character accuracy rate is the most common speech recognition evaluation metric, which evaluates the accuracy of speech recognition by calculating the difference between the speech recognition result and the reference text. Its advantages are simplicity and intuitiveness, and it has a wide range of applications. However, it cannot capture the accuracy at the semantic level, only evaluates the literal matching degree, and cannot distinguish the impact of misrecognition. For example, the case where the characters are incorrect but the semantics are still correct (which may affect the user experience), it is greatly affected by speech noise or accents, and cannot distinguish minor pronunciation deviations from key semantic errors.

[0054] (2) Character Error Rate (CER):

[0055] The Character Error Rate is similar to the WER, but the Character Error Rate evaluates at the character level rather than the word level. The calculation method of CER is the same as that of WER, except that the text is split into characters instead of words. Its advantage is that for more fine-grained evaluations in speech recognition (such as pinyin, syllables, etc.), CER is more precise and can be applied to some application scenarios that require character-level evaluations. However, like the WER, CER cannot capture semantic-level errors, focuses on literal accuracy, and may be difficult to directly reflect the overall performance of the speech recognition system in the case of long texts or low character accuracy.

[0056] (3) Sentence-level Accuracy Rate (SER):

[0057] Evaluates whether each sentence of the speech recognition exactly matches the reference sentence. If the recognition result is exactly the same as the reference sentence, then the sentence is considered to be recognized correctly. The sentence-level accuracy rate evaluation is applicable to multi-sentence recognition tasks, can evaluate the overall performance of the speech recognition system on multiple sentences, and can quickly provide an evaluation of the overall recognition ability of the speech recognition system. However, it provides less information, cannot reveal specific error types, and may be too strict for sentences with a small number of errors but correct semantics.

[0058] The scheme of evaluating the accuracy of a speech recognition system based on traditional quantitative metrics cannot capture the accuracy at the semantic level, can only reflect the recognition precision, and cannot consider the semantic correctness of the recognition result. Even if a speech recognition system has a high character accuracy rate, it may still not be able to accurately convey the actual meaning of the sentence, especially when dealing with sentences with synonyms or strong context dependence. In addition, traditional methods cannot effectively identify and distinguish the accuracy of key information. In practical applications, errors in certain characters or words do not significantly affect the overall semantics, but errors in some key information will seriously affect the effectiveness of the speech recognition system, and traditional evaluation methods often fail to reflect this situation.

[0059] To this end, the embodiments of the present invention provide a speech recognition evaluation method, apparatus, electronic device, and storage medium. By obtaining the speech recognition text and the speech recognition text label of the target speech; translating the speech recognition text and the speech recognition text label to obtain the first translation of the speech recognition text and the second translation of the speech recognition text label; aligning the speech recognition text with the first translation to obtain the first alignment information, and aligning the speech recognition text label with the second translation to obtain the second alignment information; inputting the first alignment information and the second alignment information into a pre-trained evaluation model to obtain the evaluation result of the speech recognition text, the present invention can improve the accuracy of speech recognition evaluation.

[0060] Figure 1 One of the flow diagrams of the speech recognition evaluation method provided by the embodiments of the present invention. As Figure 1 shown, a speech recognition evaluation method is provided, including the following steps: step 110, step 120, step 130, and step 140. The flow steps of this method are only one possible implementation manner of the present invention.

[0061] Step 110, obtain the speech recognition text of the target speech, and determine the speech recognition text label of the target speech.

[0062] Optionally, use a traditional speech recognition system, a speech recognition method based on deep learning, a professional speech recognition software or platform to obtain the speech recognition text of the target speech.

[0063] Optionally, capture the sound signal emitted by the user through an audio device (such as a microphone) to obtain the target speech, and recognize the target speech based on the speech recognition technology to obtain the speech recognition text.

[0064] Step 120, translate the speech recognition text and the speech recognition text label to obtain the first translation of the speech recognition text and the second translation of the speech recognition text label.

[0065] Optionally, adopt an end-to-end machine translation system based on Transformer to translate the speech recognition text and the speech recognition text label.

[0066] Among them, the end-to-end machine translation system based on Transformer utilizes the self-attention mechanism and parallel computing ability of the Transformer model to directly translate the source language sentence into the target language sentence, which can improve the translation effect and efficiency.

[0067] It should be noted that differences in speech recognition texts often lead to differences in machine translation results. If there are some problems in the speech recognition text, the machine translation system can amplify them to a certain extent through the understanding of semantics by the translation model. By comparing with the translation of the speech recognition text label, the problems can be more significantly discovered.

[0068] Figure 2 This is a schematic flowchart of translating the speech recognition text of the target speech provided by the embodiment of the present invention. As Figure 2 shown, translating the speech recognition text of the target speech includes the following steps:

[0069] Perform speech recognition on the target speech to obtain the speech recognition text of the target speech;

[0070] Perform machine translation on the speech recognition text to obtain the first translation of the speech recognition text.

[0071] For example, the target speech is "I will meet you at the park", the speech recognition text is "I willmeet you at the bark", the corresponding first translation is "I will meet you at the bark", and the second translation of the speech recognition text label is "I will meet you at the park"; the meeting place is the key information. Through machine translation, it can be found that although the character accuracy rate of this speech recognition result is relatively high, there are relatively large problems at the semantic level and the accuracy is relatively low.

[0072] Step 130: Align the speech recognition text and the first translation to obtain the first alignment information, and align the speech recognition text label and the second translation to obtain the second alignment information.

[0073] Optionally, use the Mgiza++ alignment tool to perform word alignment on the speech recognition text and the first translation to obtain the first alignment information, and perform word alignment on the speech recognition text label and the second translation to obtain the second alignment information.

[0074] It should be noted that word alignment refers to the process of corresponding each word or phrase in the source language to a word or phrase in the target language one by one.

[0075] Step 140: Input the first alignment information and the second alignment information into a pre-trained evaluation model to obtain the evaluation result of the speech recognition text output by the evaluation model;

[0076] Among them, the evaluation model is trained based on the first sample alignment information of the sample speech recognition text and the corresponding first sample translation, the second sample alignment information of the sample speech recognition text label and the corresponding second sample translation, and the evaluation result label of the sample speech recognition text.

[0077] Optionally, the first alignment information and the second alignment information are compared based on an evaluation model, the semantic similarity between the two is calculated, and the speech recognition text is evaluated according to the semantic similarity to obtain an evaluation result.

[0078] In an embodiment of the present invention, by obtaining the speech recognition text and the speech recognition text label of the target speech; translating the speech recognition text and the speech recognition text label to obtain the first translation of the speech recognition text and the second translation of the speech recognition text label; aligning the speech recognition text with the first translation to obtain the first alignment information, and aligning the speech recognition text label with the second translation to obtain the second alignment information; inputting the first alignment information and the second alignment information into a pre-trained evaluation model to obtain the evaluation result of the speech recognition text, the semantic-level evaluation of the speech recognition text can be performed, and the accuracy of the speech recognition evaluation is improved.

[0079] In some embodiments, inputting the first alignment information and the second alignment information into a pre-trained evaluation model to obtain the evaluation result of the speech recognition text output by the evaluation model includes:

[0080] Based on the evaluation model, feature extraction is performed on the first alignment information to obtain the semantic features of the first translation, feature extraction is performed on the second alignment information to obtain the semantic features of the second translation, based on the semantic features of the first translation and the semantic features of the second translation, the semantic similarity between the first translation and the second translation is calculated, and based on the semantic similarity, the intelligibility evaluation result of the speech recognition text is obtained.

[0081] Among them, intelligibility refers to a measure of the ability to understand speech under given conditions, and the intelligibility can be quantified by calculating the number of correctly recognized words, phonemes or other language units.

[0082] It can be understood that by adopting an evaluation model, based on the semantic features of the first translation and the semantic features of the second translation, the semantic similarity between the first translation and the second translation is calculated, and then based on the semantic similarity, the speech recognition text can be comprehensively evaluated to obtain the intelligibility evaluation result of the speech recognition text, improving the efficiency and accuracy of the speech recognition evaluation.

[0083] In some embodiments, based on the semantic features of the first translation and the semantic features of the second translation, calculating the semantic similarity between the first translation and the second translation, and based on the semantic similarity, obtaining the intelligibility evaluation result of the speech recognition text includes:

[0084] Align each vocabulary of each sentence of the first translation with each vocabulary of each sentence of the second translation to obtain a plurality of vocabulary pairs;

[0085] Based on the semantic features of the first translation and the semantic features of the second translation, determine the semantic features of the two vocabularies in each vocabulary pair;

[0086] Based on the semantic features of the two words in each word pair, the semantic similarity of the two words in each word pair is calculated;

[0087] The weight of each vocabulary pair is determined, and the intelligibility of the speech recognition text is calculated based on the semantic similarity of the two vocabulary words in each vocabulary pair and the weight of each vocabulary pair.

[0088] Optionally, feature extraction is performed on the first translation and the second translation to obtain semantic features of the first translation and semantic features of the second translation.

[0089] Optionally, the weight of each word pair is determined according to the importance of each word.

[0090] It can be understood that by aligning each word in each sentence of the first translation with each word in each sentence of the second translation, a plurality of word pairs are obtained, the semantic features of the two words in each word pair are determined, and the semantic similarity of the two words in each word pair is calculated, thereby quantifying the differences between different parts of the first translation and the second translation, thereby improving the accuracy and precision of speech recognition evaluation; by determining the weight of each word pair, the intelligibility of the speech recognition text is calculated based on the semantic similarity of the two words in each word pair and the weight of each word pair, and the accuracy of the key information can be identified, especially when the key information is recognized incorrectly or the word accuracy is low but does not affect the overall semantics, misjudgment can be avoided.

[0091] Figure 3 This is an example of the first alignment information and the second alignment information provided by the embodiment of the present invention. Figure 3 As shown, the speech recognition text is “I went to the market to sell vegetables”, the first translation is “I went to the market to sell vegetables”, and the first alignment information includes “I / I, went to / go, the market / market, to sell / sell, vegetables / vegetables”; the speech recognition text label is “I went to the market to buy vegetables”, the second translation is “I went to the market to buyvegetables”, and the second alignment information includes “I / I, went to / go, the market / market, to buy / buy, vegetables / vegetables”.

[0092] In this example, "I go to the market to sell vegetables" and "I go to the market to buy vegetables" differ by only one word, and the word accuracy is relatively high, at 83%. However, key information recognition errors result in completely different semantics, and the intelligibility score is only 50 points.

[0093] In some embodiments, aligning the speech recognition text with the first translation to obtain first alignment information, including:

[0094] Inputting the speech recognition text and the first translation of the speech recognition text into a pre-trained first alignment model to obtain the first alignment information output by the first alignment model;

[0095] Wherein, the first alignment model is trained based on the sample speech recognition text and the first sample translation of the sample speech recognition text, and the first alignment information label of the sample speech recognition text and the first sample translation.

[0096] Optionally, the first alignment model includes a first alignment layer, a first discrimination layer, and a first mapping layer; the first alignment layer is used to perform word alignment on the speech recognition text and the first translation, the first discrimination layer is used to distinguish the speech recognition text and the first translation, and the first mapping layer is used to determine the mapping relationship between each vocabulary of the speech recognition text and each vocabulary of the first translation.

[0097] Optionally, the training process of the first alignment model includes:

[0098] Obtaining the sample speech recognition text and the first sample translation of the sample speech recognition text;

[0099] Determining the first alignment information label of the sample speech recognition text and the first sample translation;

[0100] Using the sample speech recognition text and the first sample translation as training samples, and the first alignment information label of the sample speech recognition text and the first sample translation as sample labels to train an initial first alignment model, and after the training is completed, obtaining the first alignment model.

[0101] It can be understood that, based on the first alignment model, aligning the speech recognition text and the first translation of the speech recognition text to obtain the first alignment information improves the alignment efficiency, facilitates the alignment of large-scale data, and reduces the alignment cost.

[0102] In some embodiments, aligning the speech recognition text label with the second translation to obtain second alignment information, including:

[0103] Inputting the speech recognition text label and the second translation of the speech recognition text label into a pre-trained second alignment model to obtain the second alignment information output by the second alignment model;

[0104] Wherein, the second alignment model is trained based on the sample speech recognition text label and the second sample translation of the sample speech recognition text label, and the second alignment information label of the sample speech recognition text label and the second sample translation.

[0105] Optionally, the second alignment model includes a second alignment layer, a second discrimination layer, and a second mapping layer; the second alignment layer is used to perform word alignment on the speech recognition text label and the second translation; the second discrimination layer is used to discriminate the speech recognition text label and the second translation; and the second mapping layer is used to determine the mapping relationship between each vocabulary of the speech recognition text label and each vocabulary of the second translation.

[0106] Optionally, the training process of the second alignment model includes:

[0107] Obtain a sample speech recognition text label and a second sample translation of the sample speech recognition text label;

[0108] Determine the second alignment information label of the sample speech recognition text label and the second sample translation;

[0109] Use the sample speech recognition text label and the second sample translation as training samples, and use the second alignment information label of the sample speech recognition text label and the second sample translation as sample labels to train the initial second alignment model. After training is completed, the second alignment model is obtained.

[0110] It can be understood that, based on the second alignment model, the speech recognition text label and the second translation of the speech recognition text label are aligned to obtain the second alignment information output by the second alignment model, which improves the alignment efficiency, facilitates the alignment of large-scale data, and reduces the alignment cost.

[0111] In some embodiments, translating the speech recognition text and the speech recognition text label to obtain the first translation of the speech recognition text and the second translation of the speech recognition text label includes:

[0112] Use a pre-trained machine translation model to translate the speech recognition text and the speech recognition text label to obtain the first translation of the speech recognition text and the second translation of the speech recognition text label;

[0113] Among them, the machine translation model is trained based on the sample speech recognition text and the sample speech recognition text label, as well as the first translation label of the sample speech recognition text and the second translation label of the sample speech recognition text label.

[0114] It can be understood that translating the speech recognition text and the speech recognition text label based on the machine translation model improves the translation efficiency.

[0115] Optionally, the training process of the machine translation model includes:

[0116] Obtain a sample speech recognition text and a sample speech recognition text label;

[0117] Determine the first translation label of the sample speech recognition text and the second translation label of the sample speech recognition text label;

[0118] Using the sample speech recognition text and the sample speech recognition text label as training samples, and using the first translation label of the sample speech recognition text and the second translation label of the sample speech recognition text label as sample labels, train the initial machine translation model. After training is completed, a machine translation model is obtained.

[0119] Figure 4 This is the second flowchart of the speech recognition evaluation method provided by the embodiment of the present invention. As Figure 4 shown, in some embodiments, a speech recognition evaluation method is provided, including the following steps:

[0120] Obtain the target speech;

[0121] Determine the speech recognition text and the speech recognition text label of the target speech;

[0122] Based on the machine translation model, translate the speech recognition text and the speech recognition text label to obtain the first translation of the speech recognition text and the second translation of the speech recognition text label;

[0123] Based on the first alignment model, perform word alignment on the speech recognition text and the first translation to obtain the first alignment information. Based on the second alignment model, align the speech recognition text label and the second translation to obtain the second alignment information;

[0124] Based on the evaluation model, compare the first alignment information and the second alignment information to obtain the automatic evaluation result of the speech recognition text;

[0125] Obtain the expert evaluation result of the speech recognition text;

[0126] Based on the expert evaluation result, correct the automatic evaluation result.

[0127] Among them, the expert evaluation result is obtained by performing semantic rationality evaluation based on the first translation of the speech recognition text and the speech recognition text label.

[0128] It can be understood that by comparing the automatic evaluation result of the speech recognition text with the expert evaluation result, the parameters of the evaluation model, the first alignment model, and the second alignment model can be optimized, thereby improving the accuracy of the automatic evaluation.

[0129] Figure 5 This is the third flowchart of the speech recognition evaluation method provided by the embodiment of the present invention. As Figure 5 shown, in some embodiments, a speech recognition evaluation method is provided, including the following steps:

[0130] Obtain the target speech of the speaker;

[0131] Based on a speech recognition system, perform speech recognition on the target speech to obtain the speech recognition text (i.e., the true original text of speech recognition);

[0132] Determine the speech recognition text label of the target speech (i.e., the labeled original text of speech recognition);

[0133] Based on a large model machine translation engine, translate the speech recognition text to obtain the first translation (i.e., the true translation of speech recognition), and translate the speech recognition text label to obtain the second translation (i.e., the labeled translation of speech recognition);

[0134] Based on a pre-trained intelligibility evaluation model, process the speech recognition text and the first translation, as well as the speech recognition text label and the second translation, to obtain the intelligibility score of the speech recognition text.

[0135] Optionally, the intelligibility evaluation model includes a first alignment layer, a second alignment layer, a similarity calculation layer, and an intelligibility evaluation layer.

[0136] Among them, the first alignment layer is used to align the speech recognition text and the first translation to obtain the first alignment information; the second alignment layer is used to align the speech recognition text label and the second translation to obtain the second alignment information; the similarity calculation layer is used to calculate the semantic similarity between the first translation and the second translation based on the first alignment information and the second alignment information; the intelligibility evaluation layer is used to evaluate the intelligibility of the speech recognition text based on the semantic similarity to obtain the intelligibility score of the speech recognition text.

[0137] Among them, the intelligibility evaluation model is trained based on the sample speech recognition text and its corresponding sample first translation, the sample speech recognition text label and its corresponding sample second translation, and the intelligibility score label of the sample speech recognition text.

[0138] Figure 6 It is a schematic flowchart of the training process of the evaluation model provided by the embodiments of the present invention. As Figure 6 shown, in some embodiments, the training process of the evaluation model includes:

[0139] Step 610, obtain the sample speech recognition text of the sample speech, determine the sample speech recognition text label of the sample speech, and determine the evaluation result label of the sample speech recognition text;

[0140] Step 620, translate the sample speech recognition text and the sample speech recognition text label to obtain the first sample translation of the sample speech recognition text and the second sample translation of the sample speech recognition text label;

[0141] Step 630: Align the sample speech recognition text with the first sample translation to obtain the first sample alignment information, and align the sample speech recognition text label with the second sample translation to obtain the second sample alignment information;

[0142] Step 640: Use the first sample alignment information and the second sample alignment information as training samples, and use the evaluation result label of the sample speech recognition text as the sample label to train the initial evaluation model. After the training is completed, the evaluation model is obtained.

[0143] Optionally, based on the machine translation model, translate the sample speech recognition text and the sample speech recognition text label to obtain the first sample translation of the sample speech recognition text and the second sample translation of the sample speech recognition text label.

[0144] Optionally, based on the first alignment model, align the sample speech recognition text with the first sample translation to obtain the first sample alignment information.

[0145] Optionally, based on the second alignment model, align the sample speech recognition text label with the second sample translation to obtain the second sample alignment information.

[0146] Optionally, input the first sample alignment information and the second sample alignment information into the initial evaluation model to obtain the predicted evaluation result of the sample speech recognition text output by the initial evaluation model.

[0147] Optionally, based on the predicted evaluation result of the sample speech recognition text and the evaluation result label of the sample speech recognition text, calculate the loss function value, and based on the loss function value, iteratively optimize the parameters of the initial evaluation model to obtain the evaluation model.

[0148] Next, the speech recognition evaluation device provided by the embodiments of the present invention will be described. The speech recognition evaluation device described below can be correspondingly referred to the speech recognition evaluation method described above.

[0149] Figure 7 It is a schematic structural diagram of the speech recognition evaluation device provided by the embodiments of the present invention. As Figure 7 shown, the speech recognition evaluation device 700 includes:

[0150] An acquisition unit 710, configured to acquire the speech recognition text of the target speech and determine the speech recognition text label of the target speech;

[0151] A translation unit 720, configured to translate the speech recognition text and the speech recognition text label to obtain the first translation of the speech recognition text and the second translation of the speech recognition text label;

[0152] An alignment unit 730, configured to align the speech recognition text with the first translation to obtain first alignment information, and align the speech recognition text tags with the second translation to obtain second alignment information;

[0153] An evaluation unit 740, configured to input the first alignment information and the second alignment information into a pre-trained evaluation model to obtain an evaluation result of the speech recognition text output by the evaluation model;

[0154] Wherein, the evaluation model is trained based on the first sample alignment information of the sample speech recognition text and the corresponding first sample translation, the second sample alignment information of the sample speech recognition text tags and the corresponding second sample translation, and the evaluation result tags of the sample speech recognition text.

[0155] Optionally, inputting the first alignment information and the second alignment information into a pre-trained evaluation model to obtain an evaluation result of the speech recognition text output by the evaluation model includes:

[0156] Based on the evaluation model, extracting features from the first alignment information to obtain the semantic features of the first translation, extracting features from the second alignment information to obtain the semantic features of the second translation, calculating the semantic similarity between the first translation and the second translation based on the semantic features of the first translation and the semantic features of the second translation, and obtaining the intelligibility evaluation result of the speech recognition text based on the semantic similarity.

[0157] Optionally, calculating the semantic similarity between the first translation and the second translation based on the semantic features of the first translation and the semantic features of the second translation, and obtaining the intelligibility evaluation result of the speech recognition text based on the semantic similarity includes:

[0158] Aligning each word of each sentence of the first translation with each word of each sentence of the second translation to obtain a plurality of word pairs;

[0159] Based on the semantic features of the first translation and the semantic features of the second translation, determining the semantic features of the two words in each word pair;

[0160] Calculating the semantic similarity between the two words in each word pair based on the semantic features of the two words in each word pair;

[0161] Determining the weight of each word pair, and calculating the intelligibility of the speech recognition text based on the semantic similarity between the two words in each word pair and the weight of each word pair.

[0162] Optionally, aligning the speech recognition text with the first translation to obtain the first alignment information includes:

[0163] Input the speech recognition text and the first translation of the speech recognition text into a pre-trained first alignment model to obtain the first alignment information output by the first alignment model;

[0164] Among them, the first alignment model is trained based on the sample speech recognition text and the first sample translation of the sample speech recognition text, as well as the first alignment information labels of the sample speech recognition text and the first sample translation.

[0165] Optionally, align the speech recognition text label and the second translation to obtain the second alignment information, including:

[0166] Input the speech recognition text label and the second translation of the speech recognition text label into a pre-trained second alignment model to obtain the second alignment information output by the second alignment model;

[0167] Among them, the second alignment model is trained based on the sample speech recognition text label and the second sample translation of the sample speech recognition text label, as well as the second alignment information labels of the sample speech recognition text label and the second sample translation.

[0168] Optionally, translate the speech recognition text and the speech recognition text label to obtain the first translation of the speech recognition text and the second translation of the speech recognition text label, including:

[0169] Use a pre-trained machine translation model to translate the speech recognition text and the speech recognition text label to obtain the first translation of the speech recognition text and the second translation of the speech recognition text label;

[0170] Among them, the machine translation model is trained based on the sample speech recognition text and the sample speech recognition text label, as well as the first translation labels of the sample speech recognition text and the second translation labels of the sample speech recognition text label.

[0171] Optionally, the training process of the evaluation model includes:

[0172] Obtain the sample speech recognition text of the sample speech, determine the sample speech recognition text label of the sample speech, and determine the evaluation result label of the sample speech recognition text;

[0173] Translate the sample speech recognition text and the sample speech recognition text label to obtain the first sample translation of the sample speech recognition text and the second sample translation of the sample speech recognition text label;

[0174] Align the sample speech recognition text and the first sample translation to obtain the first sample alignment information, and align the sample speech recognition text label and the second sample translation to obtain the second sample alignment information;

[0175] Using the first sample alignment information and the second sample alignment information as training samples, and using the evaluation result label of the sample speech recognition text as the sample label, train the initial evaluation model. After the training is completed, an evaluation model is obtained.

[0176] It should be noted here that the speech recognition evaluation device provided in the embodiments of the present invention can implement all the method steps implemented in the above-mentioned embodiments of the speech recognition evaluation method, and can achieve the same technical effects. Therefore, the same parts and beneficial effects as those in the method embodiments will not be specifically described in this embodiment.

[0177] Figure 8 It is a schematic structural diagram of an electronic device provided in an embodiment of the present invention. As Figure 8 shown, the electronic device may include: a processor 810, a communication interface 820, a memory 830, and a communication bus 840. Among them, the processor 810, the communication interface 820, and the memory 830 communicate with each other through the communication bus 840. The processor 810 can call the logical instructions in the memory 830 to execute the speech recognition evaluation method, which includes: obtaining the speech recognition text of the target speech and determining the speech recognition text label of the target speech; translating the speech recognition text and the speech recognition text label to obtain the first translation of the speech recognition text and the second translation of the speech recognition text label; aligning the speech recognition text with the first translation to obtain the first alignment information, and aligning the speech recognition text label with the second translation to obtain the second alignment information; inputting the first alignment information and the second alignment information into a pre-trained evaluation model to obtain the evaluation result of the speech recognition text output by the evaluation model; wherein, the evaluation model is trained based on the first sample alignment information of the sample speech recognition text and the corresponding first sample translation, the second sample alignment information of the sample speech recognition text label and the corresponding second sample translation, and the evaluation result label of the sample speech recognition text.

[0178] In addition, when the logical instructions in the above-mentioned memory 830 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.

[0179] In another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is configured to execute the speech recognition evaluation method provided by the above-mentioned various methods. The method includes: obtaining the speech recognition text of the target speech and determining the speech recognition text label of the target speech; translating the speech recognition text and the speech recognition text label to obtain the first translation of the speech recognition text and the second translation of the speech recognition text label; aligning the speech recognition text with the first translation to obtain the first alignment information, and aligning the speech recognition text label with the second translation to obtain the second alignment information; inputting the first alignment information and the second alignment information into a pre-trained evaluation model to obtain the evaluation result of the speech recognition text output by the evaluation model; wherein, the evaluation model is trained based on the first sample alignment information between the sample speech recognition text and the corresponding first sample translation, the second sample alignment information between the sample speech recognition text label and the corresponding second sample translation, and the evaluation result label of the sample speech recognition text.

[0180] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative efforts.

[0181] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0182] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features. And these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A speech recognition evaluation method, characterized in that: include: Acquire the speech recognition text of the target speech, and determine the speech recognition text label of the target speech; Translating the speech recognition text and the speech recognition text label to obtain a first translation of the speech recognition text and a second translation of the speech recognition text label; Aligning the speech recognition text with the first translation to obtain first alignment information, and aligning the speech recognition text label with the second translation to obtain second alignment information; Inputting the first alignment information and the second alignment information into a pre-trained evaluation model to obtain an evaluation result of the speech recognition text output by the evaluation model; The evaluation model is trained based on first sample alignment information between a sample speech recognition text and a corresponding first sample translation, second sample alignment information between a sample speech recognition text label and a corresponding second sample translation, and an evaluation result label of the sample speech recognition text; The step of aligning the speech recognition text with the first translation to obtain first alignment information includes: Inputting the speech recognition text and the first translation of the speech recognition text into a pre-trained first alignment model to obtain the first alignment information output by the first alignment model; The step of aligning the speech recognition text label with the second translation to obtain second alignment information includes: The speech recognition text label and the second translation of the speech recognition text label are input into a pre-trained second alignment model to obtain the second alignment information output by the second alignment model.

2. The speech recognition evaluation method according to claim 1, characterized in that: The step of inputting the first alignment information and the second alignment information into a pre-trained evaluation model to obtain an evaluation result of the speech recognition text output by the evaluation model includes: Based on the evaluation model, feature extraction is performed on the first alignment information to obtain semantic features of the first translation, feature extraction is performed on the second alignment information to obtain semantic features of the second translation, semantic similarity between the first translation and the second translation is calculated based on the semantic features of the first translation and the semantic features of the second translation, and an intelligibility evaluation result of the speech recognition text is obtained based on the semantic similarity.

3. The speech recognition evaluation method according to claim 2, characterized in that: The calculating the semantic similarity between the first translation and the second translation based on the semantic features of the first translation and the semantic features of the second translation, and obtaining the intelligibility evaluation result of the speech recognition text based on the semantic similarity, includes: Aligning each word of each sentence of the first translation with each word of each sentence of the second translation to obtain a plurality of word pairs; Determining semantic features of two words in each word pair based on the semantic features of the first translation and the semantic features of the second translation; Calculating the semantic similarity between the two words in each word pair based on the semantic features of the two words in each word pair; The weight of each vocabulary pair is determined, and the intelligibility of the speech recognition text is calculated based on the semantic similarity of the two vocabulary words in each vocabulary pair and the weight of each vocabulary pair.

4. The speech recognition evaluation method according to claim 1, characterized in that: The first alignment model is trained based on a sample speech recognition text and a first sample translation of the sample speech recognition text, and a first alignment information label of the sample speech recognition text and the first sample translation.

5. The speech recognition evaluation method according to claim 1, characterized in that: The second alignment model is obtained by training based on the sample speech recognition text label and the second sample translation of the sample speech recognition text label, and the second alignment information label of the sample speech recognition text label and the second sample translation.

6. The speech recognition evaluation method according to any one of claims 2 to 5, characterized in that: The translating the speech recognition text and the speech recognition text label to obtain a first translation of the speech recognition text and a second translation of the speech recognition text label includes: Using a pre-trained machine translation model, the speech recognition text and the speech recognition text label are translated to obtain a first translation of the speech recognition text and a second translation of the speech recognition text label; The machine translation model is trained based on the sample speech recognition text and the sample speech recognition text label, as well as the first translation label of the sample speech recognition text and the second translation label of the sample speech recognition text label.

7. The speech recognition evaluation method according to claim 1, characterized in that: The training process of the evaluation model includes: Acquire a sample speech recognition text of the sample speech, determine a sample speech recognition text label of the sample speech, and determine an evaluation result label of the sample speech recognition text; Translating the sample speech recognition text and the sample speech recognition text label to obtain a first sample translation of the sample speech recognition text and a second sample translation of the sample speech recognition text label; Align the sample speech recognition text with the first sample translation to obtain first sample alignment information, and align the sample speech recognition text label with the second sample translation to obtain second sample alignment information; The first sample alignment information and the second sample alignment information are used as training samples, and the evaluation result labels of the sample speech recognition texts are used as sample labels to train an initial evaluation model. After the training is completed, the evaluation model is obtained.

8. A speech recognition evaluation device, characterized in that: include: An acquisition unit, used to acquire a speech recognition text of a target speech and determine a speech recognition text label of the target speech; A translation unit, configured to translate the speech recognition text and the speech recognition text label to obtain a first translation of the speech recognition text and a second translation of the speech recognition text label; an alignment unit, configured to align the speech recognition text with the first translation to obtain first alignment information, and to align the speech recognition text label with the second translation to obtain second alignment information; An evaluation unit, configured to input the first alignment information and the second alignment information into a pre-trained evaluation model to obtain an evaluation result of the speech recognition text output by the evaluation model; The evaluation model is trained based on first sample alignment information between a sample speech recognition text and a corresponding first sample translation, second sample alignment information between a sample speech recognition text label and a corresponding second sample translation, and an evaluation result label of the sample speech recognition text; The step of aligning the speech recognition text with the first translation to obtain first alignment information includes: Inputting the speech recognition text and the first translation of the speech recognition text into a pre-trained first alignment model to obtain the first alignment information output by the first alignment model; The step of aligning the speech recognition text label with the second translation to obtain second alignment information includes: The speech recognition text label and the second translation of the speech recognition text label are input into a pre-trained second alignment model to obtain the second alignment information output by the second alignment model.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, the speech recognition evaluation method according to any one of claims 1 to 7 is implemented.

10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the speech recognition evaluation method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Text semantic matching method and apparatus

    CN108132931A

  • Voice recognition evaluating method and device

    CN109493852A

  • Speech recognition result detection method and device, and storage medium

    CN114846543A