A method, device and readable storage medium for sentence pronunciation evaluation

By building a decoding network containing multiple read/miss read/miss read paths and using high-frequency words as misread paths, the problem of misjudgment of high-score words in the prior art is solved, and a more accurate and experienced sentence pronunciation evaluation is achieved.

CN114678013BActive Publication Date: 2025-05-16SUZHOU QIMENGZHE NETWORK TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202210278116.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-21
Publication Date
2025-05-16
Estimated Expiration
2042-03-21

AI Technical Summary

Technical Problem

Existing sentence pronunciation evaluation systems are prone to misjudgment when evaluating high-score words, resulting in poor subjective experience for language learners.

Method used

By building a decoding network containing multiple read/miss read/miss read paths, and using high-frequency words as misread paths, combining the language model cost scores of the target word, a weighted inter-word decoding network is generated to reduce the scoring errors of high-score words.

Benefits of technology

While taking into account the ratings of multiple reads/misreads/misreads, the rating errors of high-scoring words are reduced, improving the accuracy of the evaluation and the learner's experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114678013B_ABST
    Figure CN114678013B_ABST
Patent Text Reader

Abstract

The present invention provides a method, device and readable storage medium for sentence pronunciation evaluation. The method comprises: constructing a weighted inter-word decoding network based on a target word sequence and a high-frequency word set; performing speech recognition on the audio to be evaluated to obtain a set of candidate decoding paths; traversing all possible word sequences corresponding to the current candidate decoding path set to obtain a set of new candidate word sequences with the minimum edit distance to the target text, and further selecting the path with the highest decoding score from the candidate decoding paths corresponding to the candidate word sequences as the identification optimal path output. The present invention can minimize the scoring errors of high-scoring words while taking into account the scoring of over-read / missed / misread words.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of speech recognition, and in particular to a method, device and readable storage medium for sentence pronunciation evaluation. Background Art

[0002] Oral pronunciation assessment technology is being accepted by more and more language learners because of its ability to stably and quickly evaluate pronunciation accuracy. In the use of sentence pronunciation assessment systems, misjudgment of high-scoring words as low scores often leads to poor subjective experience for language learners.

[0003] In existing solutions, the target word-to-word decoding network containing multiple read / missed read paths is usually used to decode the audio to be tested to obtain the optimal recognition path, and the word sequence represented by the optimal path and its corresponding likelihood are used for subsequent pronunciation evaluation. The decoding path score includes two parts: the acoustic score and the language model score. Since the language model cost score of a word sequence is often the accumulation of the connection probability between words, when the acoustic scores are similar, the recognition process will tend to be short word sequences. This will cause high-scoring words to have low scores in sentence evaluation. Summary of the invention

[0004] The object of the present invention is to provide a method, device and readable storage medium for sentence pronunciation evaluation that can minimize the number of high-scoring word scoring errors while taking into account the scoring of over-read / omitted / misread words.

[0005] A brief summary of one or more aspects is given below to provide a basic understanding of these aspects. This summary is not an exhaustive overview of all conceived aspects, and is neither intended to identify the key or critical elements of all aspects nor to define the scope of any or all aspects. Its only purpose is to give some concepts of one or more aspects in a simplified form as a prelude to a more detailed description that will be given later.

[0006] According to one aspect of the present invention, a method for sentence pronunciation evaluation is provided, comprising:

[0007] Step 100, constructing a decoding network based on the target word set and the high-frequency word set;

[0008] Step 200, performing speech recognition on the audio to be tested to obtain a set of candidate decoding paths;

[0009] Select the word sequence with the minimum edit distance with the text to be evaluated from the candidate path set as the candidate word sequence set;

[0010] Select the path with the highest decoding score from the candidate decoding paths corresponding to the candidate word sequence as the identification optimal path output;

[0011] Step 300, obtaining the pronunciation score of each word in the recognition word sequence according to the GOP formula;

[0012] Step 400, parse the recognition word sequence and the target word sequence to obtain the pronunciation score of each word in the target word sequence.

[0013] In one embodiment, in step 100, the target word related path in the decoding network includes the multiple reading / missed reading path of the target text, and the decoding network uses the high-frequency word part as the misreading path of the target text.

[0014] In one embodiment, in step 100, the language model cost score of the target word is set to a constant value that is the same for each word.

[0015] In one embodiment, in step 100, high-frequency words are obtained from training data statistics, and the language model cost score corresponding to the high-frequency words is the probability of the word appearing in the training data multiplied by a weight coefficient less than 1, and the language model cost score of the high-frequency word is greater than the language model cost score of the target word.

[0016] In one embodiment, the step 100 includes:

[0017] Step 101, compile and generate a top-level state network before the evaluation begins;

[0018] Step 102, constructing a sub-state network using the target text during evaluation;

[0019] Step 103, nesting the sub-state network in the top-level state network to obtain the final state decoding network.

[0020] In one embodiment, step 101 includes: adding a target word symbol according to high-frequency words and their corresponding language model cost scores, constructing a word-level decoding network for inter-word jumps, and then constructing a state decoding network in combination with a conventional pronunciation dictionary and state binding information of phonemes as a top-level state network.

[0021] In one embodiment, step 102 includes: constructing a state decoding network according to the target word and its corresponding language model cost score in combination with state binding information of a specific pronunciation dictionary and phonemes.

[0022] In one embodiment, in step 200, each candidate decoding path includes a state sequence with the same length as the time frame, the likelihood / jump probability of each state, the correspondence between the word sequence and the state sequence, and the acoustic score / language model score cost of the word.

[0023] In one embodiment, in step 200, the word sequence having the minimum edit distance with the text to be evaluated is selected from the candidate path set as the candidate word sequence set, including: extracting the word sequence in the current candidate path set, and searching for the word sequence set having the minimum edit distance with the target word sequence in the word sequence after removing duplicates as the candidate word sequence set.

[0024] In one embodiment, the selection of the minimum edit distance word path is performed based on a weighted finite state transition machine.

[0025] In one embodiment, the method for selecting the minimum edit distance word path specifically includes:

[0026] Construct a finite state receiver for the current candidate word sequence set and the target word sequence respectively;

[0027] Establish a finite state receiver corresponding to the edit cost of the target word and the candidate word. There is an arc between any target word and any candidate word. When the candidate word and the target word are the same, the weight of the arc is set to 0, and when the candidate word and the target word are different, the weight of the arc is set to 1; the weight of the arc corresponding to the empty input and the candidate word output is set to 1, and the weight of the arc between the target word and the empty output is set to 0;

[0028] The FSA corresponding to the target word sequence and the WFST corresponding to the edit cost function are compounded, and then the FSA corresponding to the candidate word sequence set and the new WFST are compounded to obtain the edit distance corresponding to each candidate word path, and the path / path set with the minimum cost is selected for output.

[0029] In one embodiment, step 400 includes: performing text alignment on the recognition word sequence and the target word sequence, setting the scores of the words corresponding to "deletion" and "replacement" errors in the target word sequence to the lowest score, and keeping the scores of the remaining words unchanged.

[0030] According to a second aspect of the present invention, there is provided a sentence pronunciation evaluation device, comprising:

[0031] The decoding network building module is configured to input a target word set and a high-frequency word set, generate an inter-word decoding network, and then combine the pronunciation dictionary and the HMM model to output a state-level decoding network;

[0032] A decoding module is configured to use a state-level decoding network to identify the audio to be tested and output a set of candidate decoding paths;

[0033] The optimal recognition path selection module is configured to select a word path with the minimum edit distance with the target word sequence from the candidate decoding path set as the candidate word path set, and select the path with the highest decoding score from the candidate decoding paths corresponding to the candidate word path set for output;

[0034] A word GOP scoring module is configured to input a word time boundary and a word likelihood corresponding to an optimal recognition path, and output a GOP score for each word in the recognition word path;

[0035] The recognition word sequence parsing module is configured to input the GOP score of each word in the recognition word path, rewrite the score according to the alignment result of the recognition word sequence and the target word sequence, and output the final word pronunciation score.

[0036] According to a third aspect of the present invention, there is provided a sentence pronunciation evaluation device, comprising a memory and a processor;

[0037] The memory is used to store computer programs;

[0038] The processor is used to implement the sentence pronunciation evaluation method as described in the first aspect when executing the computer program.

[0039] According to a fourth aspect of the present invention, there is provided a readable storage medium having a program stored thereon, and when the program is executed by a processor, the sentence pronunciation evaluation method as described in the first aspect is implemented.

[0040] The beneficial effects of the embodiments of the present invention are: 1. A weighted inter-word decoding network is constructed using high-frequency words and target words. This decoding network covers the multiple reading / missed reading / wrong reading phenomena that may occur in sentence evaluation, and gives the target words a smaller language model cost score and gives the high-frequency words used as misreading absorption a larger language model cost score. While taking into account the misreading path, it also ensures that high-scoring words are not misjudged as much as possible.

[0041] ② After the decoding search process is completed, the word path set that meets the "minimum edit distance with the target text" is selected from the optimal and suboptimal candidate path sets, which further increases the possibility of the target word path being selected and further reduces the scoring errors of high-scoring words. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for use in the embodiments are briefly introduced below. It should be understood that the following drawings only show certain embodiments of the present invention and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without creative work.

[0043] The above features and advantages of the present invention can be better understood after reading the detailed description of the embodiments of the present disclosure in conjunction with the following drawings. In the drawings, the components are not necessarily drawn to scale, and components with similar related properties or features may have the same or similar reference numerals.

[0044] Figure 1 This is a flow chart of the sentence pronunciation evaluation method according to an embodiment of the present application;

[0045] Figure 2 This is a schematic diagram of building a decoding network based on a target word sequence and a set of high-frequency words;

[0046] Figure 3 It is a schematic diagram of the word-level decoding space corresponding to the target text;

[0047] Figure 4 This is a flow chart of establishing a state decoding network for a target word sequence in an embodiment of the present application;

[0048] Figure 5 It is a schematic diagram of the set of candidate paths;

[0049] Figure 6 is an example of a sentence pronunciation score display;

[0050] Figure 7 It is a module schematic diagram of the device of the embodiment of the present application. DETAILED DESCRIPTION

[0051] The present invention is described in detail below in conjunction with the accompanying drawings and specific embodiments. Note that the aspects described below in conjunction with the accompanying drawings and specific embodiments are only exemplary and should not be construed as limiting the scope of protection of the present invention in any way.

[0052] In a first aspect of the embodiments of the present invention, a method for sentence pronunciation evaluation is provided, comprising:

[0053] Step S100, constructing a weighted inter-word decoding network based on a target word set and a high-frequency word set.

[0054] After obtaining the target text (i.e. the text to be evaluated), a word-level network is generated using the target word set and the high-frequency word set (refer to Figure 2 ), the target word related path in this network includes the multiple reading / missing reading path of the target text, and the high-frequency word part is used as the misreading path of the target text. Using high-frequency words as the misreading path of the target text meets the requirement of "if the target word is misread as other words, it will be judged as the lowest score" in the pronunciation scoring rules, and is easy to implement.

[0055] The language model cost score of the target word is set to a constant value for each word. The high-frequency words are obtained from the training data. The language model cost score corresponding to the high-frequency words is the probability of the word appearing in the training data multiplied by a weight coefficient less than 1, lm_cost(w top )=-scale*log(P(w top)), and the language model cost score of the high-frequency word needs to be greater than the language model cost score of the target word. The purpose is to ensure that the path containing the misread word is selected as the optimal decoding path only when the acoustic score of the misread word is high enough.

[0056] HMM-based speech recognition refers to the process of finding the state sequence that is most likely to correspond to a given audio in a given state search space, also known as decoding. The word sequence corresponding to the state sequence is the decoded text. This process can be expressed as the following formula:

[0057]

[0058] Where O represents the observed audio, W represents the word sequence, S w Represents the state sequence set corresponding to the word sequence W, t m Represents word w m The end time of the corresponding state Among them, P(O, S|W) corresponds to the acoustic score, P(W) is the language model score (the negative corresponds to the language model cost), and the product of the two is the decoding score (after taking the logarithm, it is the sum of the two).

[0059] If the target word sequence in the sentence is read twice / missed / misread, and there is no corresponding path in the search space, it will affect the time boundary of the correct word and the corresponding acoustic score, and then affect the word score that depends on this acoustic information.

[0060] Figure 3 This is a schematic diagram of the decoding space corresponding to the target text "nice to meet you". The arcs between the nodes represent the output words and the corresponding language model cost scores. The thick circle represents the end node, where: (0, 1, nice / alpha)-(1, 2, to / alpha)-(2, 3, meet / alpha)-(3, 4, you / alpha) represents the target word sequence nice-to-meet-you, and the corresponding language model cost score is 4alpha.

[0061] (0, 1, nice / alpha)-(1, 2, to / alpha)-(3, 4, you / alpha) represents the missed word sequence nice-to-you, and the corresponding language model cost score is 3alpha.

[0062] 0, 1, nice / alpha)-(1, 2, nice / alpha)-(2, 3, to / alpha)-(3, 4, meet / alpha)-(4, 5, you / alpha) represents the multi-word sequence nice-nice-to-meet-to-you, and the corresponding language model cost score is 5alpha.

[0063] (0, 1, top_w1 / beta1)-(1, 2, top_w2 / beta2)-(2, 3, meet / alpha)-(3, 4, you / alpha) represents the misread word sequence top_w1-top_w2-meet-you, and the corresponding language model cost score is beta1+beta2+2*alpha.

[0064] In some embodiments, in order to reduce the time consumption of decoding network compilation, a top-level state network can be compiled and generated before the evaluation begins. During the evaluation, the target text is used to construct a sub-state network, and the sub-state network is nested in the top-level state network to obtain the final state decoding network. See the flowchart. Figure 4 . The compilation of the top-level state network in step S101 is as follows: according to the aforementioned high-frequency words and their corresponding language model cost scores, a target word symbol is added to construct a word-level decoding network for inter-word jumps, and then the state decoding network is constructed in combination with the pronunciation dictionary and the state binding information of the phonemes as the top-level state network. The compilation process of the sub-state network in step S102 is to construct a state decoding network according to the aforementioned high-frequency words and their corresponding language model cost scores, combined with the specific pronunciation dictionary and the state binding information of the phonemes.

[0065] Step S200, perform speech recognition on the audio to be evaluated to obtain a set of candidate decoding paths. Traverse all possible word sequences corresponding to the current candidate decoding path set to obtain a new set of candidate word sequences with the minimum edit distance to the text to be evaluated, and select the path with the highest decoding score from the candidate decoding paths corresponding to the candidate word sequences as the output of the optimal recognition path.

[0066] 2.1) Each candidate path contains a state sequence with the same length as the time frame and the likelihood / jump probability of each state, the correspondence between the word sequence and the state sequence, and the acoustic score / language model score cost of the word. At this time, a word sequence may correspond to multiple state sequences. Figure 5 This is a schematic diagram of a candidate path set. There are three corresponding paths: A1-B1-C1-D1; A2-B2-C2-D2; A3-B3-C3-D3-E3; each node contains 6 pieces of information, namely, word / corresponding state sequence / start time / end time / acoustic score / language model score cost.

[0067] 2.2) Extract the word sequence in the current candidate path set, and find the word sequence set with the minimum edit distance with the target word sequence in the word sequence after removing duplicates as the candidate word sequence set. The minimum edit distance refers to the minimum number of operations required to transform the symbol sequence S1 into the symbol sequence S2.

[0068] The selection of the minimum edit distance word path can be based on a weighted finite state transducer (WFST). A finite state acceptor (FSA-finite state acceptor) is constructed for the current candidate word sequence set and the target word sequence. An FSA corresponding to the edit cost of the target word and the candidate word is established. There is an arc between any target word and any candidate word. When the candidate word and the target word are the same, the weight of the arc is 0, and when the candidate word and the target word are different, the weight of the arc is 1. In addition, the weight of the arc corresponding to the empty input and the candidate word output is 1, and the weight of the arc between the target word and the empty output is 0. The FSA corresponding to the target word sequence and the WFST corresponding to the edit cost function are compounded, and then the FSA corresponding to the candidate word sequence set and the new WFST are used to compound the edit distance corresponding to each candidate word path, and the path / path set with the minimum cost is selected for output.

[0069] 2.3) Select the path set corresponding to the candidate word sequence set in the original candidate paths, and select the path with the highest decoding score as the identification path output.

[0070] Assume that the audio content is "How do you do", the target text is "how do you do", and the audio has connected reading at doyou. The recognition text corresponding to the path with the highest decoding score is how-do-do. If the pronunciation evaluation is performed based on this path information, the word you will result in a score error due to a deletion error. Although the deletion errors in the recognition process can be solved by setting an insertion penalty during decoding, it will also lead to an increase in insertion errors. The words corresponding to the insertion errors will occupy the time boundary of the correct words, resulting in misjudgment of the correct pronunciation score. This patent does not directly use the state sequence with the highest decoding score, but selects the path with the smallest edit distance from the target word path from the optimal and suboptimal candidate state sequences, which maximizes the possibility of the target word path being selected, that is, reduces the possibility of high-scoring words being misjudged.

[0071] Step S300, obtaining the pronunciation score of each word in the recognized word sequence according to the GOP (Goodness of Pronunciation) formula, which can be referred to in the paper Witt SM, FSJ Y. Phone-level pronunciation scoring and assessment for interactive language learning [J]. Speech Communication, 2000, 30 (2 / 3): 95-108.

[0072] Step S400: parse the recognition word sequence and the target word sequence to obtain the final pronunciation score of each word in the target word sequence.

[0073] Text alignment is performed on the recognition word sequence and the target word sequence. The scores of the words corresponding to "deletion" and "replacement" errors in the target word sequence are set to the lowest score, and the scores of the remaining words remain unchanged. Figure 6 Here is an example of a sentence pronunciation score display.

[0074] like Figure 7 As shown, the present invention also provides a sentence pronunciation evaluation device, the sentence pronunciation evaluation device comprising:

[0075] The decoding network construction module 701 is configured to input a target word set and a high-frequency word set, generate an inter-word decoding network, and then combine the pronunciation dictionary and the HMM model to output a state-level decoding network;

[0076] A decoding module 702 is configured to use a state-level decoding network to identify the audio to be tested and output a set of candidate decoding paths;

[0077] The optimal recognition path selection module 703 is configured to select a word path with the minimum edit distance with the target word sequence from the candidate decoding path set as the candidate word path set, and select the path with the highest decoding score from the candidate decoding paths corresponding to the candidate word path set for output;

[0078] The word GOP scoring module 704 is configured to input the word time boundary and word likelihood corresponding to the optimal recognition path, and output the GOP score of each word in the recognition word path;

[0079] The recognition word sequence parsing module 705 is configured to input the GOP score of each word in the recognition word path, rewrite the score according to the alignment result of the recognition word sequence and the target word sequence, and output the final word pronunciation score.

[0080] It is easy to understand that the embodiment of the present application also provides a sentence pronunciation evaluation device, including a memory and a processor; wherein the memory can be used to store instructions, programs, codes, code sets or instruction sets. The memory can include a program storage area and a data storage area, wherein the program storage area can store instructions for implementing an operating system, instructions for at least one function, and instructions for implementing the above-mentioned sentence pronunciation evaluation method, etc.; the data storage area can store data involved in the above-mentioned sentence pronunciation evaluation method, etc.

[0081] The processor may include one or more processing cores. The processor calls the data stored in the memory by running or executing the instructions, programs, code sets or instruction sets stored in the memory, performs various functions of the present application and processes data. The processor may be at least one of a special purpose integrated circuit, a digital signal processor, a digital signal processing device, a programmable logic device, a field programmable gate array, a central processing unit, a controller, a microcontroller and a microprocessor. It is understandable that for different devices, the electronic device used to implement the above-mentioned processor function can also be other, and the embodiments of the present application are not specifically limited.

[0082] If the above method of the embodiment of the present application is implemented in the form of a software function module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiment of the present invention can be essentially or partly reflected in the form of a software product that contributes to the prior art. The computer software product is stored in a storage medium, including a number of instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the methods described in each embodiment of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as a U disk, a mobile hard disk, a read-only memory (ROM, Read Only Memory), a magnetic disk or an optical disk. In this way, the embodiment of the present invention is not limited to any specific combination of hardware and software.

[0083] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other.

[0084] The previous description of the disclosure is provided to enable any person skilled in the art to make or use the disclosure. Various modifications to the disclosure will be apparent to those skilled in the art, and the general principles defined herein may be applied to other variations without departing from the spirit or scope of the disclosure. Thus, the disclosure is not intended to be limited to the examples and designs described herein, but should be granted the widest scope consistent with the principles and novel features disclosed herein.

[0085] The above description is only a preferred example of the present application and is not intended to limit the present application. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present application shall be included in the scope of protection of the present application.

Claims

1. A method for sentence pronunciation assessment, characterized in that: include: Step 100, constructing a decoding network based on a target word set and a high-frequency word set, wherein the target word-related path in the decoding network includes a multiple-reading / missed-reading path of the target text, and the decoding network uses the high-frequency word part as the misreading path of the target text, and the language model cost score of the target word is set to a constant value that is the same for each word, and the high-frequency words are obtained by statistics of the training data, and the language model cost score corresponding to the high-frequency word is the probability of the word appearing in the training data multiplied by a weight coefficient less than 1, and the language model cost score of the high-frequency word is greater than the language model cost score of the target word; Step 200, performing speech recognition on the audio to be tested to obtain a set of candidate decoding paths; Select the word sequence with the minimum edit distance with the text to be evaluated from the candidate path set as the candidate word sequence set; Select the path with the highest decoding score from the candidate decoding paths corresponding to the candidate word sequence as the identification optimal path output; Step 300, obtaining the pronunciation score of each word in the recognition word sequence according to the GOP formula; Step 400, parse the recognition word sequence and the target word sequence to obtain the pronunciation score of each word in the target word sequence.

2. The method for sentence pronunciation evaluation according to claim 1, characterized in that: The step 100 comprises: Step 101, compile and generate a top-level state network before the evaluation begins; Step 102, constructing a sub-state network using the target text during evaluation; Step 103, nesting the sub-state network in the top-level state network to obtain the final state decoding network.

3. The method for sentence pronunciation evaluation according to claim 2, characterized in that: The step 101 includes: adding a target word symbol according to the high-frequency words and their corresponding language model cost scores, constructing a word-level decoding network for inter-word jumps, and then constructing a state decoding network as a top-level state network in combination with a conventional pronunciation dictionary and state binding information of phonemes.

4. The method for sentence pronunciation evaluation according to claim 3, characterized in that: The step 102 includes: constructing a state decoding network according to the target word and its corresponding language model cost score, combined with the state binding information of the specific pronunciation dictionary and the phoneme.

5. The method for sentence pronunciation evaluation according to claim 1, characterized in that: In step 200, each candidate decoding path includes a state sequence with the same length as the time frame, the likelihood / jump probability of each state, the correspondence between the word sequence and the state sequence, and the acoustic score / language model score cost of the word.

6. The method for sentence pronunciation evaluation according to claim 1, characterized in that: In step 200, the word sequence having the minimum edit distance with the text to be evaluated is selected from the candidate path set as the candidate word sequence set, including: extracting the word sequence in the current candidate path set, and searching for the word sequence set having the minimum edit distance with the target word sequence in the word sequence after removing duplicates as the candidate word sequence set.

7. The method for sentence pronunciation evaluation according to claim 6, characterized in that: The selection of the minimum edit distance word path is performed based on a weighted finite state transition machine.

8. The method for sentence pronunciation evaluation according to claim 7, characterized in that: The method for selecting the minimum edit distance word path specifically includes: Construct a finite state receiver for the current candidate word sequence set and the target word sequence respectively; Establish a finite state receiver corresponding to the edit cost of the target word and the candidate word. There is an arc between any target word and any candidate word. When the candidate word and the target word are the same, the weight of the arc is set to 0, and when the candidate word and the target word are different, the weight of the arc is set to 1; the weight of the arc corresponding to the empty input and the candidate word output is set to 1, and the weight of the arc between the target word and the empty output is set to 0; The FSA corresponding to the target word sequence and the WFST corresponding to the edit cost function are compounded, and then the FSA corresponding to the candidate word sequence set and the new WFST are compounded to obtain the edit distance corresponding to each candidate word path, and the path / path set with the minimum cost is selected for output.

9. The method for sentence pronunciation evaluation according to claim 1, characterized in that: The step 400 includes: performing text alignment on the recognition word sequence and the target word sequence, setting the scores of the words corresponding to "deletion" and "replacement" errors in the target word sequence to the lowest score, and keeping the scores of the remaining words unchanged.

10. A sentence pronunciation evaluation device, characterized in that it comprises: A decoding network construction module is configured to input a target word set and a high-frequency word set, generate an inter-word decoding network, and then combine the pronunciation dictionary and the HMM model to output a state-level decoding network, wherein the target word-related path in the decoding network includes a multiple reading / missed reading path of the target text, and the decoding network uses the high-frequency word part as the misreading path of the target text. The language model cost score of the target word is set to a constant value that is the same for each word, and the high-frequency word is obtained by statistics of the training data. The language model cost score corresponding to the high-frequency word is the probability of the word appearing in the training data multiplied by a weight coefficient less than 1, and the language model cost score of the high-frequency word is greater than the language model cost score of the target word; A decoding module is configured to use a state-level decoding network to identify the audio to be tested and output a set of candidate decoding paths; The optimal recognition path selection module is configured to select a word path with the minimum edit distance with the target word sequence from the candidate decoding path set as the candidate word path set, and select the path with the highest decoding score from the candidate decoding paths corresponding to the candidate word path set for output; A word GOP scoring module is configured to input a word time boundary and a word likelihood corresponding to an optimal recognition path, and output a GOP score for each word in the recognition word path; The recognition word sequence parsing module is configured to input the GOP score of each word in the recognition word path, rewrite the score according to the alignment result of the recognition word sequence and the target word sequence, and output the final word pronunciation score.

11. A sentence pronunciation evaluation device, characterized in that: including memory and processor; The memory is used to store computer programs; The processor is used to implement the sentence pronunciation evaluation method as described in any one of claims 1 to 9 when executing the computer program.

12. A readable storage medium, characterized in that: The storage medium stores a program, and when the program is executed by the processor, the sentence pronunciation evaluation method according to any one of claims 1 to 9 is implemented.

Citation Information

Patent Citations

  • Text processing method and device

    CN113095066A

  • Audio evaluation method and device and non-instantaneous storage medium

    CN113707178A

  • Speech recognition method and device, storage medium and electronic equipment

    CN113990293A

  • Speech recognition method, related equipment and readable storage medium

    CN114155836A