Pronunciation evaluation method and apparatus, medium, and electronic device

By combining a deep learning scoring model with pre-defined pronunciation evaluation rules, the problem of insensitivity to pronunciation error recognition in existing technologies is solved, and more accurate pronunciation evaluation is achieved.

CN114792528BActive Publication Date: 2026-01-16BEIJING YOUZHUJU NETWORK TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210488534.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-06
Publication Date
2026-01-16
Estimated Expiration
2042-05-06

AI Technical Summary

Technical Problem

Existing pronunciation assessment methods are not sensitive to a small number of errors in pronunciation and cannot identify individual typical pronunciation errors, resulting in inaccurate assessments even when pronunciation is mostly good.

Method used

A pre-trained deep learning scoring model is used to initially score user audio, and words with insufficient confidence are evaluated a second time by judging confidence. Pre-set pronunciation evaluation rules are used to make more accurate pronunciation evaluations.

Benefits of technology

This improves the overall accuracy of pronunciation evaluation, ensuring that the scoring of user audio is more accurate and reliable.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114792528B_ABST
    Figure CN114792528B_ABST
Patent Text Reader

Abstract

The present disclosure relates to a pronunciation evaluation method, device, medium and electronic equipment, comprising: obtaining user audio and follow-up reading text corresponding to the user audio; obtaining first word scores corresponding to each word in the user audio and confidence of the first word scores through a pre-trained deep learning scoring model according to the user audio and the follow-up reading text; if the confidence of the first word score corresponding to the word is less than a preset threshold, determining a second word score of the word as a target word score according to a preset pronunciation evaluation rule, the target word score representing the pronunciation accuracy of the word. In this way, not only can a round of scoring be performed through the deep learning scoring model, but also the pronunciation of the word with insufficient confidence in the round of scoring can be evaluated more accurately through the preset pronunciation evaluation rule again, so that the scoring evaluation of the user audio can be more accurate, and the overall accuracy of the pronunciation evaluation of the user audio is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to the field of audio technology, in particular, to a pronunciation evaluation method and device, medium and electronic equipment. BACKGROUND

[0002] It is a common scenario in the pronunciation evaluation field to evaluate the pronunciation of audio read by a user and then provide pronunciation evaluation scores accurate to the word level. Currently, common pronunciation evaluation is usually performed using GOP scores, for example, the time stamp and likelihood value of each phoneme corresponding to the read text are first obtained, thereby calculating the GOP score of each phoneme, and finally the word GOP score is obtained as the final score by means of weighted average. However, such a scheme is not sensitive to a small number of errors in pronunciation, cannot identify individual typical pronunciation errors in most good pronunciation cases, and cannot evaluate all read content, and so on, which needs to be solved urgently. SUMMARY

[0003] This summary is provided to introduce a selection of concepts, which are further described below in the detailed description section. This summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter.

[0004] In a first aspect, the present disclosure provides a pronunciation evaluation method, comprising: obtaining user audio and read text corresponding to the user audio; obtaining, according to the user audio and the read text, a first word score corresponding to each word in the user audio and a confidence of the first word score by using a pre-trained deep learning scoring model; if the confidence of the first word score corresponding to the word is less than a preset threshold, determining a second word score of the word as a target word score according to a preset pronunciation evaluation rule, the target word score representing the pronunciation accuracy of the word.

[0005] In a second aspect, the present disclosure provides a pronunciation evaluation device, comprising: an obtaining module configured to obtain user audio and read text corresponding to the user audio; a first scoring module configured to obtain, according to the user audio and the read text, a first word score corresponding to each word in the user audio and a confidence of the first word score by using a pre-trained deep learning scoring model; and a second scoring module configured to, if the confidence of the first word score corresponding to the word is less than a preset threshold, determine a second word score of the word as a target word score according to a preset pronunciation evaluation rule, the target word score representing the pronunciation accuracy of the word.

[0006] In a third aspect, the present disclosure provides a computer readable medium having stored thereon a computer program which, when executed by a processing device, implements the steps of the method of the first aspect.

[0007] In a fourth aspect, the present disclosure provides an electronic device comprising: a storage device having stored thereon a computer program; and a processing device configured to execute the computer program stored in the storage device to implement the steps of the method of the first aspect.

[0008] By the above technical solution, when the pronunciation of the user audio is evaluated, not only can the first word score obtained by one round of scoring by the pre-trained deep learning scoring model be evaluated, but also the confidence of the first word score obtained by one round of scoring can be judged, and the words with insufficient confidence in one round of scoring can be more accurately evaluated by the pre-set pronunciation evaluation rule again, so that the score evaluation of the user audio can be more accurate, and the overall accuracy of the pronunciation evaluation of the user audio is improved.

[0009] Other features and advantages of the present disclosure will be described in detail in the following detailed description. BRIEF DESCRIPTION OF DRAWINGS

[0010] The above and other features, advantages, and aspects of embodiments of the present disclosure will become more apparent by describing in detail exemplary embodiments thereof with reference to the attached drawings in which:

[0011] Figure 1 is a flowchart of a pronunciation evaluation method according to an exemplary embodiment of the present disclosure.

[0012] Figure 2 is a flowchart of a pronunciation evaluation method according to another exemplary embodiment of the present disclosure.

[0013] Figure 3 is a flowchart of a pronunciation evaluation method according to another exemplary embodiment of the present disclosure.

[0014] Figure 4 is a flowchart of a pronunciation evaluation method according to another exemplary embodiment of the present disclosure.

[0015] Figure 5 is a flowchart of a pronunciation evaluation method according to another exemplary embodiment of the present disclosure.

[0016] Figure 6 is a structural block diagram of a pronunciation evaluation device according to an exemplary embodiment of the present disclosure.

[0017] Figure 7 A structural diagram of an electronic device suitable for implementing embodiments of the present disclosure is shown. DETAILED DESCRIPTION

[0018] Embodiments of the present disclosure will be described in more detail with reference to the drawings. While certain embodiments of the present disclosure are shown in the drawings, it is understood that the present disclosure can be embodied in various forms and should not be interpreted as being limited to the embodiments set forth herein. Rather, these embodiments are provided so that the present disclosure will be thorough and complete, and fully convey the scope of the present disclosure to those skilled in the art.

[0019] It should be understood that each step recited in the method embodiments of the present disclosure can be executed in different orders and / or in parallel. In addition, the method embodiments can include additional steps and / or omit the execution of the steps shown. The scope of the present disclosure is not limited in this respect.

[0020] The term "comprising" and variations thereof as used herein are used inclusively, i.e., "comprising, but not limited to". The term "based on" is "based, at least in part, on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Related definitions will be given in the description below.

[0021] It should be noted that the terms "first", "second", and the like in the present disclosure are merely used to distinguish different devices, modules or units, and do not imply the order or interdependence of the functions performed by these devices, modules or units.

[0022] It should be noted that the terms "one", "multiple" in the present disclosure are illustrative and not restrictive, and those skilled in the art should understand that "one or more" should be understood unless otherwise explicitly indicated in the context.

[0023] The names of the messages or information exchanged between the devices in the embodiments of the present disclosure are merely for illustrative purposes, and are not intended to limit the scope of the messages or information.

[0024] It can be understood that before using the technical solutions disclosed in the embodiments of the present disclosure, the type of personal information involved in the present disclosure, the scope of use, the use scenario, etc. should be informed to the user and the authorization of the user should be obtained in accordance with relevant laws and regulations.

[0025] For example, in response to receiving an active request of a user, a prompt information is sent to the user to explicitly prompt the user that the operation requested to be performed by the user will require obtaining and using personal information of the user. Thus, the user can autonomously select whether to provide personal information to the software or hardware such as an electronic device, an application program, a server or a storage medium performing the operation of the technical solution of the present disclosure according to the prompt information.

[0026] As an optional but non-limiting implementation manner, in response to receiving an active request of a user, the manner of sending a prompt information to the user may, for example, be a pop-up window manner, in which the prompt information can be presented in a text manner. In addition, the pop-up window can also carry a selection control for the user to select “agree” or “disagree” to provide personal information to the electronic device.

[0027] It can be understood that the above notification and obtaining user authorization process is only illustrative and does not limit the implementation manner of the present disclosure, and other manners meeting relevant laws and regulations can also be applied to the implementation manner of the present disclosure.

[0028] Meanwhile, it can be understood that the data (including but not limited to the data itself, the obtaining or use of the data) involved in the technical solution should comply with the requirements of relevant laws and regulations and relevant provisions.

[0029] Figure 1 is a flowchart of a pronunciation evaluation method according to an example embodiment of the present disclosure. As shown in Figure 1 , the method comprises steps 101 to 103.

[0030] In step 101, a user audio and a follow-up reading text corresponding to the user audio are obtained. The obtaining approach of the user audio and the follow-up reading text can be any application or scenario, and the present disclosure does not limit the specific obtaining approach thereof, as long as the user audio and the follow-up reading text are needed to perform user pronunciation evaluation in a scenario.

[0031] In step 102, according to the user audio and the follow-up reading text, a first word score corresponding to each word in the user audio and a confidence of the first word score are obtained by a pre-trained deep learning scoring model.

[0032] The pre-trained deep learning scoring model may, for example, be a deep learning model based on long short-term memory (LSTM). In the training process, an artificial pronunciation evaluation result can be used as supervision information, so that the first word score output by the deep learning scoring model can be closer to the artificial evaluation result.

[0033] In a possible implementation, when the user inputs speech according to the read-aloud text, it is possible that the user repeatedly reads the same word in the read-aloud text for multiple times, and thus one word in the read-aloud text can correspond to multiple word pronunciations in the user audio. When scoring the words by using the deep learning scoring model, in addition to outputting the first word score of each word in the read-aloud text, the deep learning scoring model can also output the first word score and the confidence of the first word score of each of the multiple repeated word pronunciations in the user audio.

[0034] In step 103, if the confidence of the first word score of the word is less than a preset threshold, a second word score of the word is determined as a target word score according to a preset pronunciation evaluation rule, where the target word score represents the pronunciation accuracy of the word.

[0035] After obtaining the first word score and the confidence of the first word score of each word in the user audio output by the deep learning scoring model, the confidence of the first word score of each word can be compared with the preset threshold. If the confidence is less than the preset threshold, it indicates that the pronunciation evaluation represented by the first word score of the word is relatively inaccurate. Therefore, for the word whose confidence is less than the preset threshold, the second scoring is performed again by using the preset pronunciation evaluation rule, so as to ensure that the pronunciation evaluation of the word is more accurate and the overall accuracy of the pronunciation evaluation of the user audio is improved.

[0036] The preset pronunciation evaluation rule can be obtained according to the scoring rules of teachers, language experts and other professionals. The specific scoring method will be described below. In this embodiment, the specific content and use method of the preset pronunciation evaluation rule are not limited, as long as the words with low confidence of the first word score can be re-evaluated more accurately.

[0037] According to the above technical solution, when the pronunciation of the user audio is evaluated, not only the deep learning scoring model pre-trained can be used for one round of scoring, but also the confidence of the first word score obtained by the one round of scoring can be judged, and the words with insufficient confidence in the one round of scoring can be re-evaluated more accurately by using the preset pronunciation evaluation rule, so as to ensure that the scoring evaluation of the user audio is more accurate and the overall accuracy of the pronunciation evaluation of the user audio is improved.

[0038] The following Figure 2 is a flowchart illustrating a detailed method of determining the second word score of the word according to the preset pronunciation evaluation rule. As shown in Figure 2 , Figure 2is a flow chart of a pronunciation evaluation method according to an example embodiment of the present disclosure, the method further comprising steps 201 and 202.

[0039] In step 201, if the confidence of the first word score corresponding to the word is less than a preset threshold, the phoneme pronunciation type of each phoneme in the word is determined according to the user audio and the follow-up text by a preset phoneme pronunciation type judgment rule. The phoneme pronunciation type is used to represent the acceptable degree of the pronunciation of each phoneme in the word.

[0040] In step 202, the second word score of the word is determined according to the phoneme pronunciation type of each phoneme in the word, and the second word score is taken as the target word score.

[0041] The phoneme pronunciation type can include: standard, non-standard, tolerable error, intolerable error. The following gives a judgment example of determining the phoneme pronunciation type according to the preset phoneme pronunciation type judgment rule.

[0042] First, for the phoneme pronunciation type "standard", the judgment condition that can be judged as standard can be that a certain phoneme is read correctly, which is very standard in terms of American English, such as (1) phoneme is read as corresponding American English variant, such as word tail darkL is read as similar to OW, word initial T is read as tap-flap (e.g. water, turtle); (2) word non-stress syllable vowel is weakened to AH, such as the vowel of words like to and is is weakened to AH, which is standard; (3) "unclear nasal sound" in the word (non-initial and non-final), even without separate M / N / NG; (4) word tail plosive, fricative, affricate voiced sound is read as clear and little aspirated; (5) correct linking.

[0043] Second, for the phoneme pronunciation type "non-standard", the judgment condition that can be judged as non-standard can be that as long as it is not a definite error, it can be judged as non-standard, such as (1) a certain phoneme is read not enough correctly, which is like correct pronunciation A, but also like B similar to A, and it is difficult to distinguish (e.g. kite sounds like kite and tight); (2) unclear addition and subtraction sound (e.g. the phoneme lH of beautiful is read very lightly, almost not; the phoneme T of kite is followed by AH, as if added, as if not added); (3) the plosive (P, T, K) after S is aspirated (e.g. the T of stop is aspirated); (4) American phoneme is read as corresponding British phoneme, such as American R / ER is read without tongue curling (e.g. are is read as AA, far is read as F AA, ER of first is not tongue curled), and such as American bath is read as British BAA TH.

[0044] Again, for the phoneme pronunciation type "tolerable error", the determination condition that can be determined as a tolerable error can be that a certain phoneme is obviously read as another phoneme that is easily confused with it, or there is an increase or decrease in the phoneme, but it can be tolerated due to perceptual reasons, such as (1) the beginning of a word is truncated by 1 consonant by audio; (2) after the end of a word consonant, in the middle of a consonant cluster, there is an obvious increase in AH / UW and Chinese vowels (e.g. kite is read as ki, guess is read as ge, H is read as EY, and relax is read as rela k s); (3) the voiceless and voiced consonants (B-P, D-T, G-K) in the middle of a word (not at the beginning or end of a word) are mixed.

[0045] Finally, for the phoneme pronunciation type "intolerable error", the determination condition that can be determined as an intolerable error can be all phoneme pronunciation errors other than the above-mentioned tolerable errors, such as (1) all other increases in phonemes are intolerable errors except for the increase in AH / UW and Chinese vowels at the end of a syllable; (2) all decreases in phonemes are intolerable errors.

[0046] After determining the phoneme pronunciation type of each phoneme in the word, determining the second word score of the word according to the phoneme pronunciation type of each phoneme in the word can also be achieved according to a preset scoring rule. For example, in the scoring rule, different phoneme pronunciation types can correspond to different phoneme scores, and then the second word score of the word can be obtained by combining the phoneme scores corresponding to the phoneme pronunciation types of all phonemes in the word.

[0047] Figure 3 is a flowchart of a pronunciation evaluation method according to an example embodiment of the present disclosure. As shown in Figure 3 The method further includes steps 301 to 305.

[0048] In step 301, acoustic features are extracted from the user audio by a pre-trained acoustic model, and the acoustic features are frame-level likelihood values.

[0049] In step 302, a weighted finite state transducer is constructed by the follow-up text.

[0050] In step 303, the acoustic features and the weighted finite state transducer are decoded and aligned to obtain phoneme time stamps corresponding to each phoneme in the word.

[0051] The weighted finite state transducer is a sentence cycle-based weighted finite state transducer (WFST), so as to ensure that each reading in the user audio is an independent word through the alignment when the same word in the reading text is read twice or more times in the user audio, and each phoneme in the corresponding word of each reading is obtained.

[0052] The steps 301 to 303 can be performed again in the case that the confidence of the first word score corresponding to the word in the user audio is less than the preset threshold, or can be directly performed after the user audio and the reading text are obtained in the step 101 as shown in the embodiment. Figure 3

[0053] In the step 304, the GOP score, the phoneme type and the position of each phoneme in the word corresponding to each phoneme in the word are determined according to the phoneme timestamp and the acoustic feature.

[0054] In the step 305, the phoneme pronunciation type of each phoneme in the word is determined according to the GOP score, the phoneme type and the position of each phoneme in the word corresponding to each phoneme in the word through the preset phoneme pronunciation type determination rule.

[0055] When it is necessary to determine the phoneme pronunciation type of each phoneme in the word according to the user audio and the reading text through the preset phoneme pronunciation type determination rule, the GOP score, the phoneme type and the position of each phoneme in the word corresponding to each phoneme in the word can be determined according to the user audio and the reading text, and then the phoneme pronunciation type of each phoneme in the word is determined according to the GOP score, the phoneme type and the position of each phoneme in the word corresponding to each phoneme in the word through the preset phoneme pronunciation type determination rule.

[0056] Specifically, in the case that the steps 301 to 303 are performed, the method for determining the GOP score, the phoneme type and the position of each phoneme in the word corresponding to each phoneme in the word can be to determine according to the acoustic feature constituted by the frame-level likelihood value extracted in the step 301 and the phoneme timestamp corresponding to each phoneme in the word obtained through the alignment in the step 303.

[0057] ​The corresponding GOP score of each phoneme can be calculated by a preset GOP calculation formula and the corresponding acoustic features of each phoneme. The phoneme type corresponding to each phoneme in the word can be any conventional phoneme type mentioned in the foregoing description of how to determine the phoneme pronunciation type, such as stop, vowel, fricative, nasal, and voiced. The position of each phoneme in the word can be the beginning, middle, or end of the word.

[0058] If the steps 301 to 303 are performed again in the case where the confidence of the first word score corresponding to a word in the user audio is less than the preset threshold, the step of obtaining the first word score corresponding to each word in the user audio and the confidence of the first word score from the user audio and the follow-up text by using the pre-trained deep learning scoring model can be implemented in other ways. If the steps 301 to 303 are performed directly after obtaining the user audio and the follow-up text in the step 101 shown in Figures 1-3 , the implementation method of obtaining the first word score corresponding to each word in the user audio and the confidence of the first word score from the user audio and the follow-up text by using the pre-trained deep learning scoring model can be as shown in Figure 4 .

[0059] Figure 4 is a flowchart of a pronunciation evaluation method according to an example embodiment of the present disclosure. As shown in Figure 4 , the method further includes steps 401 and 402.

[0060] In step 401, the phoneme likelihood value vector of the phonemes included in each word in the user audio and the phoneme ID of the phonemes included in each word in the user audio are determined from the phoneme timestamp and the acoustic features. The phoneme ID is used to determine the actual phonemes included in each word in the user audio. The phoneme ID can be pre-assigned. For example, there are 44 phonemes included in English, and the 44 English phonemes can be numbered by using the phoneme ID in advance, so that the phoneme ID can be used to represent the identity of the phoneme.

[0061] In step 402, the phoneme ID and the phoneme likelihood value vector are used as inputs of the deep learning scoring model to obtain the first word score corresponding to each word in the user audio and the confidence of the first word score.

[0062] Figure 5 is a flowchart of a pronunciation evaluation method according to an example embodiment of the present disclosure. As shown in Figure 5 , the method further includes steps 501 and 502.

[0063] In step 501, it is judged whether the confidence of the first word score corresponding to the word is less than a preset threshold. If yes, go to step 304; if no, go to step 502.

[0064] In step 502, that is, in the case where the confidence of the first word score corresponding to the word is not less than the preset threshold, the first word score is taken as the target word score.

[0065] After obtaining the confidence of the first word score corresponding to each word in the user audio through the deep learning scoring model, if the first word score corresponding to all words in the user audio is not less than the preset threshold, it indicates that the accuracy of obtaining the first word score corresponding to each word in the user audio through the deep learning scoring model is high enough, and the first word score can be directly taken as the target word score to realize pronunciation evaluation of the user audio.

[0066] Figure 6 is a structural block diagram of a pronunciation evaluation device according to an exemplary embodiment of the present disclosure. As shown in FIG. 6, the device comprises: an acquisition module 10, configured to acquire user audio and a follow-up reading text corresponding to the user audio; a first scoring module 20, configured to obtain a first word score corresponding to each word in the user audio and a confidence of the first word score through a pre-trained deep learning scoring model according to the user audio and the follow-up reading text; and a second scoring module 30, configured to, if the confidence of the first word score corresponding to the word is less than a preset threshold, determine a second word score of the word as a target word score according to a preset pronunciation evaluation rule, the target word score representing an accuracy degree of pronunciation of the word.

[0067] Through the above technical solution, when performing pronunciation evaluation on the user audio, not only one round of scoring can be performed through the pre-trained deep learning scoring model, but also the confidence of the first word score obtained in the one round of scoring can be judged, and the words with insufficient confidence in the one round of scoring can be further evaluated more accurately through the preset pronunciation evaluation rule, so that the scoring evaluation of the user audio can be more accurate, and the overall accuracy of the pronunciation evaluation of the user audio is improved.

[0068] In a possible implementation, the second scoring module 30 is further configured to: determine a phoneme pronunciation type of each phoneme in the word according to the user audio and the follow-up reading text through a preset phoneme pronunciation type judgment rule, the phoneme pronunciation type being used to represent an acceptable degree of pronunciation of each phoneme in the word; determine the second word score of the word according to the phoneme pronunciation type of each phoneme in the word, and take the second word score as the target word score.

[0069] In a possible implementation, the phoneme pronunciation type includes: standard, non-standard, tolerable error, and intolerable error.

[0070] In a possible implementation, the second scoring module 30 is further configured to: determine, according to the user audio and the read-along text, a GOP score corresponding to each phoneme in the word, a phoneme type corresponding to each phoneme in the word, and a position of each phoneme in the word; and determine, according to the GOP score corresponding to each phoneme in the word, the phoneme type corresponding to each phoneme in the word, and the position of each phoneme in the word, the phoneme pronunciation type of each phoneme in the word by using the preset phoneme pronunciation type judgment rule.

[0071] In a possible implementation, the second scoring module 30 is further configured to: extract, by using a pre-trained acoustic model, acoustic features from the user audio, the acoustic features being frame-level likelihood values; construct a weighted finite state transducer by using the read-along text; perform decoding alignment on the acoustic features and the weighted finite state transducer to obtain a phoneme timestamp corresponding to each phoneme in the word; and determine, according to the phoneme timestamp and the acoustic features, the GOP score corresponding to each phoneme in the word, the phoneme type corresponding to each phoneme in the word, and the position of each phoneme in the word.

[0072] In a possible implementation, the first scoring module 20 is further configured to: determine, according to the phoneme timestamp and the acoustic features, a phoneme likelihood value vector of a phoneme included in each word in the user audio, and a phoneme ID of the phoneme included in each word in the user audio, the phoneme ID being used to determine an actual phoneme included in each word in the user audio; and input the phoneme ID and the phoneme likelihood value vector into the deep learning scoring model as input, to obtain a first word score corresponding to each word in the user audio and a confidence of the first word score.

[0073] In a possible implementation, the first scoring module 20 is further configured to: if the confidence of the first word score corresponding to the word is not less than a preset threshold, take the first word score as the target word score.

[0074] Reference will be made to the following description of embodiments of the disclosure in conjunction with the accompanying drawings. Figure 7The diagram illustrates a structural schematic of an electronic device 700 suitable for implementing embodiments of the present disclosure. Terminal devices in embodiments of the present disclosure may include, but are not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 7 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.

[0075] like Figure 7 As shown, the electronic device 700 may include a processing unit (e.g., a central processing unit, a graphics processor, etc.) 701, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 702 or a program loaded from a storage device 708 into a random access memory (RAM) 703. The RAM 703 also stores various programs and data required for the operation of the electronic device 700. The processing unit 701, ROM 702, and RAM 703 are interconnected via a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.

[0076] Typically, the following devices can be connected to I / O interface 705: input devices 706 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 707 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 708 including, for example, magnetic tapes, hard disks, etc.; and communication devices 709. Communication device 709 allows electronic device 700 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 7 An electronic device 700 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.

[0077] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication device 709, or installed from storage device 708, or installed from ROM 702. When the computer program is executed by processing device 701, it performs the functions defined in the methods of embodiments of this disclosure.

[0078] It is noted that the aforementioned computer-readable medium of the present disclosure can be a computer-readable signal medium or a computer-readable storage medium or any combination thereof. The computer-readable storage medium can be, for example and without limitation, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the computer-readable storage medium can include, but are not limited to, an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing. In the present disclosure, the computer-readable storage medium can be any tangible medium that contains or stores a program used by or in connection with an instruction execution system, apparatus, or device. In the present disclosure, the computer-readable signal medium can include a computer-readable program code transmitted by a computer-readable medium or a carrier wave in a baseband or as part of a carrier wave. Such a propagated computer-readable signal medium can take many forms, including but not limited to, an electromagnetic signal, an optical signal, or any suitable combination of the foregoing. The computer-readable signal medium can also be any computer-readable medium that is not a computer-readable storage medium and that can be used to carry or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained in the computer-readable medium can be transmitted by any suitable medium, including but not limited to, wire, cable, RF (radio frequency), or the like, or any suitable combination of the foregoing.

[0079] In some embodiments, the client, server, or both can communicate using any current known or future developed network protocol, such as HTTP (HyperText Transfer Protocol), and can be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include local area networks ("LANs"), wide area networks ("WANs"), the Internet, and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any current known or future developed networks.

[0080] The aforementioned computer-readable medium can be included within the aforementioned electronic device; or can exist separately from the electronic device without being incorporated therein.

[0081] The computer readable medium described above carries one or more programs, when the one or more programs are executed by the electronic device, cause the electronic device to: acquire user audio and follow-up text corresponding to the user audio; acquire, according to the user audio and the follow-up text, a first word score corresponding to each word in the user audio and a confidence degree of the first word score by using a pre-trained deep learning scoring model; and if the confidence degree of the first word score corresponding to the word is less than a preset threshold, determine a second word score of the word as a target word score according to a preset pronunciation evaluation rule, the target word score representing a pronunciation accuracy degree of the word.

[0082] Computer program code for carrying out operations of the present disclosure can be written in any of one or more programming languages or combinations of languages including object or visual programming languages such as Java, Smalltalk, C++ or conventional procedural programming languages such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider).

[0083] The flow diagrams and the block diagrams in the drawings are illustrations of possible architectures, functions, and operations for systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flow diagrams or block diagrams can represent a module, a segment, or a portion of code, which comprises one or more executable instructions for implementing the specified logical functions. It should also be noted that in some alternative implementations, the functions noted in the blocks can occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently or the blocks may

[0084] The modules described in the embodiments of the present disclosure can be implemented in the form of software, or can be implemented in the form of hardware. In some cases, the name of a module does not constitute a limitation on the module itself. For example, a obtaining module can also be described as a module that obtains user audio and read-along text corresponding to the user audio.

[0085] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, and without limitation, exemplary types of hardware logic components that can be used include Field-programmable Gate Arrays (FPGAs), Application-specific Integrated Circuits (ASICs), Application-specific Standard Products (ASSPs), System-on-a-chip systems (SOCs), Complex Programmable Logic Devices (CPLDs), etc.

[0086] In the context of the present disclosure, a machine-readable medium can be a tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. Machine-readable storage media can include, without limitation, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media can include one or more lines of electrical connections, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or Flash memory), optical storage devices, portable compact disc read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0087] According to one or more embodiments of the present disclosure, example 1 provides a pronunciation evaluation method, the method comprising: obtaining user audio and read-along text corresponding to the user audio; obtaining, according to the user audio and the read-along text, a first word score corresponding to each word in the user audio and a confidence of the first word score by using a pre-trained deep learning scoring model; if the confidence of the first word score corresponding to the word is less than a preset threshold, determining a second word score of the word as a target word score according to a preset pronunciation evaluation rule, the target word score representing a pronunciation accuracy degree of the word.

[0088] According to one or more embodiments of the present disclosure, example 2 provides the method of example 1, wherein the determining the second word score of the word as the target word score according to the preset pronunciation evaluation rule comprises: determining, according to the user audio and the follow-up text, a phoneme pronunciation type of each phoneme in the word by a preset phoneme pronunciation type judgment rule, the phoneme pronunciation type being used to represent an acceptable degree of pronunciation of each phoneme in the word; and determining the second word score of the word according to the phoneme pronunciation type of each phoneme in the word, and taking the second word score as the target word score.

[0089] According to one or more embodiments of the present disclosure, example 3 provides the method of example 2, wherein the phoneme pronunciation type comprises: standard, non-standard, tolerable error, and intolerable error.

[0090] According to one or more embodiments of the present disclosure, example 4 provides the method of example 2, wherein the determining, according to the user audio and the follow-up text, a phoneme pronunciation type of each phoneme in the word by a preset phoneme pronunciation type judgment rule comprises: determining, according to the user audio and the follow-up text, a GOP score corresponding to each phoneme in the word, a phoneme type corresponding to each phoneme in the word, and a position of each phoneme in the word; and determining, according to the GOP score corresponding to each phoneme in the word, the phoneme type corresponding to each phoneme in the word, and the position of each phoneme in the word, the phoneme pronunciation type of each phoneme in the word by the preset phoneme pronunciation type judgment rule.

[0091] According to one or more embodiments of the present disclosure, example 5 provides the method of example 4, wherein the determining, according to the user audio and the follow-up text, a GOP score corresponding to each phoneme in the word, a phoneme type corresponding to each phoneme in the word, and a position of each phoneme in the word comprises: extracting, by an acoustic model trained in advance, acoustic features from the user audio, the acoustic features being frame-level likelihood values; constructing a weighted finite state transducer by the follow-up text; performing decoding alignment on the acoustic features and the weighted finite state transducer to obtain a phoneme timestamp corresponding to each phoneme in the word; and determining, according to the phoneme timestamp and the acoustic features, the GOP score corresponding to each phoneme in the word, the phoneme type corresponding to each phoneme in the word, and the position of each phoneme in the word.

[0092] According to one or more embodiments of the present disclosure, example 6 provides the method of example 5, wherein the obtaining, according to the user audio and the follow-up text, the first word score corresponding to each word in the user audio and the confidence of the first word score by a pre-trained deep learning scoring model comprises: determining, according to the phoneme timestamp and the acoustic feature, a phoneme likelihood value vector of a phoneme included in each word in the user audio and a phoneme ID of the phoneme included in each word in the user audio, the phoneme ID being used to determine an actual phoneme included in each word in the user audio; and taking the phoneme ID and the phoneme likelihood value vector as an input of the deep learning scoring model to obtain the first word score corresponding to each word in the user audio and the confidence of the first word score.

[0093] According to one or more embodiments of the present disclosure, example 7 provides the method of example 1, wherein the method further comprises: if the confidence of the first word score corresponding to the word is not less than a preset threshold, taking the first word score as the target word score.

[0094] According to one or more embodiments of the present disclosure, example 8 provides a pronunciation evaluation device, comprising: an obtaining module configured to obtain user audio and follow-up text corresponding to the user audio; a first scoring module configured to obtain, according to the user audio and the follow-up text, a first word score corresponding to each word in the user audio and a confidence of the first word score by a pre-trained deep learning scoring model; and a second scoring module configured to, if the confidence of the first word score corresponding to the word is less than a preset threshold, determine a second word score of the word according to a preset pronunciation evaluation rule as a target word score, the target word score representing a pronunciation accuracy of the word.

[0095] According to one or more embodiments of the present disclosure, example 9 provides a computer readable medium having a computer program stored thereon, the program being executed by a processing device to implement the steps of the method of any one of examples 1-7.

[0096] According to one or more embodiments of the present disclosure, example 10 provides an electronic device, comprising: a storage device having a computer program stored thereon; and a processing device configured to execute the computer program in the storage device to implement the steps of the method of any one of examples 1-7.

[0097] The above description merely illustrates the preferred embodiment of the disclosure and a principle of applied technologies. It should be understood by those skilled in the art that the disclosed range of the disclosure is not limited to the technical solutions formed by the specific combinations of the technical features described above, and should also cover other technical solutions formed by the combinations of the technical features described above or their equivalent features without departing from the disclosed concept. For example, the technical solutions formed by the mutual replacement of the above-described features and the technical features with similar functions disclosed in the disclosure (but not limited to) can be formed.

[0098] Furthermore, although each operation is depicted in a particular order, this should not be understood as requiring these operations to be performed in the particular order shown or in a sequential order. Under certain circumstances, multitasking and parallel processing can be advantageous. Likewise, although specific implementation details are included in the above discussion, these should not be interpreted as limiting the scope of the disclosure. Certain features described in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment can also be implemented separately or in any suitable subcombination. It will be appreciated that various features described herein can form one or more embodiments of the disclosure. The disclosure covers all possible combinations of the features described herein.

[0099] Although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely illustrative of the example forms of implementing the claims. As to the means for performing the operations of the apparatus in the above-described embodiments, the specific manner in which the operations are performed has been described in detail in the embodiments related to the method, and will not be described here in detail.

Claims

1. A method of evaluating a pronunciation, characterized by, The method comprises: obtaining user audio and a read-along text corresponding to the user audio; obtaining, according to the user audio and the read-along text, a first word score corresponding to each word in the user audio and a confidence of the first word score by using a pre-trained deep learning scoring model; if the confidence of the first word score corresponding to the word is less than a preset threshold, determining a second word score of the word as a target word score according to a preset pronunciation evaluation rule, the target word score representing a pronunciation accuracy of the word; the determining of the second word score of the word as the target word score according to the preset pronunciation evaluation rule comprises: determining, according to the user audio and the read-along text, a phoneme pronunciation type of each phoneme in the word by using a preset phoneme pronunciation type judgment rule, the phoneme pronunciation type being used to represent an acceptable degree of pronunciation of each phoneme in the word; determining the second word score of the word according to the phoneme pronunciation type of each phoneme in the word, and taking the second word score as the target word score; the determining of the phoneme pronunciation type of each phoneme in the word according to the user audio and the read-along text by using the preset phoneme pronunciation type judgment rule comprises: determining, according to the user audio and the read-along text, a GOP score corresponding to each phoneme in the word, a phoneme type corresponding to each phoneme in the word, and a position of each phoneme in the word in the word; determining, according to the GOP score corresponding to each phoneme in the word, the phoneme type corresponding to each phoneme in the word, and the position of each phoneme in the word in the word, the phoneme pronunciation type of each phoneme in the word by using the preset phoneme pronunciation type judgment rule.

2. The method of claim 1, wherein, The phoneme pronunciation type comprises: standard, non-standard, tolerable error, and intolerable error.

3. The method of claim 1, wherein, The determining of the GOP score corresponding to each phoneme in the word, the phoneme type corresponding to each phoneme in the word, and the position of each phoneme in the word in the word according to the user audio and the read-along text comprises: extracting, from the user audio, acoustic features by using a pre-trained acoustic model, the acoustic features being frame-level likelihood values; constructing a weighted finite state transducer by using the read-along text; performing decoding alignment on the acoustic features and the weighted finite state transducer to obtain a phoneme timestamp corresponding to each phoneme in the word; determining, according to the phoneme timestamp and the acoustic features, the GOP score corresponding to each phoneme in the word, the phoneme type corresponding to each phoneme in the word, and the position of each phoneme in the word in the word.

4. The method of claim 3, wherein, The obtaining of the first word score corresponding to each word in the user audio and the confidence of the first word score by using the pre-trained deep learning scoring model according to the user audio and the read-along text comprises: determine, according to the phoneme time stamp and the acoustic feature, a phoneme likelihood value vector of a phoneme included in each word in the user audio and a phoneme ID of the phoneme included in each word in the user audio, the phoneme ID being used to determine an actual phoneme included in each word in the user audio; input the phoneme ID and the phoneme likelihood value vector as an input of the deep learning scoring model to obtain a first word score corresponding to each word in the user audio and a confidence of the first word score.

5. The method of claim 1, wherein, The method further includes: if the confidence of the first word score corresponding to the word is not less than a preset threshold, taking the first word score as the target word score.

6. A voice evaluation apparatus characterized by comprising: The apparatus includes: an obtaining module configured to obtain user audio and read-aloud text corresponding to the user audio; a first scoring module configured to obtain, according to the user audio and the read-aloud text, a first word score corresponding to each word in the user audio and a confidence of the first word score by using a pre-trained deep learning scoring model; a second scoring module configured to, if the confidence of the first word score corresponding to the word is less than a preset threshold, determine a second word score of the word according to a preset pronunciation evaluation rule, the second word score being taken as a target word score and representing a pronunciation accuracy degree of the word. The second scoring module is further configured to: determine, according to the user audio and the read-aloud text, a phoneme pronunciation type of each phoneme in the word by using a preset phoneme pronunciation type determination rule, the phoneme pronunciation type being used to represent an acceptable degree of pronunciation of each phoneme in the word; and determine the second word score of the word according to the phoneme pronunciation type of each phoneme in the word, and take the second word score as the target word score. The second scoring module is further configured to: determine, according to the user audio and the read-aloud text, a GOP score corresponding to each phoneme in the word, a phoneme type corresponding to each phoneme in the word, and a position of each phoneme in the word; and determine, according to the GOP score corresponding to each phoneme in the word, the phoneme type corresponding to each phoneme in the word, and the position of each phoneme in the word, the phoneme pronunciation type of each phoneme in the word by using the preset phoneme pronunciation type determination rule.

7. A computer readable medium having stored thereon a computer program, characterized in that The program is executed by a processing apparatus to implement steps of the method of any one of claims 1-5.

8. An electronic device, comprising: comprise: a storage device having a computer program stored thereon; a processing apparatus configured to execute the computer program in the storage device to implement steps of the method of any one of claims 1-5.

Citation Information

Patent Citations

  • Spoken language pronunciation evaluation method based on deep neural network posterior probability algorithm

    CN108364634A

  • Voice evaluation method and device, storage medium and electronic device

    CN110782921A