Pronunciation correction method and device, electronic equipment and storage medium

By acquiring and analyzing the text and speech to be evaluated, identifying the control and competing phonetic sequences, and utilizing the GOP scoring model and diphthong processing technology, the problem of low pronunciation correction success rate and accuracy in existing speech recognition systems is solved, thereby improving user trust and experience.

CN116682435BActive Publication Date: 2026-07-21NETEASE YOUDAO (HANGZHOU) SMART TECH CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
NETEASE YOUDAO (HANGZHOU) SMART TECH CO LTD
Filing Date
2023-06-30
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

Existing automatic speech recognition systems have a high tolerance for pronunciation errors, resulting in low success and accuracy rates for pronunciation correction. Furthermore, timestamp algorithms cannot accurately align when users have poor pronunciation, affecting users' trust in the system's reliability.

Method used

By obtaining the text to be evaluated, a sequence of reference phonetic symbols is determined, and the speech to be evaluated is processed to obtain a sequence of competing phonetic symbols. Using the GOP scoring model and diphthong processing technology, the text for pronunciation correction is determined, thereby improving the accuracy of pronunciation correction.

Benefits of technology

It improves the accuracy of pronunciation correction, enhances user trust in the system and user experience, and avoids the negative impact of speech recognition systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116682435B_ABST
    Figure CN116682435B_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide a pronunciation correction method, device, electronic equipment and storage medium. The pronunciation correction method comprises: obtaining a text to be evaluated and a speech to be evaluated; determining a reference phonetic symbol sequence according to the text to be evaluated; performing speech analysis processing on the speech to be evaluated to obtain a competitive phonetic symbol sequence; and determining a pronunciation correction text based on the reference phonetic symbol sequence and the competitive phonetic symbol sequence. The technical solution of the present application can improve the accuracy of pronunciation correction and improve the user experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of this application generally relate to the field of speech processing technology, and more specifically, the embodiments of this application relate to pronunciation correction methods, apparatus, electronic devices, and storage media. Background Technology

[0002] This section is intended to provide background or context for the embodiments of this application set forth in the claims. The description herein may include concepts that may be explored, but not necessarily concepts that have been previously conceived or explored. Therefore, unless otherwise stated, what is described in this section is not prior art for the purposes of the specification and claims of this application, and is not acknowledged as prior art simply by virtue of its inclusion in this section.

[0003] During automatic speech recognition, when a categorized error exists between the user's actual pronunciation and the standard pronunciation, the incorrect pronunciation can be corrected. When a non-categorized error exists between the user's actual pronunciation and the standard pronunciation, pronunciation prompts can be provided to the user. The existence of an error between the user's actual pronunciation and the standard pronunciation can be determined by comparing and analyzing the standard pronunciation sequence and the user's pronunciation sequence obtained from the text conversion based on automatic speech recognition.

[0004] However, because existing automatic speech recognition systems generally have a high tolerance for errors and can correctly recognize some pronunciation deviations, the success rate and accuracy of pronunciation correction may be relatively low. Furthermore, the timestamp algorithm of speech recognition systems not only fails to determine a reasonable timestamp when the user's pronunciation is poor, but also disrupts the time alignment of adjacent phonemes, causing well-pronounced phonemes to be judged as incorrect by the system. Moreover, speech recognition systems only support pronunciation guidance for CMU phonemes, and their support for the International Phonetic Alphabet (IPA) is incomplete. When there is a significant difference between CMU phonemes and IPA, the speech recognition system struggles to perfectly correct pronunciation. In summary, both missed and excessive corrections significantly reduce users' trust in the system's reliability.

[0005] Therefore, there is an urgent need to propose an innovative pronunciation correction method in order to improve the accuracy of pronunciation correction and enhance the user experience. Summary of the Invention

[0006] To overcome the problems existing in related technologies, the embodiments of this application aim to provide a pronunciation correction method, apparatus, electronic device, and storage medium. This pronunciation correction method can improve the accuracy of pronunciation correction and enhance the user experience.

[0007] In a first aspect of the embodiments of this application, a pronunciation correction method is provided, comprising: acquiring a text to be evaluated and a speech to be evaluated; determining a reference phonetic symbol sequence based on the text to be evaluated; performing speech analysis processing on the speech to be evaluated to obtain a competing phonetic symbol sequence; and determining a pronunciation correction text based on the reference phonetic symbol sequence and the competing phonetic symbol sequence.

[0008] In one embodiment of this application, the speech analysis processing of the speech to be evaluated to obtain a competitive phonetic symbol sequence includes: parsing the speech frames of the speech to be evaluated using a preset acoustic processing model to obtain multiple phonemes and phoneme probability vectors corresponding to each speech frame in the speech to be evaluated; wherein, the phoneme probability vector contains the phoneme probability corresponding to each phoneme in the speech frame; determining the corresponding reference phonetic symbol in the reference phonetic symbol sequence for each speech frame based on the multiple phonemes and phoneme probability vectors corresponding to each speech frame, so as to divide all speech frames in the speech to be evaluated into multiple speech frame groups; and determining the competitive phonetic symbol for each speech frame group based on the reference phonetic symbol sequence and the multiple phonemes and phoneme probability vectors corresponding to each speech frame, so as to form a competitive phonetic symbol sequence.

[0009] In one embodiment of this application, determining the competing phonetic symbols for each speech frame group based on the reference phonetic symbol sequence, multiple phonemes corresponding to each speech frame, and phoneme probability vectors includes: determining the phoneme score value corresponding to each phoneme in each speech frame group based on the multiple phonemes corresponding to each speech frame and phoneme probability vectors in each speech frame group; and determining the competing phonetic symbols for each speech frame group based on the phoneme score value corresponding to each phoneme in each speech frame group and the phoneme occurrence frequency of each phoneme in each speech frame group.

[0010] In one embodiment of this application, after determining the phoneme score value corresponding to each phoneme in each speech frame group based on multiple phonemes and phoneme probability vectors corresponding to each speech frame in each speech frame group, the method further includes: determining that there are diphthongs in the text to be evaluated based on the reference phonetic symbol sequence; merging the two speech frame groups corresponding to the diphthongs into a diphthong group; determining the phonetic symbol score value corresponding to each phoneme in the diphthong group based on the phoneme score value corresponding to each phoneme in the two speech frame groups corresponding to the diphthongs; and determining the competing phonetic symbols of the diphthong group based on the phonetic symbol score value corresponding to each phoneme in the diphthong group and the phonetic symbol occurrence frequency of each phoneme in the diphthong group.

[0011] In one embodiment of this application, determining the phoneme score value corresponding to each phoneme in each speech frame group based on multiple phonemes and phoneme probability vectors corresponding to each speech frame in each speech frame group includes: determining the phoneme score value corresponding to each phoneme in each speech frame group based on multiple phonemes and phoneme probability vectors corresponding to each speech frame in each speech frame group using a GOP scoring model; the GOP scoring model is as follows:

[0012]

[0013] Where GOP(p) represents the phoneme score for each phoneme in each speech frame group, T is the total number of speech frames in each speech frame group, and P(s) t | t ) represents the t-th speech frame O in each speech frame group. t The corresponding phoneme is s t The probability value.

[0014] In one embodiment of this application, before determining the pronunciation correction text based on the comparative phonetic symbol sequence and the competing phonetic symbol sequence, the method further includes: determining the word score value corresponding to each word in the text to be evaluated based on the phonetic score value corresponding to each phoneme in each speech frame group; and determining whether to perform pronunciation correction on the words in the text to be evaluated based on the word score value.

[0015] In one embodiment of this application, determining the pronunciation correction text based on the reference phonetic symbol sequence and the competing phonetic symbol sequence includes: determining the evaluation result based on the reference phonetic symbol sequence and the competing phonetic symbol sequence, wherein the evaluation result includes the competing phonetic symbol and the reference phonetic symbol corresponding to the competing phonetic symbol; and determining the pronunciation correction text based on a preset text format, the competing phonetic symbol and the reference phonetic symbol through a text output model.

[0016] In one embodiment of this application, after determining the pronunciation correction text based on the contrasting phonetic symbol sequence and the competing phonetic symbol sequence, the method further includes: converting the pronunciation correction text into an audio file using TTS speech synthesis technology; and playing the audio file.

[0017] In one embodiment of this application, determining the reference phonetic sequence based on the text to be evaluated includes: determining whether each word in the text to be evaluated exists in the International Phonetic Alphabet (IPA) dictionary; if it exists, determining the reference phonetic sequence based on the IPA of each word in the text to be evaluated in the IPA dictionary; and if it does not exist, parsing the pronunciation sequence of the text to be evaluated using a G2P dictionary tool, and performing IPA mapping based on the pronunciation sequence to obtain the reference phonetic sequence.

[0018] In a second aspect of the embodiments of this application, a pronunciation correction apparatus is provided for performing the pronunciation correction method as described in any one of the first aspects, comprising:

[0019] The data acquisition module is used to acquire the text and audio to be evaluated.

[0020] The phonetic symbol sequence determination module is used to determine the corresponding phonetic symbol sequence based on the text to be evaluated;

[0021] The analysis and processing module is used to perform speech analysis and processing on the speech to be evaluated to obtain the competing phonetic symbol sequence;

[0022] The correction text determination module is used to determine pronunciation correction text based on the reference phonetic symbol sequence and the competing phonetic symbol sequence.

[0023] A third aspect of this application provides an electronic device, comprising:

[0024] Processor; and

[0025] A memory that stores executable code, which, when executed by the processor, causes the processor to perform the method described above.

[0026] A fourth aspect of this application provides a non-transitory machine-readable storage medium having executable code stored thereon, which, when executed by a processor of an electronic device, causes the processor to perform the method described above.

[0027] The technical solution provided by the embodiments of this application has the following beneficial effects:

[0028] The pronunciation correction method, apparatus, electronic device, and storage medium provided in this application acquire the text to be evaluated and the speech to be evaluated. On the one hand, they determine a reference phonetic symbol sequence based on the text to be evaluated; on the other hand, they perform speech analysis processing on the speech to be evaluated to obtain a competing phonetic symbol sequence. This avoids the negative impact of speech recognition systems that could reduce the success rate and accuracy of pronunciation correction. Furthermore, based on the reference phonetic symbol sequence and the competing phonetic symbol sequence, they determine the pronunciation correction text and use this text to provide feedback to the user on the phonetic symbols the user mispronounced, thereby improving the accuracy of pronunciation correction and enhancing the user's trust in pronunciation correction and overall user experience. Attached Figure Description

[0029] The above and other objects, features, and advantages of exemplary embodiments of this application will become readily apparent from the following detailed description taken in conjunction with the accompanying drawings. Several embodiments of this application are illustrated in the drawings by way of example and not limitation, in which:

[0030] Figure 1A block diagram schematically illustrates an exemplary computing system 100 suitable for implementing embodiments of this application;

[0031] Figure 2 A schematic flowchart of a pronunciation correction method according to another embodiment of this application is shown;

[0032] Figure 3 A schematic flowchart of a pronunciation correction method according to yet another embodiment of this application is shown.

[0033] Figure 4 A schematic flowchart of a pronunciation correction method according to another embodiment of this application is shown;

[0034] Figure 5 A schematic diagram of the structure of a pronunciation correction device according to another embodiment of this application is shown;

[0035] Figure 6 A schematic block diagram of an electronic device according to an embodiment of this application is shown.

[0036] In the accompanying drawings, the same or corresponding reference numerals indicate the same or corresponding parts. Detailed Implementation

[0037] The principles and spirit of this application will now be described with reference to several exemplary embodiments. It should be understood that these embodiments are provided merely to enable those skilled in the art to better understand and implement this application, and are not intended to limit the scope of this application in any way. Rather, these embodiments are provided to make this disclosure more thorough and complete, and to fully convey the scope of this disclosure to those skilled in the art.

[0038] Figure 1 A block diagram of an exemplary computing system 100 suitable for implementing embodiments of this application is shown. Figure 1As shown, the computing system 100 may include: a central processing unit (CPU) 101, random access memory (RAM) 102, read-only memory (ROM) 103, a system bus 104, a hard disk controller 105, a keyboard controller 106, a serial interface controller 107, a parallel interface controller 108, a display controller 109, a hard disk 110, a keyboard 111, a serial external device 112, a parallel external device 113, and a display 114. Among these devices, the CPU 101, RAM 102, ROM 103, hard disk controller 105, keyboard controller 106, serial controller 107, parallel controller 108, and display controller 109 are coupled to the system bus 104. The hard disk 110 is coupled to the hard disk controller 105, the keyboard 111 is coupled to the keyboard controller 106, the serial external device 112 is coupled to the serial interface controller 107, the parallel external device 113 is coupled to the parallel interface controller 108, and the display 114 is coupled to the display controller 109. It should be understood that... Figure 1 The structural diagrams described are for illustrative purposes only and are not intended to limit the scope of this application. In some cases, certain devices may be added or removed depending on the specific circumstances.

[0039] Those skilled in the art will recognize that embodiments of this application can be implemented as a system, method, or computer program product. Therefore, this disclosure can be specifically implemented as entirely hardware, entirely software (including firmware, resident software, microcode, etc.), or a combination of hardware and software, generally referred to herein as a "circuit," "module," or "system." Furthermore, in some embodiments, this application can also be implemented as a computer program product contained in one or more computer-readable media, which includes computer-readable program code.

[0040] Any combination of one or more computer-readable media may be used. A computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. A computer-readable storage medium can be, for example,, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples (not exhaustive) of a computer-readable storage medium may include: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this document, a computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in connection with an instruction execution system, apparatus, or device.

[0041] Computer-readable signal media may include data signals propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. Computer-readable signal media may also be any computer-readable medium other than computer-readable storage media, capable of sending, propagating, or transmitting programs for use by or in connection with an instruction execution system, apparatus, or device.

[0042] Program code contained on a computer-readable medium may be transmitted using any suitable medium, including but not limited to wireless, wire, optical fiber, RF, etc., or any suitable combination thereof.

[0043] Computer program code for performing the operations of this application can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, and C++, and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network (including a local area network (LAN) or a wide area network (WAN)), or it can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0044] The embodiments of this application will now be described with reference to flowchart illustrations and block diagrams of apparatus (or systems) according to embodiments of this application. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine that, when executed by a computer or other programmable data processing apparatus, creates means for implementing the functions / operations specified in the blocks of the flowcharts and / or block diagrams.

[0045] These computer program instructions may also be stored in a computer-readable medium that enables a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer-readable medium produce a product comprising an instruction apparatus that implements the functions / operations specified in the boxes of a flowchart and / or block diagram.

[0046] Computer program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, such that the instructions that execute on the computer or other programmable apparatus can provide a process for implementing the functions / operations specified in the boxes of a flowchart and / or block diagram.

[0047] According to an embodiment of this application, a pronunciation correction method and apparatus are proposed.

[0048] In this article, it is important to understand that any number of elements in the accompanying figures is for illustrative purposes and not for limitation, and any naming is for distinction only and has no limiting meaning.

[0049] The principles and spirit of this application will be explained in detail below with reference to several representative embodiments. Invention Overview

[0051] The applicant found that existing automatic speech recognition systems generally have a high tolerance for errors, correctly recognizing even some pronunciation deviations. This may result in a lower success rate and accuracy of pronunciation correction. Furthermore, the timestamp algorithm in speech recognition systems not only fails to determine a reasonable timestamp when the user's pronunciation is poor, but also disrupts the time alignment of adjacent phonemes, causing well-pronounced phonemes to be misidentified by the system. In summary, both missed and excessive corrections significantly reduce users' trust in the system's reliability.

[0052] Based on this, the technical solution of this application acquires the text to be evaluated and the speech to be evaluated. On the one hand, it determines the reference phonetic symbol sequence based on the text to be evaluated; on the other hand, it performs speech analysis processing on the speech to be evaluated to obtain the competing phonetic symbol sequence. Furthermore, it determines the pronunciation correction text based on the reference phonetic symbol sequence and the competing phonetic symbol sequence. This avoids the negative impact of speech recognition systems that could reduce the success rate and accuracy of pronunciation correction, thereby improving the accuracy of pronunciation correction, increasing user trust in pronunciation correction, and enhancing the user experience.

[0053] After introducing the basic principles of this application, the various non-limiting embodiments of this application will be described in detail below.

[0054] Application Scenarios Overview

[0055] The pronunciation correction method described in this application is applicable to various educational electronic products that can correct a user's pronunciation, such as smart learning tablets and smart learning desk lamps. It can also be applied to expand functionality for electronic devices such as smart mobile devices, personal computers, and servers.

[0056] Exemplary methods

[0057] The following is for reference. Figure 2 This application describes a pronunciation correction method according to exemplary embodiments thereof. It should be noted that the above application scenarios are shown only to facilitate understanding of the spirit and principles of this application, and the embodiments of this application are not limited in any way. Rather, the embodiments of this application can be applied to any applicable scenario.

[0058] Figure 2 A schematic flowchart of a pronunciation correction method according to another embodiment of this application is shown. Please refer to [link / reference]. Figure 2 The pronunciation correction method shown in the embodiments of this application may include:

[0059] In step S201, the text to be evaluated and the audio to be evaluated are acquired. The aforementioned audio to be evaluated is an audio file input by the user through a recording device such as a microphone, awaiting evaluation and pronunciation correction. Additionally, the aforementioned text to be evaluated can be a text file displayed to the user for reading aloud; this text file can be pre-recorded by the processor or downloaded from a cloud server.

[0060] It is understood that there are various ways to obtain the text and speech to be evaluated. In practical applications, the method of obtaining the text and speech to be evaluated needs to be determined according to the actual application situation. This application does not impose any restrictions in this regard.

[0061] In step S202, a reference phonetic sequence is determined based on the text to be evaluated. It is understood that the aforementioned reference phonetic sequence can be considered as the standard pronunciation sequence of the text to be evaluated. In practical applications, the text to be evaluated can be phonetically mapped to determine the reference phonetic sequence. For example, assuming the text to be evaluated is the word "relax," the standard phonetic sequence for "relax" can be determined based on a dictionary or similar resource. Therefore, this standard phonetic symbol sequence It can then serve as the corresponding phonetic sequence for the word "relax".

[0062] In step S203, speech analysis processing is performed on the speech to be evaluated to obtain a competing phonetic symbol sequence. It can be understood that the aforementioned competing phonetic symbol sequence can be considered as the phonetic symbol sequence of the user's actual pronunciation. For example, assuming the text to be evaluated is the word "relax," the user's pronunciation of "relax" may deviate slightly, and the competing phonetic symbol sequence obtained after speech analysis processing might be...

[0063] In step S204, pronunciation correction text is determined based on the reference phonetic symbol sequence and the competing phonetic symbol sequence. In this embodiment, the reference phonetic symbol sequence and the competing phonetic symbol sequence can be compared. If the phonetic symbols at corresponding positions in the reference and competing phonetic symbol sequences are consistent, it indicates that the user has pronounced the phonetic symbol correctly; otherwise, it indicates that the user has mispronounced the phonetic symbol. Further, the phonetic symbols mispronounced by the user can be filtered out from the reference and competing phonetic symbol sequences to obtain the reference and competing phonetic symbols, thereby constructing the pronunciation correction text based on the reference and competing phonetic symbols. The aforementioned pronunciation correction text is used to display to the user, enabling the user to correct their pronunciation based on the information in the pronunciation correction text.

[0064] As an example, suppose the text to be evaluated is "relax", and the corresponding phonetic sequence is: The competing phonetic symbol sequence is Then the pronunciation correction text could be " Don't read it as / k / should not be read as / t / . It is understood that pronunciation correction texts come in various forms. In practical applications, the specific form of the pronunciation correction text needs to be determined according to the actual application situation. This application does not impose any restrictions in this regard.

[0065] By acquiring the text and speech to be evaluated, a matching phonetic symbol sequence is determined based on the text, while the speech is analyzed to obtain a competing phonetic symbol sequence. This avoids the negative impact of the speech recognition system on pronunciation correction success and accuracy. Furthermore, based on the matching and competing phonetic symbol sequences, a pronunciation correction text is determined. This text is then used to provide feedback to the user on the phonetic symbols the user mispronounced, improving the accuracy of pronunciation correction and enhancing user trust and overall user experience.

[0066] In some embodiments, the method for determining the contrasting phonetic symbol sequence and the competing phonetic symbol sequence can be further designed. The following will combine... Figure 3 This section will provide a detailed explanation of the process for determining the comparative and competing phonetic symbol sequences. Figure 3 A schematic flowchart of a pronunciation correction method according to another embodiment of this application is shown. Please refer to [link / reference]. Figure 3 The pronunciation correction method shown in the embodiments of this application may include:

[0067] In step S301, a corresponding phonetic sequence is determined based on the text to be evaluated. In this embodiment, an International Phonetic Alphabet (IPA) dictionary is pre-established, which can simultaneously retrieve the CMU phoneme sequence and the IPA sequence for the same word. For example, assuming the current word is "relax", the IPA dictionary can retrieve the CMU phoneme sequence of "relax" as "rih0 lae1 ks". Since the IPA for "relax" is... The IPA sequence for "relax" is: This International Phonetic Alphabet sequence is the standard pronunciation sequence of the text to be evaluated.

[0068] Specifically, it can be determined whether each word in the text to be evaluated exists in the International Phonetic Alphabet (IPA) dictionary. If it exists, a corresponding IPA sequence is determined based on the IPA symbol of each word in the text. In other words, the aforementioned IPA sequence is used as the corresponding IPA sequence.

[0069] If the phoneme sequence is not found, a G2P (grapheme to phoneme) dictionary tool is used to parse the pronunciation sequence of the text to be evaluated, and then the International Phonetic Alphabet (IPA) is mapped to obtain the corresponding phonetic transcription sequence. Commonly used G2P dictionary tools achieve this by querying the CMU pronunciation dictionary. Therefore, a G2P dictionary tool can parse the phoneme pronunciation sequence of the text to be evaluated, and then map the phoneme pronunciation sequence to an IPA sequence to obtain the corresponding phonetic transcription sequence.

[0070] In step S302, the speech frame of the speech to be evaluated is parsed using a preset acoustic processing model to obtain multiple phonemes and phoneme probability vectors corresponding to each speech frame in the speech to be evaluated. In this embodiment, the aforementioned preset acoustic processing model can be the TDNN model of Kaldi. Kaldi is a speech recognition tool that integrates multiple speech recognition models. The TDNN model refers to the Time Delay Neural Network model, which is one of the speech recognition models in Kaldi applied to speech recognition problems. It is understood that the implementation of the preset acoustic processing model can be diverse. In practical applications, a suitable speech recognition model should be selected as the preset acoustic processing model according to the actual application situation. The technical solution of this application is not bound by the underlying model and makes no restrictions in this regard.

[0071] After parsing the speech frames of the speech to be evaluated through a preset acoustic processing model, multiple pronunciation phonemes and phoneme probability vectors corresponding to each speech frame in the speech to be evaluated can be obtained, where the phoneme probability vector contains the phoneme probability corresponding to each pronunciation phoneme in the speech frame. As an example, assume that the text to be evaluated is "你" (nǐ), the speech to be evaluated formed after reading aloud lasts for 100 ms, and the duration of each speech frame is 10 ms. It can be understood that 10 speech frames can be formed at this time. Exemplarily, if there are deviations in the pronunciation used, then the pronunciation phonemes corresponding to the first frame are [ / n / , / l / ], and the phoneme probability vector corresponding to the first frame is [0.2, 0.8], which indicates that the pronunciation phonemes corresponding to the first frame may be / n / or / l / , the phoneme probability corresponding to / n / is 0.2, and the phoneme probability corresponding to / l / is 0.8. The pronunciation phonemes corresponding to the second frame are [ / n / , / l / ], and the phoneme probability vector corresponding to the second frame is [0.3, 0.7]. It is not until the pronunciation of the user changes in the sixth frame that the pronunciation phonemes corresponding to the sixth frame are [ / i / , / y / ], and the phoneme probability vector corresponding to the sixth frame is [0.9, 0.1].

[0072] In step S303, the corresponding reference phonetic symbols in the reference phonetic symbol sequence for each speech frame are determined according to the multiple pronunciation phonemes and phoneme probability vectors corresponding to each speech frame. As in the example in step S302, the pronunciation phonemes corresponding to the first frame to the fifth frame may all be [ / n / , / l / ] and all have specific phoneme probability vectors, and the pronunciation phonemes corresponding to the sixth frame to the tenth frame may all be [ / i / , / y / ] and all have specific phoneme probability vectors. Then, through the above TDNN model, the first frame to the fifth frame can be mapped to the reference phonetic symbol / n / in the reference phonetic symbol sequence of the text to be evaluated "你" (nǐ), and the sixth frame to the tenth frame can be mapped to the reference phonetic symbol / i / in the reference phonetic symbol sequence of the text to be evaluated "你" (nǐ), so as to divide all the speech frames in the speech to be evaluated into multiple speech frame groups. It can be understood that the number of reference phonetic symbols in the reference phonetic symbol sequence determines the number of speech frame groups. Based on the above operations, the alignment of speech frames and reference phonetic symbols can be achieved, effectively reducing the probability that phonemes with good pronunciation are judged as mispronounced, and improving the success rate and accuracy of pronunciation correction.

[0073] In step S304, competing phonetic symbols for each speech frame group are determined based on the reference phonetic symbol sequence, multiple phonemes corresponding to each speech frame, and phoneme probability vectors, to form a competing phonetic symbol sequence. In this embodiment, the phoneme score value corresponding to each phoneme in each speech frame group can be determined first based on the multiple phonemes corresponding to each speech frame in each speech frame group and phoneme probability vectors. Specifically, the phoneme score value corresponding to each phoneme in each speech frame group can be determined using the GOP scoring model based on the multiple phonemes corresponding to each speech frame in each speech frame group and phoneme probability vectors. As an example, the GOP scoring model can be:

[0074]

[0075] Where GOP(p) represents the phoneme score for each phoneme in each speech frame group, T is the total number of speech frames in each speech frame group, and P(s) t | t ) represents the t-th speech frame O in each speech frame group. t The corresponding phoneme is s t The probability value. For example, in step S302, the phoneme score value corresponding to / n / .

[0076] Then, the presence of diphthongs in the text to be evaluated can be determined by comparing the phonetic symbols. Diphthongs are represented by two letters, such as / ɑːr / . There is a smooth transition between diphthongs, that is, the pronunciation of a diphthong involves two different tongue positions, and the pronunciation slides from one tongue position to the other.

[0077] If there are no diphthongs in the text to be evaluated, the competing phonetic symbol for each speech frame group is determined based on the phoneme score value and the phoneme occurrence frequency of each phoneme in each speech frame group. For example, in step S302, assuming that the phoneme score value of the phoneme / l / in frames 1 to 5 is higher than that of the phoneme / n / , and the occurrence frequency of the phoneme / l / is higher than that of the phoneme / n / , then the competing phonetic symbol for the speech frame group corresponding to frames 1 to 5 can be determined to be / l / . Similarly, assuming that the phoneme score value of the phoneme / i / in frames 6 to 10 is higher than that of the phoneme / y / , and the occurrence frequency of the phoneme / i / is higher than that of the phoneme / y / , then the competing phonetic symbol for the speech frame group corresponding to frames 6 to 10 can be determined to be / i / . Understandably, the method for determining the competing phonetic symbols examines the phonetic symbols corresponding to the phonemes with the highest phoneme score and the highest frequency of occurrence. Furthermore, the competing phonetic symbols and their corresponding reference phonetic symbols can be the same. If the competing phonetic symbols and their corresponding reference phonetic symbols are the same, it means that the user's pronunciation is correct; otherwise, it means that the pronunciation is incorrect, and it is necessary to determine whether to correct the pronunciation on the spot.

[0078] If the text to be evaluated contains diphthongs, the two speech frames corresponding to the diphthongs are merged into a single diphthong group. The phonetic transcription score for each phoneme in the diphthong group is then determined based on the phoneme score for each phoneme in the two speech frames corresponding to the diphthongs. For example, assuming the text to be evaluated is the word "remark," its corresponding phonetic transcription sequence is as follows: Its CMU phoneme sequence is "r ih0 m aa1 rk". It contains the diphthong / ɑ:r / , corresponding to the phonemes aa1 and r. In this case, the two phonemes aa1 and r need to be merged into a single phonetic symbol to avoid misalignment between the speech frame and the corresponding phonetic symbol. For example, if phoneme aa1 corresponds to speech frames 16 to 20, and phoneme r corresponds to speech frames 21 to 25, then the speech frame groups corresponding to frames 16 to 20 and 21 to 25 are merged into a diphthong group. Furthermore, the average phoneme score calculated based on the phoneme score values ​​corresponding to phoneme aa1 and r can be used as the phonetic symbol score value for the diphthong / ɑ:r / in the diphthong group (it is understood that when the user's pronunciation is inaccurate, other possible phonetic symbols may exist in the diphthong group, such as / er / , etc.). Specifically, if the phoneme score for phoneme aa1 and the phoneme score for phoneme r are both higher than the preset score, then the average phoneme score can be multiplied by a preset coefficient, and the result can be used as the phonetic score for the diphthong / ɑ:r / . The preset score can be set to 60 points (the phoneme score is converted to a percentage before comparison), and the preset coefficient can be set to 1.5. In practical applications, the preset score and preset coefficient need to be set according to the actual application situation; this application does not impose any restrictions in this regard.

[0079] Furthermore, competing phonetic symbols for a diphthong group can be determined based on the phonetic score of each phonetic symbol within the diphthong group and the frequency of occurrence of each phonetic symbol within the diphthong group. For example, if the phonetic score of / ɑ:r / in the diphthong group is higher than that of / er / , and the frequency of occurrence of / ɑ:r / is higher than that of / er / , then / ɑ:r / can be identified as a competing phonetic symbol for the current diphthong group.

[0080] Specifically, in some applications, the preceding phonemes of the current phoneme are considered. If the preceding phoneme is pronounced incorrectly, the score for the current phoneme needs to be recalculated. Specifically, the range of frames used for calculation can be expanded for the speech frames corresponding to the current phoneme, for example, expanding forward 3 to 5 frames from the first frame. Simultaneously, after calculating the new score, competing phonemes are re-checked and identified.

[0081] In some embodiments, when determining the pronunciation correction text, the format of the pronunciation correction text is further designed, and after the pronunciation correction text is determined, it is fed back to the user so that the user can receive personalized pronunciation correction feedback and improve the user's human-computer interaction experience. The following will combine... Figure 4 This section will provide a detailed explanation of the process for determining and processing pronunciation correction texts. Figure 4A schematic flowchart of a pronunciation correction method according to another embodiment of this application is shown. Please refer to [link / reference]. Figure 4 The pronunciation correction method shown in the embodiments of this application may include:

[0082] In step S401, the word score for each word in the text to be evaluated is determined based on the phoneme score for each phoneme in each speech frame group. In this embodiment, the word score for each word can be the average of the phoneme scores for each phoneme in each speech frame group. Further, the calculated word score can be converted to a percentage system. It is understood that the above method for determining word scores is merely exemplary. In practical applications, the method for determining word scores should be reasonably selected based on the actual application situation, and this application makes no restrictions in this regard.

[0083] In step S402, it is determined whether to perform pronunciation correction on words in the text to be evaluated based on the word score. To avoid over-correction during pronunciation correction and to ensure no omissions, this application embodiment establishes a series of rules to determine whether to perform pronunciation correction on words in the text to be evaluated. Specifically, the above series of rules may include, but are not limited to: First, if the word score is within a preset score range, pronunciation correction is performed; otherwise, pronunciation correction is not performed. The preset score range can be set to 30 to 90 points, depending on the actual application; this application does not impose any restrictions in this regard. Second, if the number of phonemes requiring pronunciation correction in the current word exceeds the preset number of corrections, then the one or two phonemes with the lowest phoneme scores are selected for pronunciation correction. The preset number of corrections can be set to three, depending on the actual application; this application does not impose any restrictions in this regard. Third, some phonemes that are difficult to distinguish can be left uncorrected, such as / s / and / θ / . If the competing phoneme for / s / is / θ / , then no pronunciation correction is needed, and vice versa.

[0084] It is understood that the above description of the judgment rules for whether to correct the pronunciation of words in the evaluation text is merely exemplary. In practical applications, the judgment rules need to be set according to the actual application situation, and this application does not impose any restrictions in this regard.

[0085] In step S403, the pronunciation correction text is determined based on the reference phonetic symbol sequence and the competing phonetic symbol sequence. Specifically, the evaluation result can be determined based on the reference phonetic symbol sequence and the competing phonetic symbol sequence. The evaluation result includes the competing phonetic symbols and the corresponding reference phonetic symbols. In this embodiment, the evaluation result can be a JSON format return result. This return result records the current pronunciation evaluation status, and in addition to including the competing phonetic symbols and the corresponding reference phonetic symbols, it can also include, but is not limited to, the current text to be evaluated and the score value, etc. Then, the pronunciation correction text is determined based on the preset text format, competing phonetic symbols, and reference phonetic symbols using the LLM text output model. Taking "relax" as an example, the pronunciation correction text can be as follows:

[0086] Not bad! The pronunciation of / k / needs further attention. Don't read it as Don't pronounce / k / as / t / . Follow me: / l / , / l / ; / k / , / k / . Come on, let's try again.

[0087] When the word score is above 90, "Not bad" can be replaced with "Very good". When the word score is between 60 and 90, "Not bad" will be displayed.

[0088] In step S404, the pronunciation correction text is converted into an audio file using TTS speech synthesis technology, and the audio file is played. Specifically, a speech synthesis technology that supports complex pronunciations and phonetics can be selected to convert the pronunciation correction text into an audio file, which is then played on the user interface of a user terminal such as a smart learning tablet or a smart learning lamp.

[0089] Exemplary device

[0090] After introducing the methods of exemplary embodiments of this application, the following references are made. Figure 5 and Figure 6 Products related to the pronunciation correction method according to exemplary embodiments of the present invention will be described.

[0091] Figure 5 A schematic diagram of a speech correction device according to another embodiment of this application is shown. Please refer to... Figure 5 The pronunciation correction device shown in the embodiments of this application may include:

[0092] The data acquisition module 501 is used to acquire the text and speech to be evaluated.

[0093] The phonetic symbol sequence determination module 502 is used to determine the reference phonetic symbol sequence based on the text to be evaluated.

[0094] Analysis and processing module 503 is used to perform speech analysis and processing on the speech to be evaluated to obtain a competitive phonetic symbol sequence;

[0095] The correction text determination module 504 is used to determine the pronunciation correction text based on the reference phonetic symbol sequence and the competing phonetic symbol sequence.

[0096] The pronunciation correction device disclosed in this application acquires the text to be evaluated and the speech to be evaluated. On the one hand, it determines a reference phonetic symbol sequence based on the text to be evaluated; on the other hand, it performs speech analysis processing on the speech to be evaluated to obtain a competing phonetic symbol sequence. This avoids the negative impact of speech recognition systems that could reduce the success rate and accuracy of pronunciation correction. Furthermore, it determines the pronunciation correction text based on the reference phonetic symbol sequence and the competing phonetic symbol sequence, and uses the pronunciation correction text to provide feedback to the user on the phonetic symbols that the user mispronounced, thereby improving the accuracy of pronunciation correction and increasing the user's trust in the pronunciation correction and the user experience.

[0097] Figure 6 A schematic block diagram of an electronic device according to an embodiment of this application is shown. Please refer to... Figure 6 The electronic device 600 may include a processor 601. Furthermore, the electronic device may also include a memory 602 storing computer instructions that, when executed by the processor 601, cause the electronic device 600 to perform the methods described in the foregoing embodiments or implementations.

[0098] In some implementation scenarios, electronic device 600 may include server or terminal devices, such as physical servers, cloud servers, server clusters, data processing devices, application testing robots, computer terminals, smart terminals, PC devices, and Internet of Things terminals, etc.

[0099] Depending on the implementation scenario, the processor 601 mentioned above can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor can be a microprocessor or any conventional processor.

[0100] Based on the foregoing, this application also discloses a computer-readable storage medium containing program instructions that, when executed by a processor, cause the methods described according to the foregoing embodiments or implementations to be performed.

[0101] In some implementation scenarios, the aforementioned computer-readable storage medium can be any suitable magnetic or magneto-optical storage medium, such as Resistive Random Access Memory (RRAM), Dynamic Random Access Memory (DRAM), Static Random-Access Memory (SRAM), Enhanced Dynamic Random Access Memory (EDRAM), High-Bandwidth Memory (HBM), Hybrid Memory Cube (HMC), etc., or any other medium that can be used to store required information and can be accessed by an application, module, or both. Any such computer storage medium can be part of a device or accessible to or connected to a device. Any application or module described in this invention can be implemented using computer-readable / executable instructions that can be stored or otherwise maintained by such a computer-readable medium.

[0102] It should be noted that although several devices or sub-devices of the pronunciation correction device have been mentioned in the detailed description above, this division is not mandatory. In fact, according to the embodiments of this application, the features and functions of two or more devices described above can be embodied in one device. Conversely, the features and functions of one device described above can be further divided and embodied by multiple devices.

[0103] Furthermore, although the operations of the method of this application are described in a specific order in the accompanying drawings, this does not require or imply that these operations must be performed in that specific order, or that all of the operations shown must be performed to achieve the desired result. On the contrary, the steps depicted in the flowchart may be performed in a different order. Additionally or alternatively, certain steps may be omitted, multiple steps may be combined into one step, and / or one step may be broken down into multiple steps.

[0104] The use of the verbs "including" and "contains" and their inflections in the application documents does not preclude the existence of elements or steps other than those described in the application documents. The article "a" or "one" preceding an element does not preclude the existence of multiple such elements.

[0105] While the spirit and principles of this application have been described with reference to several specific embodiments, it should be understood that this application is not limited to the disclosed specific embodiments, and the division of aspects does not imply that features in these aspects cannot be combined for benefit; such division is merely for convenience of expression. This application is intended to cover various modifications and equivalent arrangements included within the spirit and scope of the appended claims. The scope of the appended claims is to be interpreted in the broadest sense, thereby encompassing all such modifications and equivalent structures and functions.

Claims

1. A pronunciation correction method, characterized in that, include: Obtain the text and audio to be evaluated; Determine the corresponding phonetic symbol sequence based on the text to be evaluated; The speech to be evaluated is subjected to speech analysis processing to obtain a competing phonetic symbol sequence; as well as The pronunciation correction text is determined based on the reference phonetic sequence and the competing phonetic sequence; The step of performing speech analysis processing on the speech to be evaluated to obtain a competing phonetic symbol sequence includes: The speech to be evaluated is parsed by a preset acoustic processing model to obtain multiple phonemes and phoneme probability vectors corresponding to each speech frame in the speech to be evaluated; wherein, the phoneme probability vector contains the phoneme probability corresponding to each phoneme in the speech frame, and the preset acoustic processing model is a TDNN model. Based on multiple phonemes and phoneme probability vectors corresponding to each speech frame, the corresponding reference phonetic symbol in the reference phonetic symbol sequence is determined for each speech frame, so as to divide all speech frames in the speech to be evaluated into multiple speech frame groups; and The phoneme score for each phoneme in each speech frame group is determined based on multiple phonemes and phoneme probability vectors corresponding to each speech frame in each speech frame group; and The competing phonetic symbols for each speech frame group are determined based on the phoneme score value corresponding to each phoneme in each speech frame group and the phoneme occurrence frequency of each phoneme in each speech frame group, so as to form a competing phonetic symbol sequence. After determining the phoneme score value corresponding to each phoneme in each speech frame group based on multiple phonemes and phoneme probability vectors corresponding to each speech frame in each speech frame group, the method further includes: Based on the aforementioned phonetic symbol sequence, it is determined that diphthongs exist in the text to be evaluated. The two speech frames corresponding to the diphthongs are merged into a single diphthong group. Furthermore, based on the phoneme score values ​​corresponding to each phoneme in the two speech frame groups corresponding to the diphthongs, the phonetic symbol score values ​​corresponding to each phoneme in the diphthong group are determined. The competing phonetic symbols of the diphthong group are determined based on the phonetic score value corresponding to each phonetic symbol in the diphthong group and the phonetic frequency of each phonetic symbol in the diphthong group.

2. The pronunciation correction method according to claim 1, characterized in that, The step of determining the phoneme score value corresponding to each phoneme in each speech frame group based on multiple phonemes and phoneme probability vectors corresponding to each speech frame in each speech frame group includes: The GOP scoring model determines the phoneme score value corresponding to each phoneme in each speech frame group based on multiple phonemes and phoneme probability vectors corresponding to each speech frame in each speech frame group. The GOP scoring model is as follows: in, This represents the phoneme score corresponding to each phoneme in each speech frame group, where T is the total number of speech frames in each speech frame group. This represents the t-th speech frame in each speech frame group. The corresponding phonemes are The probability value.

3. The pronunciation correction method according to claim 1, characterized in that, Before determining the pronunciation correction text based on the reference phonetic sequence and the competing phonetic sequence, the method further includes: The word score for each word in the text to be evaluated is determined based on the phoneme score for each phoneme in each speech frame group; and Whether to perform pronunciation correction on the words in the text to be evaluated is determined based on the word score.

4. The pronunciation correction method according to claim 3, characterized in that, The process of determining the pronunciation correction text based on the reference phonetic sequence and the competing phonetic sequence includes: The evaluation result is determined based on the reference phonetic symbol sequence and the competing phonetic symbol sequence, the evaluation result including the competing phonetic symbol and the reference phonetic symbol corresponding to the competing phonetic symbol; and The pronunciation correction text is determined by the text output model based on a preset text format, the competing phonetic symbols, and the reference phonetic symbols.

5. The pronunciation correction method according to claim 1, characterized in that, After determining the pronunciation correction text based on the reference phonetic sequence and the competing phonetic sequence, the method further includes: Converting pronunciation-corrected text into audio files using TTS speech synthesis technology; and Play the audio file.

6. The pronunciation correction method according to claim 1, characterized in that, The step of determining the reference phonetic sequence based on the text to be evaluated includes: Determine whether each word in the text to be evaluated exists in the International Phonetic Alphabet dictionary; If present, the corresponding phonetic sequence is determined based on the International Phonetic Alphabet (IPA) representation of each word in the text to be evaluated in the IPA dictionary; and If it does not exist, the pronunciation sequence of the text to be evaluated is parsed using the G2P dictionary tool, and the International Phonetic Alphabet is mapped based on the pronunciation sequence to obtain the reference phonetic symbol sequence.

7. A pronunciation correction device, characterized in that, For performing the pronunciation correction method as described in any one of claims 1-6, comprising: The data acquisition module is used to acquire the text and audio to be evaluated. A phonetic symbol sequence determination module is used to determine a reference phonetic symbol sequence based on the text to be evaluated. The analysis and processing module is used to perform speech analysis and processing on the speech to be evaluated to obtain a competitive phonetic symbol sequence; The correction text determination module is used to determine the pronunciation correction text based on the reference phonetic symbol sequence and the competing phonetic symbol sequence; Specifically, the analysis and processing module is used for: The speech to be evaluated is parsed by a preset acoustic processing model to obtain multiple phonemes and phoneme probability vectors corresponding to each speech frame in the speech to be evaluated; wherein, the phoneme probability vector contains the phoneme probability corresponding to each phoneme in the speech frame, and the preset acoustic processing model is a TDNN model. Based on multiple phonemes and phoneme probability vectors corresponding to each speech frame, the corresponding reference phonetic symbol in the reference phonetic symbol sequence is determined for each speech frame, so as to divide all speech frames in the speech to be evaluated into multiple speech frame groups; and The phoneme score for each phoneme in each speech frame group is determined based on multiple phonemes and phoneme probability vectors corresponding to each speech frame in each speech frame group; and The competing phonetic symbols for each speech frame group are determined based on the phoneme score value corresponding to each phoneme in each speech frame group and the phoneme occurrence frequency of each phoneme in each speech frame group, so as to form a competing phonetic symbol sequence. After determining the phoneme score value corresponding to each phoneme in each speech frame group based on multiple phonemes and phoneme probability vectors corresponding to each speech frame in each speech frame group, the method further includes: Based on the aforementioned phonetic symbol sequence, it is determined that diphthongs exist in the text to be evaluated. The two speech frames corresponding to the diphthongs are merged into a single diphthong group. Furthermore, based on the phoneme score values ​​corresponding to each phoneme in the two speech frame groups corresponding to the diphthongs, the phonetic symbol score values ​​corresponding to each phoneme in the diphthong group are determined. The competing phonetic symbols of the diphthong group are determined based on the phonetic score value corresponding to each phonetic symbol in the diphthong group and the phonetic frequency of each phonetic symbol in the diphthong group.

8. An electronic device, characterized in that, include: processor; as well as A memory having stored executable code for pronunciation correction, which, when executed by the processor, causes the processor to perform the method as described in any one of claims 1-6.

9. A non-transitory machine-readable storage medium having stored executable code for pronunciation correction, which, when executed by a processor of an electronic device, causes the processor to perform the method as described in any one of claims 1-6.