Pronunciation Error Detection Method, Device, Computer Equipment and Storage Medium

By combining pronunciation information and standard pronunciation text information for speech recognition, the problem of insufficient accuracy of pronunciation error detection in traditional technology is solved, more accurate pronunciation error detection and correction is achieved, and learners' oral skills are improved.

CN114283852BActive Publication Date: 2025-06-13TENCENT TECHNOLOGY (SHENZHEN) CO LTD +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202111004261.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-08-30
Publication Date
2025-06-13
Estimated Expiration
2041-08-30

AI Technical Summary

Technical Problem

Traditional speech recognition technology has insufficient accuracy in pronunciation error detection, making it difficult to effectively help learners correct pronunciation problems.

Method used

By obtaining the pronunciation information and its corresponding standard pronunciation text information, speech recognition is performed to obtain predicted pronunciation text information, and pronunciation bias detection is performed based on the predicted and standard text information.

Benefits of technology

Improve the accuracy of pronunciation error detection, allowing learners to more accurately identify and correct pronunciation errors, thereby improving their oral skills.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114283852B_ABST
    Figure CN114283852B_ABST
Patent Text Reader

Abstract

The present application relates to a pronunciation error detection method, apparatus, computer device, and storage medium. The method includes: obtaining speech information and standard pronunciation text information corresponding to the speech information; performing speech recognition based on the speech information and the standard pronunciation text information to obtain predicted pronunciation text information corresponding to the speech information; and detecting pronunciation errors of the speech information based on the predicted pronunciation text information and the standard pronunciation text information. Using this method can improve the accuracy of pronunciation error detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of speech recognition technology, and particularly to a pronunciation error detection method, device, computer device, and storage medium. Background Art

[0002] With the development of speech technology and the popularization of online learning, Computer-Aided Pronunciation Training (CAPT) has been increasingly applied in language teaching. As an important part of computer-aided pronunciation teaching, automatic pronunciation error detection is mainly used to detect the pronunciation errors of learners, so as to help learners timely discover their pronunciation problems and correct them during the language learning process.

[0003] In traditional technologies, when a learner reads aloud according to a follow-up text in a follow-up reading scenario, only the speech of the learner is used to recognize the reading content, and then pronunciation error detection is performed based on the recognition result and the original follow-up text, and the detection effect needs to be improved. Summary of the Invention

[0004] Based on this, in view of the above technical problems, it is necessary to provide a pronunciation error detection method, device, computer device, and storage medium that can improve the accuracy of pronunciation error detection.

[0005] A pronunciation error detection method, the method includes:

[0006] Obtain speech information and the standard pronunciation text information corresponding to the speech information;

[0007] Perform speech recognition according to the speech information and the standard pronunciation text information to obtain the predicted pronunciation text information corresponding to the speech information;

[0008] Detect the pronunciation error of the speech information according to the predicted pronunciation text information and the standard pronunciation text information.

[0009] A pronunciation error detection device, the device includes:

[0010] An obtaining module, configured to obtain speech information and the standard pronunciation text information corresponding to the speech information;

[0011] A recognition module, configured to perform speech recognition according to the speech information and the standard pronunciation text information to obtain the predicted pronunciation text information corresponding to the speech information;

[0012] A detection module, configured to detect the pronunciation error of the speech information according to the predicted pronunciation text information and the standard pronunciation text information.

[0013] A computer device, comprising a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:

[0014] Obtain voice information and standard pronunciation text information corresponding to the voice information;

[0015] Perform speech recognition based on the voice information and the standard pronunciation text information to obtain predicted pronunciation text information corresponding to the voice information;

[0016] Detect pronunciation errors of the voice information according to the predicted pronunciation text information and the standard pronunciation text information.

[0017] A computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the following steps are implemented:

[0018] Obtain voice information and standard pronunciation text information corresponding to the voice information;

[0019] Perform speech recognition based on the voice information and the standard pronunciation text information to obtain predicted pronunciation text information corresponding to the voice information;

[0020] Detect pronunciation errors of the voice information according to the predicted pronunciation text information and the standard pronunciation text information.

[0021] A pronunciation error detection method, the method comprising:

[0022] Display a first page, where the first page includes target content;

[0023] In response to a pronunciation operation on the target content on the first page, receive voice information corresponding to the pronunciation operation;

[0024] When a pronunciation operation end condition is met, display a second page, where the second page includes a pronunciation error detection result, and mark the content corresponding to the pronunciation error part in the target content in the pronunciation error detection result;

[0025] Wherein, the pronunciation error detection result is obtained by detecting pronunciation errors of the voice information according to the predicted pronunciation text information corresponding to the voice information and standard pronunciation text information, and the predicted pronunciation text information is obtained by performing speech recognition according to the voice information and the standard pronunciation text information.

[0026] A pronunciation error detection device, the device comprising:

[0027] A display module, configured to display a first page, where the first page includes target content;

[0028] A receiving module, configured to receive voice information corresponding to the pronunciation operation in response to a pronunciation operation on the target content on the first page;

[0029] The display module is further configured to display a second page when a pronunciation operation end condition is met. The second page includes a pronunciation error detection result, and the content in the target content corresponding to the pronunciation error part is marked in the pronunciation error detection result;

[0030] Wherein, the pronunciation error detection result is obtained by performing pronunciation error detection on the voice information according to predicted pronunciation text information and standard pronunciation text information corresponding to the voice information, and the predicted pronunciation text information is obtained by performing speech recognition according to the voice information and the standard pronunciation text information.

[0031] A computer device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:

[0032] Display a first page, where the first page includes target content;

[0033] In response to a pronunciation operation on the target content on the first page, receive voice information corresponding to the pronunciation operation;

[0034] When a pronunciation operation end condition is met, display a second page, where the second page includes a pronunciation error detection result, and the content in the target content corresponding to the pronunciation error part is marked in the pronunciation error detection result;

[0035] Wherein, the pronunciation error detection result is obtained by performing pronunciation error detection on the voice information according to predicted pronunciation text information and standard pronunciation text information corresponding to the voice information, and the predicted pronunciation text information is obtained by performing speech recognition according to the voice information and the standard pronunciation text information.

[0036] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the following steps are implemented:

[0037] Display a first page, where the first page includes target content;

[0038] In response to a pronunciation operation on the target content on the first page, receive voice information corresponding to the pronunciation operation;

[0039] When a pronunciation operation end condition is met, display a second page, where the second page includes a pronunciation error detection result, and the content in the target content corresponding to the pronunciation error part is marked in the pronunciation error detection result;

[0040] Among them, the pronunciation error detection result is obtained by performing pronunciation error detection on the speech information according to the predicted pronunciation text information and the standard pronunciation text information corresponding to the speech information, and the predicted pronunciation text information is obtained by performing speech recognition on the speech information and the standard pronunciation text information.

[0041] The above pronunciation error detection method, device, computer device and storage medium introduce the standard pronunciation text information in speech recognition, provide richer reference information for speech recognition, make the speech recognition result more accurate, and thus can more precisely detect the pronunciation errors of learners, enabling learners to focus their attention on error correction and thus more efficiently improve their oral English ability. Brief Description of the Drawings

[0042] Figure 1 It is an application environment diagram of the pronunciation error detection method in an embodiment;

[0043] Figure 2 It is a schematic flowchart of the pronunciation error detection method in an embodiment;

[0044] Figure 3 It is a schematic flowchart of the step of performing speech recognition on the speech information and the standard pronunciation text information to obtain the predicted pronunciation text information corresponding to the speech information in an embodiment;

[0045] Figure 4 It is a schematic flowchart of the pronunciation error detection method in an embodiment;

[0046] Figure 5 It is a schematic diagram of the first page in an embodiment;

[0047] Figure 6 It is a schematic diagram of the second page in an embodiment;

[0048] Figure 7 It is a schematic diagram of the third page in an embodiment;

[0049] Figure 8 It is an interaction schematic diagram of the pronunciation error detection method in an embodiment;

[0050] Figure 9 It is a structural block diagram of the pronunciation error detection device in an embodiment;

[0051] Figure 10 It is a structural block diagram of the pronunciation error detection device in an embodiment;

[0052] Figure 11 It is an internal structure diagram of the computer device in an embodiment;

[0053] Figure 12Internal structure diagram of a computer device in an embodiment. Detailed implementation manners

[0054] To make the objectives, technical solutions and advantages of the present application more clear and understandable, the present application will be further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0055] The pronunciation error detection method provided by the present application can be applied to, for example Figure 1 the application environment shown in the figure. Among them, the terminal 102 communicates with the server 104 through the network. An application program providing language learning-related services (such as pronunciation practice, pronunciation error detection) can be installed on the terminal 102, and the server 104 can be the server where the application program is located. The user can access the application program through the terminal 102 and perform pronunciation practice operations. After the terminal 102 receives the voice information of the pronunciation practice, it transmits it to the server 104. The server 104 performs pronunciation error detection on the voice information and returns the pronunciation error detection result to the terminal 102. The terminal 102 displays the pronunciation error detection result to help the user timely discover and correct pronunciation problems during the pronunciation practice process. Among them, the terminal 102 can be but is not limited to various personal computers, laptop computers, smart phones, tablet computers and portable wearable devices, and the server 104 can be implemented by an independent server or a server cluster composed of multiple servers.

[0056] In one embodiment, as Figure 2 shown in the figure, a pronunciation error detection method is provided. Taking the method applied to Figure 1 the server in the figure as an example, it includes the following steps S202 to step S206.

[0057] S202, obtain the voice information and the standard pronunciation text information corresponding to the voice information.

[0058] The voice information refers to the information that needs to be detected for pronunciation errors, and it is in the form of voice. The standard pronunciation text information corresponding to the voice information is used to describe the standard pronunciation corresponding to the voice information, and it is in the form of text.

[0059] Specifically, the pronunciation practice page can be displayed through the terminal. The user performs pronunciation practice operations on the target content to be practiced on the pronunciation practice page. The terminal receives the voice information of the user's pronunciation practice and transmits the voice information and the corresponding standard pronunciation text information to the server. The server obtains the voice information and the corresponding standard pronunciation text information uploaded by the terminal. Among them, the target content can be but is not limited to words, phrases or sentences. It can be understood that the standard pronunciation text information corresponding to the voice information here is the standard pronunciation text information corresponding to the target content.

[0060] In one embodiment, the standard pronunciation text information may be a standard phoneme sequence. Phonemes are the smallest speech units divided according to the natural properties of speech. The pronunciation characteristics can be described more clearly through phoneme division. For example, if the speech information is speech information for pronunciation practice of the word "afternoon", the standard pronunciation text information corresponding to the speech information is

[0061] S204, performing speech recognition according to the speech information and the standard pronunciation text information, and obtaining predicted pronunciation text information corresponding to the speech information.

[0062] In the process of speech recognition of speech information, standard pronunciation text information corresponding to the speech information is introduced, and speech recognition is performed on the speech information in combination with the speech information and the standard pronunciation text information corresponding to the speech information to obtain predicted pronunciation text information corresponding to the speech information. The predicted pronunciation text information refers to the pronunciation text information obtained through speech recognition. In one embodiment, the predicted pronunciation text information may be a phoneme sequence obtained through speech recognition.

[0063] S206: Detect pronunciation errors of the speech information based on the predicted pronunciation text information and the standard pronunciation text information.

[0064] The pronunciation error of the speech information can be detected by comparing the predicted pronunciation text information with the standard pronunciation text information to obtain the pronunciation error detection result. In one embodiment, the predicted pronunciation text information includes a predicted phoneme sequence, and the standard pronunciation text information includes a standard phoneme sequence. The predicted phoneme sequence is compared with each phoneme in the standard phoneme sequence to obtain the pronunciation error detection result.

[0065] The pronunciation error detection result may include only the correct pronunciation part or the pronunciation error part, or may include both the correct pronunciation part and the pronunciation error part. The correct pronunciation part refers to the part where the comparison result between the predicted pronunciation text information and the standard pronunciation text information is consistent, and the pronunciation error part refers to the part where the comparison result between the predicted pronunciation text information and the standard pronunciation text information is inconsistent.

[0066] For example, if the standard pronunciation text information is "你好" and the predicted pronunciation text information is "你好", then it can be known that the user mistakenly pronounced "你" as "你", that is, the character corresponding to the pronunciation error in the standard pronunciation text information is "你", and the user can focus on practicing this character next time.

[0067] In the above pronunciation error detection method, by introducing standard pronunciation text information in speech recognition, richer reference information is provided for speech recognition, making the speech recognition result more accurate. Thus, the pronunciation errors of learners can be detected more precisely, enabling learners to focus their attention on error correction and thereby more efficiently improving their oral English ability.

[0068] In one embodiment, the step of performing speech recognition based on the speech information and the standard pronunciation text information to obtain the predicted pronunciation text information corresponding to the speech information may specifically include: performing speech recognition based on the speech information and the standard pronunciation text information through a trained speech recognition model to obtain the predicted pronunciation text information corresponding to the speech information.

[0069] The speech recognition model here is an end-to-end model. The model input includes the speech information and the standard pronunciation text information, and the model output is the predicted pronunciation text information corresponding to the speech information, that is, the entire speech recognition task is completed through a single model. The end-to-end model can reduce the model complexity and improve the model processing efficiency.

[0070] In one embodiment, the training method of the speech recognition model includes: obtaining sample speech information and the corresponding sample standard pronunciation text information of the sample speech information; performing speech recognition based on the sample speech information and the sample standard pronunciation text information through the speech recognition model to be trained to obtain the sample predicted pronunciation text information corresponding to the sample speech information; when the training end condition is not met, adjusting the parameters of the speech recognition model to be trained based on the difference between the sample predicted pronunciation text information and the corresponding sample standard pronunciation text information until the training end condition is met, and obtaining the trained speech recognition model.

[0071] The sample speech information refers to the information that needs to be recognized in the training process, which is in speech form. The sample standard pronunciation text information corresponding to the sample speech information is used to describe the standard pronunciation corresponding to the sample speech information, which is in text form. During the model training process, the model input includes the sample speech information and the sample standard pronunciation text information, and the model output is the sample predicted pronunciation text information corresponding to the sample speech information.

[0072] It should be noted that the sample standard pronunciation text information is not only used as the input of the model, but also used as the true label corresponding to the sample voice information. The training objective of the model is to make the predicted pronunciation text information of the sample output by the model as close as possible to the corresponding true label (i.e., the sample standard pronunciation text information). A loss function can be established to represent the difference between the predicted pronunciation text information of the sample and the corresponding sample standard pronunciation text information. In one embodiment, when the value of the loss function is less than a preset threshold, it is considered that the training end condition is satisfied. In other embodiments, it can also be considered that the training end condition is satisfied when the number of iterations reaches a preset number. Among them, the preset threshold and the preset number can both be set according to actual needs and are not limited here.

[0073] In one embodiment, as Figure 3 shown, the steps of performing speech recognition according to the voice information and the standard pronunciation text information to obtain the predicted pronunciation text information corresponding to the voice information may specifically include the following steps S302 to S310.

[0074] S302, preprocess the voice information to obtain each frame of voice segment, and extract the acoustic features of each frame of voice segment.

[0075] Voice information is macroscopically unstable and microscopically stable, with short-term stationarity. In one embodiment, it can be considered that the voice information is approximately unchanged within 10 to 30 ms. Based on this, the voice information can be divided into multiple short-term voice segments for processing.

[0076] In one embodiment, the steps of preprocessing the voice information to obtain each frame of voice segment may specifically include: performing pre-emphasis, framing, and windowing processing on the voice information to obtain each frame of voice segment.

[0077] The purpose of pre-emphasis is to enhance the high frequency of the voice information to a certain extent and remove the influence of oral radiation. The pre-emphasis formula is as follows:

[0078] y(n) = x(n) - ax(n - 1)

[0079] where x(n) represents the voice signal at the current moment, x(n - 1) represents the voice signal at the previous moment, the coefficient a can take a value of 0.98, and y(n) represents the voice signal at the current moment after pre-emphasis processing.

[0080] After performing pre-emphasis processing on the voice information, then perform framing and windowing processing on the voice information. Specifically, with a frame length of 25 ms and a frame shift of 10 ms, the voice information is decomposed into a sequence of voice segments each 25 ms long, and each frame of voice segment in the sequence is windowed, such as adding a Hamming window. After the above preprocessing, it is convenient to perform time-frequency conversion on the voice information to better extract acoustic features.

[0081] In one embodiment, the step of extracting the acoustic features of each frame of speech segment may specifically include: performing Fourier transform and Mel filtering on each frame of speech segment to obtain the acoustic features of each frame of speech segment.

[0082] Perform fast Fourier transform (FFT) on each frame of speech segment, thereby transforming the speech information from the time domain to the frequency domain. Then, perform Mel filtering on this group of speech frame sequences in the frequency domain frame by frame to extract the features available for subsequent models. This can essentially be understood as a process of information compression and abstraction.

[0083] The features that can be extracted in this stage include various types, such as spectral features (MFCC, FBANK, PLP, etc.), frequency features (fundamental frequency, formant, etc.), time-domain features (duration), energy features, etc.

[0084] In one embodiment, the acoustic features are composed of 80-dimensional FBANK features and 1-dimensional energy features, that is, 81-dimensional features. Accordingly, by combining spectral features and energy features, the features of speech information can be better described, which helps to improve the accuracy of speech recognition.

[0085] S304, Encode the acoustic features of each frame of speech segment to obtain the acoustic vector information corresponding to each moment.

[0086] Specifically, a CNN-BLSTM network can be used to encode the acoustic features of each frame of speech segment. The features input at each time step of the CNN-BLSTM network include the acoustic features of the previous frame, the current frame, and the next frame. For example, if the acoustic features are 81-dimensional features in the previous embodiment, the features input at each time step of the CNN-BLSTM network are a total of 243 dimensions (i.e., 3 * 81). The CNN-BLSTM network encodes the input features and outputs the acoustic vector information corresponding to each moment, denoted as Query.

[0087] S306, Encode each pronunciation unit in the standard pronunciation text information to obtain the text vector information corresponding to each pronunciation unit.

[0088] When the standard pronunciation text information is a phoneme sequence, the pronunciation unit is a phoneme. Specifically, each phoneme can be first mapped to a high-dimensional embedding vector (sentence embedding) through a preprocessing layer, and then encoded through a BLSTM layer to output the text vector information corresponding to each phoneme.

[0089] In one embodiment, the text vector information includes first text vector information and second text vector information, and the second text vector information is obtained by linearly processing the first text vector information.

[0090] Specifically, the first text vector information is obtained through encoding by the BLSTM layer, denoted as Value. After passing through the BLSTM layer, it then passes through a linear layer to obtain the second text vector information, denoted as Key. Accordingly, by combining the first text vector information and the second text vector information, richer pronunciation text features can be obtained, which helps to improve the accuracy of subsequent speech recognition.

[0091] S308, fuse the acoustic vector information corresponding to each moment and the text vector information corresponding to each pronunciation unit to obtain the fusion vector corresponding to each moment.

[0092] After the acoustic features and text features are both encoded, an attention mechanism is introduced to fuse the acoustic vector information and the text vector information. Specifically, according to the acoustic vector information corresponding to each moment and the text vector information corresponding to each pronunciation unit, the attention vector information corresponding to each moment is obtained; the attention vector information corresponding to each moment and the acoustic vector information are fused to obtain the fusion vector corresponding to each moment.

[0093] Among them, the step of obtaining the attention vector information corresponding to each moment according to the acoustic vector information corresponding to each moment and the text vector information corresponding to each pronunciation unit may specifically include: obtaining the weight coefficient corresponding to each moment according to the acoustic vector information corresponding to each moment and the second text vector information corresponding to each pronunciation unit; obtaining the attention vector information corresponding to each moment according to the weight coefficient corresponding to each moment and the first text vector information corresponding to each pronunciation unit.

[0094] The calculation formulas for the weight coefficient and the attention vector information corresponding to each moment are as follows:

[0095]

[0096] Among them, h represents a vector, and the superscripts Q, K, V respectively represent Query, Key, Value, and the subscript is used to indicate the moment or the position in the text. represents the acoustic vector information corresponding to the t-th moment; T represents the number of moments; represents the text vector information corresponding to the n-th phoneme; N represents the number of phonemes; score() represents a scoring function, where represents the transposed matrix of; exp() represents the exponential function; a t,n represents the weight coefficient corresponding to the t-th moment; c t represents the attention vector information corresponding to the t-th moment.

[0097] S310, perform speech recognition based on the fusion vector corresponding to each moment to obtain the predicted pronunciation text information corresponding to the speech information.

[0098] The predicted pronunciation text information is specifically the predicted phoneme sequence. After obtaining the fusion vectors corresponding to each moment, through the connection layer and the softmax function, the probabilities of each fusion vector being mapped to each target phoneme are output. Assuming the number of target phonemes is M, then M probabilities are output for each fusion vector, and the target phoneme corresponding to the maximum probability is taken as the predicted phoneme corresponding to the fusion vector. Based on this, the predicted phoneme sequence corresponding to the speech information can be obtained.

[0099] The calculation formula for the probability output by the model is as follows:

[0100]

[0101] Among them, y t′ represents the output probability corresponding to the t'-th moment; [*; *] represents concatenating two vectors together; W represents the weight, and b represents the bias.

[0102] In the above embodiments, during the speech recognition process of the model, the attention mechanism is used to fuse the speech information and the pronunciation text information, which can improve the recognition performance of the model, and then can more accurately detect the correct and incorrect pronunciations in the learner's pronunciation, making the scoring based on pronunciation quality more well-founded. When using the method of the above embodiments to detect English pronunciation error data (L2-arctic), the F-measure index can be increased to 56.08%, and there are obvious improvements in other various indexes. In addition, the end-to-end model construction also reduces the engineering quantity of data cleaning, model construction, model training, and subsequent operation and maintenance.

[0103] In one embodiment, as Figure 4 shown, a pronunciation error detection method is provided. Taking the case where this method is applied to the Figure 1 terminal as an example, it includes the following steps S402 to step S406.

[0104] S402, display the first page, and the first page includes target content.

[0105] The target content can be understood as the content that the user will perform pronunciation practice on, such as a word, phrase, or sentence. As Figure 5 shown, a schematic diagram of the first page in one embodiment is provided, where the target content is "afternoon".

[0106] S404, in response to the pronunciation operation on the target content on the first page, receive the speech information corresponding to the target content.

[0107] As Figure 5As shown, the first page also includes a function control "Click to start following reading" for triggering a pronunciation operation. After the user clicks this control, the terminal is in a state of monitoring voice information. At this time, when the user pronounces the target content, the terminal can receive the voice information corresponding to the target content through the microphone.

[0108] S406, when the pronunciation operation end condition is met, display the second page. The second page includes the pronunciation error detection result, and the content corresponding to the pronunciation error part in the target content is marked in the pronunciation error detection result.

[0109] It can be considered that the pronunciation operation end condition is met after a period of time (such as 3 seconds) after the terminal enters the state of monitoring voice information, or it can be considered that the pronunciation operation end condition is met when the terminal monitors that the user triggers the function control for ending the pronunciation operation.

[0110] When the pronunciation operation end condition is met, display the second page. The second page includes the pronunciation error detection result. The pronunciation error detection result is obtained by performing pronunciation error detection on the voice information according to the predicted pronunciation text information corresponding to the voice information and the standard pronunciation text information. The predicted pronunciation text information is obtained by performing speech recognition on the voice information and the standard pronunciation text information. The specific steps of speech recognition and pronunciation error detection can refer to the previous embodiments and will not be elaborated here.

[0111] As Figure 6 shown, a schematic diagram of the second page in an embodiment is provided. The second page includes the pronunciation error detection result, and the content corresponding to the pronunciation error part in the target content is marked in the pronunciation error detection result. Among them, the target content is "afternoon", and the pronunciation error part is "er". The pronunciation error part is highlighted by different color markings. For example, a red marking indicates a pronunciation error, that is, the user mispronounces this sound, and a green marking indicates a correct pronunciation.

[0112] In addition, when the terminal monitors a pronunciation tutoring operation triggered on the second page, display the third page. The third page includes tutoring information for the pronunciation error part. As Figure 6 shown, the user triggers the pronunciation tutoring operation by clicking "afternoon" among them, and the terminal displays the third page. As Figure 7 shown, a schematic diagram of the third page in an embodiment is provided. The third page includes the standard pronunciation of "afternoon" and the tutoring information for the pronunciation error part "er".

[0113] In the above embodiments, during the pronunciation practice process, the user can intuitively see the pronunciation error part and the corresponding tutoring information through the terminal, so that the user can timely understand their pronunciation problems and correct them, which helps to improve the user's learning efficiency.

[0114] In one embodiment, as Figure 8 shown, a pronunciation error detection method is provided, and this method is described by applying it to the interaction between a client (i.e., a terminal) and a server. When a user performs a pronunciation practice operation on the client, after the client receives the user's voice information, it transmits the voice information and its corresponding standard phoneme sequence to the server; the server recognizes the voice information to obtain a predicted phoneme sequence, then conducts pronunciation error detection by comparing the predicted phoneme sequence with the standard phoneme sequence, and transmits the pronunciation error detection result back to the client for error correction feedback; the user views the pronunciation error detection result and the corresponding pronunciation tutoring information on the client for the next practice.

[0115] Among them, the server side recognizes the voice information through a model to obtain a predicted phoneme sequence. Specifically, this model includes an audio encoder, a text encoder, and an attention encoder. The audio encoder consists of a CNN-BLSTM network and is used to encode the acoustic features of the voice information to obtain acoustic vector information. The text encoder consists of a Pre-process layer, a BLSTM layer, and a Linear layer and is used to encode each phoneme in the standard phoneme sequence to obtain text vector information. The attention encoder includes an Attention layer and a CTC layer and is used to fuse the acoustic vector information and the text vector information and recognize the predicted phoneme sequence.

[0116] In the above embodiment, in the end-to-end model, the voice information and the standard pronunciation text information are fused through an attention mechanism for speech recognition, which can detect the pronunciation errors of users more efficiently and accurately. The user can intuitively see the pronunciation error part and the corresponding pronunciation tutoring information through the terminal, so as to timely understand their own pronunciation problems and correct them, thereby improving the learning efficiency.

[0117] It should be understood that although the steps in each flowchart involved in the above embodiment are shown in sequence according to the indication of the arrows, these steps do not necessarily need to be executed in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in each flowchart involved in the above embodiment may include multiple steps or multiple stages. These steps or stages do not necessarily need to be executed at the same moment, but can be executed at different moments. The execution order of these steps or stages does not necessarily need to be sequential, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.

[0118] In one embodiment, as Figure 9As shown, a pronunciation error detection device 900 is provided. This device is applied to a server and can be implemented as a software module, a hardware module, or a combination of both as part of a computer device. The device specifically includes: an acquisition module 910, an identification module 920, and a detection module 930, where:

[0119] The acquisition module 910 is configured to acquire speech information and the corresponding standard pronunciation text information of the speech information.

[0120] The identification module 920 is configured to perform speech recognition based on the speech information and the standard pronunciation text information to obtain the predicted pronunciation text information corresponding to the speech information.

[0121] The detection module 930 is configured to detect the pronunciation error of the speech information based on the predicted pronunciation text information and the standard pronunciation text information.

[0122] In one embodiment, when the identification module 920 performs speech recognition based on the speech information and the standard pronunciation text information to obtain the predicted pronunciation text information corresponding to the speech information, it is specifically configured to perform speech recognition based on the speech information and the standard pronunciation text information through a trained speech recognition model to obtain the predicted pronunciation text information corresponding to the speech information. The training method of the speech recognition model includes: acquiring sample speech information and the corresponding sample standard pronunciation text information of the sample speech information; performing speech recognition based on the sample speech information and the sample standard pronunciation text information through the speech recognition model to be trained to obtain the sample predicted pronunciation text information corresponding to the sample speech information; when the training end condition is not met, adjusting the parameters of the speech recognition model to be trained based on the difference between the sample predicted pronunciation text information and the corresponding sample standard pronunciation text information until the training end condition is met to obtain the trained speech recognition model.

[0123] In one embodiment, when the identification module 920 performs speech recognition based on the speech information and the standard pronunciation text information to obtain the predicted pronunciation text information corresponding to the speech information, it is specifically configured to: preprocess the speech information to obtain each frame of speech segment, and extract the acoustic features of each frame of speech segment; encode the acoustic features of each frame of speech segment to obtain the acoustic vector information corresponding to each moment; encode each pronunciation unit in the standard pronunciation text information to obtain the text vector information corresponding to each pronunciation unit; fuse the acoustic vector information corresponding to each moment and the text vector information corresponding to each pronunciation unit to obtain the fusion vector corresponding to each moment; perform speech recognition based on the fusion vector corresponding to each moment to obtain the predicted pronunciation text information corresponding to the speech information.

[0124] In one embodiment, when the recognition module 920 preprocesses the voice information to obtain each frame of voice segment and extracts the acoustic features of each frame of voice segment, it is specifically configured to: perform pre-emphasis, framing, and windowing processing on the voice information to obtain each frame of voice segment; perform Fourier transform and Mel filtering on each frame of voice segment to obtain the acoustic features of each frame of voice segment.

[0125] In one embodiment, when the recognition module 920 fuses the acoustic vector information corresponding to each moment and the text vector information corresponding to each pronunciation unit to obtain the fusion vector corresponding to each moment, it is specifically configured to: obtain the attention vector information corresponding to each moment according to the acoustic vector information corresponding to each moment and the text vector information corresponding to each pronunciation unit; fuse the attention vector information corresponding to each moment and the acoustic vector information to obtain the fusion vector corresponding to each moment.

[0126] In one embodiment, the text vector information includes first text vector information and second text vector information, and the second text vector information is obtained by linearly processing the first text vector information; when the recognition module 920 obtains the attention vector information corresponding to each moment according to the acoustic vector information corresponding to each moment and the text vector information corresponding to each pronunciation unit, it is specifically configured to: obtain the weight coefficient corresponding to each moment according to the acoustic vector information corresponding to each moment and the second text vector information corresponding to each pronunciation unit; obtain the attention vector information corresponding to each moment according to the weight coefficient corresponding to each moment and the first text vector information corresponding to each pronunciation unit.

[0127] In one embodiment, the predicted pronunciation text information includes a predicted phoneme sequence, and the standard pronunciation text information includes a standard phoneme sequence; when the detection module 930 detects the pronunciation error of the voice information according to the predicted pronunciation text information and the standard pronunciation text information, it is specifically configured to: compare the predicted phoneme sequence with each phoneme in the standard phoneme sequence to obtain the pronunciation error detection result.

[0128] In one embodiment, as Figure 10 shown, a pronunciation error detection device 1000 is provided. The device can be a software module or a hardware module, or a combination of both to form a part of a computer device. The device specifically includes: a display module 1010 and a receiving module 1020, where:

[0129] The display module 1010 is configured to display a first page, and the first page includes target content.

[0130] The receiving module 1020 is configured to receive the voice information corresponding to the pronunciation operation in response to the pronunciation operation on the target content on the first page.

[0131] The display module 1010 is further configured to display a second page when the pronunciation operation end condition is met. The second page includes the pronunciation error detection result, and the content corresponding to the pronunciation error part in the target content is marked in the pronunciation error detection result. The pronunciation error detection result is obtained by performing pronunciation error detection on the voice information according to the predicted pronunciation text information and the standard pronunciation text information corresponding to the voice information. The predicted pronunciation text information is obtained by performing speech recognition on the voice information and the standard pronunciation text information.

[0132] For the specific limitations of the pronunciation error detection device, reference can be made to the limitations of the pronunciation error detection method in the above text, which will not be elaborated here. Each module in the above pronunciation error detection device can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor of the computer device in hardware form or be independent of it, or can be stored in the memory of the computer device in software form, so as to facilitate the processor to call and execute the operations corresponding to the above modules.

[0133] In one embodiment, a computer device is provided. The computer device can be a server, and its internal structure diagram can be as Figure 11 shown. The computer device includes a processor, a memory, and a network interface connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements a pronunciation error detection method.

[0134] In one embodiment, a computer device is provided. The computer device can be a terminal, and its internal structure diagram can be as Figure 12As shown in the figure. The computer device includes a processor, a memory, a communication interface, a display screen, and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The communication interface of the computer device is used to communicate with external terminals in a wired or wireless manner. The wireless manner can be implemented through WIFI, carrier networks, NFC (Near Field Communication), or other technologies. When the computer program is executed by the processor, it implements a pronunciation error detection method. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer covering the display screen, or a button, trackball, or touchpad provided on the housing of the computer device, or an external keyboard, touchpad, or mouse, etc.

[0135] Those skilled in the art can understand that Figure 11 or Figure 12 The structure shown in the figure is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.

[0136] In one embodiment, a computer device is further provided, including a memory and a processor. A computer program is stored in the memory. When the processor executes the computer program, the steps in the above method embodiments are implemented.

[0137] In one embodiment, a computer-readable storage medium is provided, storing a computer program. When the computer program is executed by the processor, the steps in the above method embodiments are implemented.

[0138] In one embodiment, a computer program product or a computer program is provided. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the steps in the above method embodiments.

[0139] It should be understood that the terms "first", "second", etc. in the above embodiments are only used for descriptive purposes, and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. In addition, in the description of this application, unless otherwise specified, the meaning of "a plurality" is at least two.

[0140] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above various methods. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided in the present application can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, or optical memory, etc. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.

[0141] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the various technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered to be within the scope described in this specification.

[0142] The above-described embodiments merely represent several implementation manners of the present application. Their descriptions are relatively specific and detailed, but they should not be construed as limiting the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application should be subject to the appended claims.

Claims

1. A method for detecting pronunciation errors, characterized in that, the method includes: obtaining speech information and standard pronunciation text information corresponding to the speech information; encoding the acoustic features of each frame of speech segment of the speech information to obtain acoustic vector information corresponding to each moment; encoding each pronunciation unit in the standard pronunciation text information to obtain text vector information corresponding to each pronunciation unit; fusing the acoustic vector information corresponding to each moment and the text vector information corresponding to each pronunciation unit to obtain a fusion vector corresponding to each moment; performing speech recognition based on the fusion vector corresponding to each moment to obtain predicted pronunciation text information corresponding to the speech information; detecting pronunciation errors of the speech information according to the predicted pronunciation text information and the standard pronunciation text information.

2. The method according to claim 1, characterized in that, the predicted pronunciation text information is obtained by performing speech recognition through a trained speech recognition model, and the training steps of the speech recognition model include: obtaining sample speech information and sample standard pronunciation text information corresponding to the sample speech information; encoding the sample acoustic features of each frame of sample speech segment of the sample speech information through a speech recognition model to be trained to obtain sample acoustic vector information corresponding to each moment; encoding each sample pronunciation unit in the sample standard pronunciation text information to obtain sample text vector information corresponding to each sample pronunciation unit; fusing the sample acoustic vector information corresponding to each moment and the sample text vector information corresponding to each sample pronunciation unit to obtain a sample fusion vector corresponding to each moment; performing speech recognition based on the sample fusion vector corresponding to each moment to obtain sample predicted pronunciation text information corresponding to the sample speech information; when the training end condition is not met, adjusting the parameters of the speech recognition model to be trained based on the difference between the sample predicted pronunciation text information and the corresponding sample standard pronunciation text information until the training end condition is met, and obtaining a trained speech recognition model.

3. The method according to claim 1, characterized in that, encoding the acoustic features of each frame of speech segment of the speech information to obtain acoustic vector information corresponding to each moment includes: preprocessing the speech information to obtain each frame of speech segment, and extracting the acoustic features of each frame of speech segment; encoding the acoustic features of each frame of speech segment to obtain acoustic vector information corresponding to each moment.

4. The method according to claim 3, characterized in that, preprocessing the speech information to obtain each frame of speech segment and extracting the acoustic features of each frame of speech segment includes: performing pre-emphasis, framing and windowing processing on the speech information to obtain each frame of speech segment; performing Fourier transform and Mel filtering on each frame of speech segment to obtain the acoustic features of each frame of speech segment.

5. The method according to claim 3, characterized in that, fusing the acoustic vector information corresponding to each moment and the text vector information corresponding to each pronunciation unit to obtain a fusion vector corresponding to each moment includes: Obtain the attention vector information corresponding to each moment according to the acoustic vector information corresponding to each moment and the text vector information corresponding to each pronunciation unit; Fuse the attention vector information corresponding to each moment and the acoustic vector information to obtain the fused vector corresponding to each moment.

6. The method according to claim 5, wherein, the text vector information includes first text vector information and second text vector information, and the second text vector information is obtained by linearly processing the first text vector information; Obtaining the attention vector information corresponding to each moment according to the acoustic vector information corresponding to each moment and the text vector information corresponding to each pronunciation unit includes: Obtain the weight coefficient corresponding to each moment according to the acoustic vector information corresponding to each moment and the second text vector information corresponding to each pronunciation unit; Obtain the attention vector information corresponding to each moment according to the weight coefficient corresponding to each moment and the first text vector information corresponding to each pronunciation unit.

7. The method according to any one of claims 1 to 6, wherein, the predicted pronunciation text information includes a predicted phoneme sequence, and the standard pronunciation text information includes a standard phoneme sequence; Detecting the pronunciation error of the speech information according to the predicted pronunciation text information and the standard pronunciation text information includes: Compare the predicted phoneme sequence with each phoneme in the standard phoneme sequence to obtain a pronunciation error detection result.

8. A pronunciation error detection method, wherein, the method includes: Display a first page, and the first page includes target content; In response to a pronunciation operation on the target content in the first page, receive the speech information corresponding to the target content; When the pronunciation operation end condition is satisfied, display a second page, and the second page includes a pronunciation error detection result, and mark the content corresponding to the pronunciation error part in the target content in the pronunciation error detection result; Wherein, the pronunciation error detection result is obtained by detecting the pronunciation error of the speech information according to the predicted pronunciation text information and the standard pronunciation text information corresponding to the speech information, the predicted pronunciation text information is obtained by performing speech recognition based on the fused vector corresponding to each moment, the fused vector corresponding to each moment is obtained by fusing the acoustic vector information corresponding to each moment and the text vector information corresponding to each pronunciation unit, the text vector information corresponding to each pronunciation unit is obtained by encoding each pronunciation unit in the standard pronunciation text information, and the acoustic vector information corresponding to each moment is obtained by encoding the acoustic features of each frame of speech segment of the speech information.

9. A pronunciation error detection device, wherein, the device includes: An acquisition module for acquiring speech information and the standard pronunciation text information corresponding to the speech information; An encoding module for encoding the acoustic features of each frame of speech segment of the speech information to obtain the acoustic vector information corresponding to each moment; encoding each pronunciation unit in the standard pronunciation text information to obtain the text vector information corresponding to each pronunciation unit; A fusion module, configured to fuse the acoustic vector information corresponding to each moment and the text vector information corresponding to each pronunciation unit to obtain the fusion vector corresponding to each moment; An identification module, configured to perform speech recognition based on the fusion vector corresponding to each moment to obtain the predicted pronunciation text information corresponding to the speech information; A detection module, configured to detect the pronunciation error of the speech information according to the predicted pronunciation text information and the standard pronunciation text information.

10. The pronunciation error detection device according to claim 9, wherein, the predicted pronunciation text information is obtained by performing speech recognition through a trained speech recognition model, and the device further includes a training module, configured to obtain sample speech information and the sample standard pronunciation text information corresponding to the sample speech information; encode the sample acoustic features of each frame of sample speech segment of the sample speech information through the speech recognition model to be trained to obtain the sample acoustic vector information corresponding to each moment; encode each sample pronunciation unit in the sample standard pronunciation text information to obtain the sample text vector information corresponding to each sample pronunciation unit; fuse the sample acoustic vector information corresponding to each moment and the sample text vector information corresponding to each sample pronunciation unit to obtain the sample fusion vector corresponding to each moment; perform speech recognition based on the sample fusion vector corresponding to each moment to obtain the sample predicted pronunciation text information corresponding to the sample speech information; when the training end condition is not satisfied, adjust the parameters of the speech recognition model to be trained based on the difference between the sample predicted pronunciation text information and the corresponding sample standard pronunciation text information until the training end condition is satisfied, and obtain a trained speech recognition model.

11. The pronunciation error detection device according to claim 9, wherein, the encoding module is further configured to preprocess the speech information to obtain each frame of speech segment, and extract the acoustic features of each frame of speech segment; encode the acoustic features of each frame of speech segment to obtain the acoustic vector information corresponding to each moment.

12. The pronunciation error detection device according to claim 11, wherein, the encoding module is further configured to perform pre-emphasis, framing and windowing processing on the speech information to obtain each frame of speech segment; perform Fourier transform and Mel filtering on each frame of speech segment to obtain the acoustic features of each frame of speech segment.

13. The pronunciation error detection device according to claim 9, wherein, the fusion module is further configured to obtain the attention vector information corresponding to each moment according to the acoustic vector information corresponding to each moment and the text vector information corresponding to each pronunciation unit; fuse the attention vector information corresponding to each moment and the acoustic vector information to obtain the fusion vector corresponding to each moment.

14. The pronunciation error detection device according to claim 13, wherein, the text vector information includes first text vector information and second text vector information, and the second text vector information is obtained by performing linear processing on the first text vector information; The fusion module is further configured to obtain weight coefficients corresponding to each moment according to the acoustic vector information corresponding to each moment and the second text vector information corresponding to each pronunciation unit; and obtain attention vector information corresponding to each moment according to the weight coefficients corresponding to each moment and the first text vector information corresponding to each pronunciation unit.

15. The pronunciation error detection device according to any one of claims 9 to 14, wherein, the predicted pronunciation text information includes a predicted phoneme sequence, and the standard pronunciation text information includes a standard phoneme sequence; the detection module is further configured to compare each phoneme in the predicted phoneme sequence with each phoneme in the standard phoneme sequence to obtain a pronunciation error detection result.

16. A pronunciation error detection device, wherein, the device includes: a display module, configured to display a first page, and the first page includes target content; a receiving module, configured to receive voice information corresponding to the target content in response to a pronunciation operation on the target content on the first page; the display module is further configured to display a second page when a pronunciation operation end condition is met, the second page includes a pronunciation error detection result, and content corresponding to the pronunciation error part in the target content is marked in the pronunciation error detection result; wherein, the pronunciation error detection result is obtained by performing pronunciation error detection on the voice information according to the predicted pronunciation text information and the standard pronunciation text information corresponding to the voice information, the predicted pronunciation text information is obtained by performing speech recognition on the voice information and the standard pronunciation text information, the predicted pronunciation text information is obtained by performing speech recognition based on the fusion vectors corresponding to each moment, the fusion vectors corresponding to each moment are obtained by fusing the acoustic vector information corresponding to each moment and the text vector information corresponding to each pronunciation unit, the text vector information corresponding to each pronunciation unit is obtained by encoding each pronunciation unit in the standard pronunciation text information, and the acoustic vector information corresponding to each moment is obtained by encoding the acoustic features of each frame of speech segment of the voice information.

17. A computer device, including a memory and a processor, the memory stores a computer program, wherein, when the processor executes the computer program, the method according to any one of claims 1 to 8 is implemented.

18. A computer-readable storage medium, on which a computer program is stored, wherein, when the computer program is executed by a processor, the method according to any one of claims 1 to 8 is implemented.

19. A computer program product, including computer instructions, wherein, when the computer instructions are executed by a processor, the method according to any one of claims 1 to 8 is implemented.

Citation Information

Patent Citations

  • Information processing method and device, medium, and computing equipment

    CN108682437A

  • Speech identification method and device, electronic equipment and storage medium

    CN110797016A

  • Voice processing and voice evaluation method and device thereof, computer equipment and storage medium

    CN111402895A