Phoneme-level pronunciation correction method, device, equipment and storage medium

By fusing the audio features of the audio frame with the phoneme features of the following text, the error read probability of phoneme data is calculated, and the problem of cumbersome and low efficiency in the prior art is solved, and more efficient and accurate phoneme error judgment is achieved.

CN114360505BActive Publication Date: 2025-06-20TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202111424491.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-26
Publication Date
2025-06-20
Estimated Expiration
2041-11-26

AI Technical Summary

Technical Problem

In oral evaluation, the prior art requires relying on external standard pronunciation data for phoneme error judgment, resulting in cumbersome evaluation process and low processing efficiency.

Method used

By obtaining the following audio data corresponding to the following text, extracting the audio characteristics and phoneme characteristics of the audio frame, and fusing them into fusion characteristics, the error probability of the phoneme data is calculated to determine the error phoneme.

Benefits of technology

The process of phoneme error judgment is simplified, the processing efficiency and accuracy of phoneme error judgment is improved, and the dependence on standard pronunciation data is avoided.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114360505B_ABST
    Figure CN114360505B_ABST
Patent Text Reader

Abstract

The present application discloses a phoneme-level pronunciation error correction method, device, equipment and storage medium, which belongs to the field of computer and Internet technology. The method includes: obtaining the follow-up audio data corresponding to the follow-up text; obtaining the audio features corresponding to each audio frame in the follow-up audio data, and obtaining the phoneme features of the phonemes contained in the follow-up text; fusing the audio features corresponding to each audio frame with the phoneme features to obtain the fusion features corresponding to each audio frame; according to the fusion features corresponding to each audio frame, obtaining at least one phoneme data contained in the follow-up audio data, and the misreading probability of each phoneme data; based on the misreading probability of each phoneme data, determining the misread phonemes in the follow-up audio data. In the present application, there is no need to generate standard pronunciation data based on the phonemes contained in the follow-up text, which simplifies the phoneme error judgment process and improves the processing efficiency of phoneme error judgment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer and Internet technologies, and particularly relates to a phoneme-level pronunciation error correction method, apparatus, device, and storage medium. Background Art

[0002] Currently, in oral evaluation, phoneme error judgment can be performed on the voice data read by the user at the phoneme level.

[0003] In the related art, in oral evaluation, after obtaining the voice data, the pronunciation segments in the voice data are compared with the pronunciation segments in the standard pronunciation data. If there are some different phonemes in the comparison, it is determined that the phoneme in the voice data is mispronounced.

[0004] However, in the above-mentioned related art, the phoneme error judgment of the voice data depends on external standard pronunciation data, and the standard pronunciation data needs to be generated before the oral evaluation, and the evaluation process is cumbersome. Summary of the Invention

[0005] Embodiments of this application provide a phoneme-level pronunciation error correction method, apparatus, device, and storage medium, which simplify the phoneme error judgment process and improve the processing efficiency of phoneme error judgment. The technical solutions are as follows.

[0006] According to one aspect of the embodiments of this application, a phoneme-level pronunciation error correction method is provided. The method includes the following steps:

[0007] Obtain the follow-up audio data corresponding to the follow-up text;

[0008] Obtain the audio features respectively corresponding to each audio frame in the follow-up audio data, and obtain the phoneme features of the phonemes included in the follow-up text;

[0009] Fuse the audio features respectively corresponding to each audio frame with the phoneme features respectively to obtain the fusion features respectively corresponding to each audio frame;

[0010] According to the fusion features respectively corresponding to each audio frame, obtain at least one phoneme data included in the follow-up audio data, and the mispronunciation probability of each phoneme data;

[0011] Based on the mispronunciation probabilities of each phoneme data, determine the mispronounced phonemes in the follow-up audio data.

[0012] According to one aspect of the embodiments of this application, a training method for a phoneme detection model is provided. The method includes the following steps:

[0013] Obtain the training samples of the phoneme detection model. The training samples include sample follow-up texts and sample follow-up audio data corresponding to the sample follow-up texts;

[0014] Obtain the audio features corresponding to each sample audio frame in the sample follow - up audio data, and obtain the phoneme features of the phonemes included in the sample follow - up text;

[0015] Fuse the audio features corresponding to each of the sample audio frames with the phoneme features respectively to obtain the fusion features corresponding to each audio frame;

[0016] According to the fusion features corresponding to each of the sample audio frames, obtain at least one phoneme data included in the sample follow - up audio data, and the mispronunciation probability of each phoneme data;

[0017] Based on the mispronunciation probabilities of each phoneme data, determine the phoneme detection result of the sample follow - up audio data;

[0018] Calculate the training loss of the phoneme detection model according to the phoneme detection result and the label of the training sample, and adjust the parameters of the phoneme detection model according to the training loss.

[0019] According to one aspect of the embodiments of the present application, there is provided a phoneme - level pronunciation correction device, and the device includes the following modules:

[0020] An audio acquisition module, configured to acquire follow - up audio data corresponding to a follow - up text;

[0021] A feature acquisition module, configured to acquire the audio features corresponding to each audio frame in the follow - up audio data, and acquire the phoneme features of the phonemes included in the follow - up text;

[0022] A feature fusion module, configured to fuse the audio features corresponding to each of the audio frames with the phoneme features respectively to obtain the fusion features corresponding to each of the audio frames;

[0023] A probability acquisition module, configured to acquire at least one phoneme data included in the follow - up audio data, and the mispronunciation probability of each phoneme data, according to the fusion features corresponding to each of the audio frames;

[0024] A phoneme misjudgment module, configured to determine the mispronounced phonemes in the follow - up audio data based on the mispronunciation probabilities of each phoneme data.

[0025] According to one aspect of the embodiments of the present application, there is provided a training device for a phoneme detection model, and the device includes the following modules:

[0026] A sample acquisition module, configured to acquire the training samples of the phoneme detection model, where the training samples include sample follow - up texts and sample follow - up audio data corresponding to the sample follow - up texts;

[0027] A feature determination module, configured to obtain audio features corresponding to each sample audio frame in the sample follow - up audio data, and obtain phoneme features of phonemes included in the sample follow - up text;

[0028] A feature processing module, configured to fuse the audio features corresponding to each of the sample audio frames with the phoneme features respectively to obtain fused features corresponding to each audio frame;

[0029] A probability determination module, configured to obtain at least one phoneme data included in the sample follow - up audio data and the mispronunciation probability of each phoneme data according to the fused features corresponding to each of the sample audio frames;

[0030] A result determination module, configured to determine a phoneme detection result of the sample follow - up audio data based on the mispronunciation probabilities of each phoneme data;

[0031] A loss acquisition module, configured to calculate a training loss of the phoneme detection model according to the phoneme detection result and the label of the training sample;

[0032] A parameter adjustment module, configured to adjust parameters of the phoneme detection model according to the training loss.

[0033] According to one aspect of the embodiments of the present application, embodiments of the present application provide a computer device, which includes a processor and a memory. At least one instruction, at least one program, a code set or an instruction set is stored in the memory, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by the processor to implement the above - mentioned phoneme - level pronunciation correction method or to implement the above - mentioned training method of the phoneme detection model.

[0034] According to one aspect of the embodiments of the present application, embodiments of the present application provide a computer - readable storage medium, in which at least one instruction, at least one program, a code set or an instruction set is stored, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by a processor to implement the above - mentioned phoneme - level pronunciation correction method or to implement the above - mentioned training method of the phoneme detection model.

[0035] According to one aspect of the embodiments of the present application, there is provided a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer - readable storage medium. A processor of a computer device reads the computer instructions from the computer - readable storage medium, and the processor executes the computer instructions, so that the computer device executes the above - mentioned phoneme - level pronunciation correction method or implements the above - mentioned training method of the phoneme detection model.

[0036] The technical solution provided by the embodiment of the present application can bring the following beneficial effects:

[0037] By fusing the audio features of audio frames and the phoneme features of phonemes included in the follow-up text, the fusion features corresponding to each audio frame are obtained, so that the fusion features can represent the audio features of the audio frames and also the phoneme features. Further, according to the fusion features, the phoneme data included in the follow-up audio data and the mispronunciation probability of the phoneme data are obtained. In the process of phoneme misjudgment for the follow-up audio data, in addition to processing the audio features of the follow-up audio data, the phoneme features of the phonemes included in the follow-up text are also considered, improving the accuracy of phoneme misjudgment; moreover, by fusing the phoneme features of the phonemes included in the follow-up text into the audio features, when phoneme misjudgment can be performed through the fusion features subsequently, it is not necessary to generate standard pronunciation data based on the phonemes included in the follow-up text, simplifying the phoneme misjudgment process and improving the processing efficiency of phoneme misjudgment. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0039] Figure 1 It is a schematic diagram of a phoneme-level pronunciation correction system provided by an embodiment of the present application;

[0040] Figure 2 Exemplarily shows a schematic diagram of a phoneme-level pronunciation correction system;

[0041] Figure 3 It is a flowchart of a phoneme-level pronunciation correction method provided by an embodiment of the present application;

[0042] Figure 4 Exemplarily shows a schematic diagram of a follow-up text display method;

[0043] Figure 5 Exemplarily shows a schematic diagram of a method for time-sequentially combining phoneme recognition results;

[0044] Figure 6 It is a flowchart of a training method for a phoneme detection model provided by an embodiment of the present application;

[0045] Figure 7 Exemplarily shows a schematic diagram of a phoneme detection model;

[0046] Figure 8 It is a block diagram of a phoneme-level pronunciation correction device provided by an embodiment of the present application;

[0047] Figure 9 It is a block diagram of a phoneme-level pronunciation error correction device provided by another embodiment of the present application;

[0048] Figure 10 It is a block diagram of a training device for a phoneme detection model provided by an embodiment of the present application;

[0049] Figure 11 It is a block diagram of the structure of a computer device provided by an embodiment of the present application. Detailed implementation manners

[0050] To make the objectives, technical solutions, and advantages of the present application clearer, the following will further describe the embodiments of the present application in detail with reference to the accompanying drawings.

[0051] Please refer to Figure 1 , which shows a schematic diagram of a phoneme-level pronunciation error correction system provided by an embodiment of the present application. The phoneme-level pronunciation error correction system may include: a terminal 10 and a server 20.

[0052] The terminal 10 includes, but is not limited to, electronic devices such as mobile phones, tablet computers, game consoles, e-book readers, multimedia playback devices, wearable devices, PCs (Personal Computers), intelligent voice interaction devices, intelligent home appliances, and vehicle terminals. The terminal 10 may include a client of an application program. Optionally, the application program may be any application program having an audio data collection function and a text display function, and the embodiments of the present application do not make any limitation thereto. Among them, the application program may be an application program that needs to be downloaded and installed, or an application program that can be used immediately upon clicking, and the embodiments of the present application do not make any limitation thereto.

[0053] The server 20 is used to provide background services for the terminal 10. The server 20 may be a single server, a server cluster composed of multiple servers, or a cloud computing service center. Optionally, the server 20 may be the background server of the above client. In an exemplary embodiment, the server 20 provides background services for multiple terminals 10.

[0054] The above terminal 10 and the above server 20 communicate with each other through a network 30.

[0055] Optionally, in the embodiments of the present application, the server 20 provides a background service for phoneme error judgment for the terminal 10. Exemplarily, as Figure 2As shown, the terminal 10 displays the text for shadowing in the user interface and collects the shadowing audio data of the user for the text for shadowing. Further, the terminal 10 sends the text for shadowing and the shadowing audio data to the server 20. The server 20 obtains the audio features corresponding to each audio frame based on the shadowing audio data, and obtains the phoneme features of the phonemes included in the text for shadowing based on the text for shadowing. After that, the audio features and the phoneme features are subjected to feature fusion, phoneme classification, and phoneme error judgment to determine the mispronounced phonemes in the shadowing audio data, and the server 20 sends the mispronounced phonemes in the shadowing audio data to the terminal 10. Correspondingly, the terminal 10 displays the text for shadowing marked with the mispronounced phonemes in the user interface.

[0056] The above phonemes are the smallest speech units divided according to the natural properties of speech. From an acoustic perspective, phonemes are the smallest speech units divided from the perspective of sound quality. Exemplarily, [ma] includes two phonemes, [m] and [a]. From a physiological perspective, one pronunciation action forms one phoneme. The sounds produced by the same pronunciation action are the same phoneme, and the sounds produced by different pronunciation actions are different phonemes. Exemplarily, in [ma - mi], the two [m] pronunciation actions are the same, so they are the same phoneme, while the pronunciation actions of [a] and [i] are different, so they are different phonemes.

[0057] It should be noted that the above Figure 2 introduction is only exemplary and explanatory. In the exemplary embodiment, the terminal 10 itself can perform each of the above Figure 2 operations on the shadowing audio data and the text for shadowing to reduce the signal transmission between the terminal 10 and the server 20. Of course, in other possible exemplary embodiments, in order to reduce the load of the server 20, multiple servers 20 can also perform each of the above Figure 2 operations on the shadowing audio data and the text for shadowing. The embodiments of the present application do not limit this.

[0058] Next, the technical solution of the present application will be introduced and described in conjunction with several embodiments.

[0059] Please refer to Figure 3 , which shows a flowchart of a phoneme - level pronunciation error correction method provided by an embodiment of the present application. This method can be applied to Figure 1 the server 20 and / or the terminal 10 of the phoneme - level pronunciation error correction system shown. For example, the execution entity of each step can be the client of the application installed in the server 20 or the terminal 10 (hereinafter, the execution entity of the phoneme - level pronunciation error correction method is collectively referred to as "computer device"). This method can include at least one of the following steps (301 - 305):

[0060] Step 301, obtaining the shadowing audio data corresponding to the text for shadowing.

[0061] The text for shadowing is the text read aloud by the user. Exemplarily, in a spoken language evaluation scenario, the text used to evaluate the user's spoken language is the text for shadowing. Optionally, in different usage scenarios, the content included in the text for shadowing is different. For example, in an English spoken language test, the text for shadowing includes English text, and in a Chinese spoken language test, the text for shadowing includes Chinese text.

[0062] The shadowing audio data refers to the sound data generated by the user when reading the above text for shadowing. In the embodiments of the present application, when the user reads the above text for shadowing, the computer device acquires the shadowing audio data corresponding to the text for shadowing.

[0063] In a possible implementation manner, when acquiring the shadowing audio data, after determining that the user's reading of the text for shadowing ends, the computer device acquires the shadowing audio data corresponding to the text for shadowing. At this time, the text for shadowing corresponding to the shadowing audio data is the complete text for shadowing presented to the user. In this case, due to the integrity of the shadowing audio data, in the subsequent phoneme judgment process, it is possible to better perform phoneme error judgment based on the context from the complete shadowing audio data, improving the accuracy of phoneme error judgment.

[0064] In another possible implementation manner, when acquiring the shadowing audio data, during the user's reading of the text for shadowing, the computer device acquires the shadowing audio data at a certain interval. At this time, the text for shadowing corresponding to the shadowing audio data is a part of the complete text for shadowing presented to the user. Optionally, when the computer device acquires the shadowing audio data, it simultaneously determines the text for shadowing corresponding to the shadowing audio data. In this case, the computer device acquires the shadowing audio data corresponding to the complete text for shadowing in segments, reducing the amount of data included in the audio data to be processed by the computer device and reducing the load on the computer device. Optionally, the above interval is at least one of the following situations: time interval, reading pause, interval determined according to the amount of data included in the reading text, and interval determined according to the amount of data included in the shadowing audio data. Exemplarily, the above interval is a time interval, and the computer device acquires the shadowing audio data every 2 s. Of course, in specific applications, the time interval can be any value; the above time interval is a reading pause, and when the computer device detects a reading pause of the user during the reading process, it acquires the shadowing audio data between the previous pause and the current pause; the above interval is an interval determined according to the amount of data included in the reading text, and when it is determined that the text read by the user reaches a certain amount of data, the computer device acquires the shadowing audio data corresponding to the text read; the above interval is an interval determined according to the amount of data included in the shadowing audio data, and when it is determined that the shadowing audio data reaches a certain amount of data, the computer device acquires the shadowing audio data. In this case, the amount of data included in the shadowing audio data acquired by the computer device each time is the same.

[0065] In yet another possible implementation, before obtaining the shadowing audio data, the computer device obtains the amount of data included in the complete shadowing text. When the amount of data included in the complete shadowing text is less than the target value, the computer device directly obtains the shadowing audio data corresponding to the complete shadowing text; when the amount of data included in the complete shadowing text is greater than or equal to the target value, the computer device obtains the shadowing audio data corresponding to the complete shadowing text at a certain interval.

[0066] Step 302: Obtain the audio features respectively corresponding to each audio frame in the shadowing audio data, and obtain the phoneme features of the phonemes included in the shadowing text.

[0067] In the embodiments of the present application, after the computer device obtains the shadowing audio data corresponding to the above-mentioned shadowing text, it obtains the audio features respectively corresponding to each audio frame in the shadowing audio data, and obtains the phoneme features of the phonemes included in the shadowing text. Among them, the above-mentioned audio features can be referred to as audio feature vectors, and the above-mentioned phoneme features can be referred to as phoneme feature vectors.

[0068] Optionally, in the embodiments of the present application, after the computer device obtains the above-mentioned shadowing audio data, it performs frame splitting on the shadowing audio data to obtain at least one audio frame of the shadowing audio data, and then performs feature extraction on each audio frame respectively to obtain the audio features respectively corresponding to each audio frame. Optionally, the computer device uses a pre-trained audio feature extraction network to perform feature extraction on each audio frame respectively to obtain the audio features respectively corresponding to each audio frame. Among them, the pre-trained audio feature extraction network refers to a feature extraction network for audio data that is pre-trained. Exemplarily, the audio feature extraction network can be Wav2vector. Optionally, during the training process, a large number of unlabeled tasks can be used to train the audio feature extraction network based on the contrast loss. It should be noted that in the embodiments of the present application, the above-mentioned audio feature extraction network can also be referred to as an audio encoder.

[0069] Optionally, in the embodiments of the present application, after the computer device obtains the above-mentioned shadowing text, it extracts features of the phonemes included in the shadowing text to obtain phoneme features, where the phoneme features include phoneme feature representations corresponding to each phoneme included in the shadowing text respectively. Optionally, the computer device uses a pre-trained phoneme feature extraction network to extract features of the phonemes included in the shadowing text to obtain phoneme features. The pre-trained phoneme feature extraction network refers to a feature extraction network for phonemes obtained by pre-training. Optionally, each phoneme is represented by a unique vector, that is, in the embodiments of the present application, the phoneme feature representations corresponding to different phonemes are different. It should be noted that, in the embodiments of the present application, the above-mentioned phoneme feature extraction network may also be referred to as a phoneme encoder.

[0070] Step 303: Fuse the audio features corresponding to each audio frame with the phoneme features respectively to obtain the fusion features corresponding to each audio frame respectively.

[0071] In the embodiments of the present application, after the computer device obtains the above-mentioned audio features and phoneme features, it fuses the audio features corresponding to each audio frame with the audio features respectively to obtain the fusion features corresponding to each audio frame respectively.

[0072] Step 304: Based on the fusion features corresponding to each audio frame, obtain at least one phoneme data included in the shadowing audio data and the mispronunciation probability of each phoneme data.

[0073] In the embodiments of the present application, after the computer device obtains the above-mentioned fusion features, based on the fusion features corresponding to each audio frame, it obtains at least one phoneme data included in the shadowing audio data and the mispronunciation probability of each audio data. Among them, the phoneme data may also be referred to as a phoneme, and the mispronunciation probability is used to indicate the pronunciation standard degree of the phoneme data in the shadowing audio data.

[0074] Step 305: Determine the mispronounced phonemes in the shadowing audio data based on the mispronunciation probabilities of each phoneme data.

[0075] In the embodiments of the present application, after the computer device obtains the mispronunciation probability, it determines the mispronounced phonemes in the shadowing audio data based on the mispronunciation probabilities of each phoneme data. Optionally, when the computer device determines the mispronounced phonemes, it determines the phoneme data whose mispronunciation probability meets the condition as the first mispronounced phonemes in the shadowing audio data. Among them, the above-mentioned condition refers to the judgment condition for mispronounced phonemes, which can be flexibly set and adjusted according to the actual situation. Exemplarily, if the mispronunciation probability is directly proportional to the above-mentioned pronunciation standard degree, the above-mentioned condition is less than a certain value; if the mispronunciation probability is inversely proportional to the above-mentioned pronunciation standard degree, the above-mentioned condition is greater than a certain value.

[0076] The above-mentioned first mispronounced phoneme refers to the phoneme with inaccurate pronunciation in the follow-up audio data. Optionally, in the embodiments of the present application, the mispronounced phonemes include the first mispronounced phoneme and the second mispronounced phoneme. Among them, the second mispronounced phoneme refers to the incorrect phoneme in the follow-up audio data relative to the follow-up text. Exemplarily, if the phonemes in the follow-up text are "A (first tone) B (second tone) C (third tone)", and the audio data in the follow-up audio data is "A (first tone) D (second tone) C (fourth tone)", then "D" is the second mispronounced phoneme in the follow-up audio data, and "D" is the mispronounced phoneme relative to "B", and "C" is the first mispronounced phoneme in the follow-up audio data.

[0077] Optionally, in the embodiments of the present application, after the computer device determines the phoneme data included in the follow-up audio data, it compares the phoneme data included in the follow-up audio data with the phonemes included in the follow-up text. If there is phoneme data in the phoneme data included in the follow-up audio data that does not match the phonemes included in the follow-up text, then the unmatched phoneme data is determined as the second mispronounced phoneme in the follow-up audio data.

[0078] Optionally, in the embodiments of the present application, after obtaining the mispronounced phonemes in the follow-up audio data, the mispronounced phonemes can be marked in the follow-up text, and the marked follow-up text can be displayed to the user. Exemplarily, as Figure 4 shown, the follow-up text 41 and the "Start Reading" button 42 are displayed in the user interface. After the user clicks the "Start Reading" button 42, the "Start Reading" button 42 changes to the "End Reading" button 43. At the same time, the computer device obtains the follow-up audio data corresponding to the follow-up text 41 and determines the mispronounced phonemes from the follow-up audio data; then, after the user clicks the "End Reading" button 43, the marked follow-up text 44 is displayed in the user interface, and the marked follow-up text 44 includes the mark 45 for the mispronounced phoneme. It should be noted that, in the embodiments of the present application, if the mispronounced phoneme is the second mispronounced phoneme, the correct phoneme corresponding to the second mispronounced phoneme needs to be provided to facilitate the marking of the follow-up text.

[0079] In summary, in the technical solution provided by the embodiments of the present application, through the fusion of the audio features of audio frames and the phoneme features of the phonemes included in the shadowing text, the fusion features corresponding to each audio frame are obtained, so that the fusion features can represent the phoneme features while representing the audio features of the audio frames. Further, the phoneme data and the mispronunciation probability of the phoneme data included in the shadowing audio data are obtained according to the fusion features. In the process of phoneme misjudgment for the shadowing audio data, in addition to processing the audio features of the shadowing audio data, the phoneme features of the phonemes included in the shadowing text are also considered, improving the accuracy of phoneme misjudgment; moreover, by fusing the phoneme features of the phonemes included in the shadowing text into the audio features, when phoneme misjudgment can be performed through the fusion features subsequently, there is no need to generate standard pronunciation data based on the phonemes included in the shadowing text, simplifying the phoneme misjudgment process and improving the processing efficiency of phoneme misjudgment.

[0080] Next, the method for obtaining the fusion features will be introduced.

[0081] In an exemplary embodiment, step 303 above includes the following steps:

[0082] 1. For the target audio frame in each audio frame, generate the to-be-concatenated feature corresponding to the target audio frame according to the phoneme feature and the audio feature corresponding to the target audio frame.

[0083] Taking the target audio frame in each of the above audio frames as an example, in the embodiments of the present application, after the computer device obtains the audio feature corresponding to the target audio frame, it generates the to-be-concatenated feature corresponding to the target audio frame according to the phoneme feature and the audio feature corresponding to the target audio frame. Among them, the to-be-concatenated feature is the feature obtained by performing feature extraction on the audio feature corresponding to the target audio frame based on the phoneme feature.

[0084] Optionally, in the embodiments of the present application, feature fusion of the audio feature and the phoneme feature is performed through an attention mechanism. After the computer device obtains the audio feature corresponding to the target audio frame, it determines the audio feature corresponding to the target audio frame as the query vector, and determines the above phoneme feature as the key vector and the value vector. Further, based on the attention mechanism, the query vector, the key vector, and the value vector are processed to generate the to-be-concatenated feature corresponding to the target audio frame.

[0085] 2. Concatenate the audio feature corresponding to the target audio frame and the to-be-concatenated feature corresponding to the target audio frame to obtain the fusion feature corresponding to the target audio frame.

[0086] In an embodiment of the present application, after the computer device obtains the above-mentioned features to be spliced, it splices the audio features corresponding to the target audio frame and the features to be spliced corresponding to the target audio frame to obtain the fused features corresponding to the target audio frame. Exemplarily, if the feature to be spliced is an a-dimensional feature vector and the audio feature corresponding to the target audio frame is a b-dimensional feature vector, the fused feature obtained after splicing is an (a + b)-dimensional feature vector.

[0087] Exemplarily, assume that the audio feature corresponding to the target audio frame is The phoneme feature is H phone , then the fused feature is:

[0088]

[0089] The relationships among the query vector Q, the key vector K, and the value vector V in the attention mechanism are shown by the following two formulas:

[0090] Attention(Q, K, V) = AttentionScore(Q, K) * V;

[0091]

[0092] where d is the network dimension corresponding to the attention mechanism.

[0093] Next, the acquisition methods of phoneme data and mispronunciation probability are introduced.

[0094] In an exemplary embodiment, step 304 above includes the following steps:

[0095] 1. According to the fused features respectively corresponding to each audio frame, perform phoneme classification and phoneme error judgment on the shadowing audio data to obtain the phoneme recognition results respectively corresponding to each audio frame.

[0096] In an embodiment of the present application, after the computer device obtains the above-mentioned fused features, according to the fused features respectively corresponding to each audio frame, it performs phoneme classification and phoneme error judgment on the shadowing audio data to obtain the phoneme recognition results respectively corresponding to each audio frame. Among them, the phoneme recognition result includes the single-frame phoneme data corresponding to the audio frame and the mispronunciation probability corresponding to the single-frame phoneme data; the single-frame phoneme data refers to the phonemes included in the audio frame; the mispronunciation probability corresponding to the single-frame phoneme data is used to indicate the pronunciation standard degree of the single-frame phoneme data in the audio frame.

[0097] Optionally, when the computer device obtains the phoneme recognition result, taking the target audio frame in each of the above audio frames as an example, according to the fusion feature corresponding to the target audio frame and the adjacent fusion feature of the fusion feature corresponding to the target audio frame in time series, the local feature corresponding to the target audio frame is obtained; further, the computer device classifies and judges the phoneme errors of the shadowing audio data according to the local features corresponding to each audio frame, and obtains the phoneme recognition results corresponding to each audio frame. Among them, the adjacent fusion feature refers to the fusion feature corresponding to the adjacent audio frame of the target audio frame. Exemplarily, if the time order of the audio frames in the shadowing audio data is audio frame S1, audio frame S2, audio frame S3..., then the adjacent audio frame of audio frame S1 is audio frame S2, and the adjacent fusion feature of the fusion feature corresponding to audio frame S1 is the fusion feature corresponding to audio frame S2.

[0098] 2. Perform a time series merging process on each phoneme recognition result to obtain at least one phoneme data included in the shadowing audio data and the mispronunciation probability of each phoneme data.

[0099] In the embodiment of the present application, after the computer device obtains the above phoneme recognition result, it performs a time series merging process on the phoneme recognition results corresponding to each audio frame to obtain at least one audio data included in the shadowing audio data and the mispronunciation probability of each audio data.

[0100] Optionally, when the computer device obtains the audio data and the mispronunciation probability, based on the time order of each audio frame in the shadowing audio data, it sorts the single-frame phoneme data included in each phoneme recognition result to obtain the sorted single-frame phoneme data; further, it merges the adjacent and identical phoneme data in the sorted single-frame phoneme data into the same phoneme data to obtain at least one phoneme data included in the shadowing audio data. Then, for the target phoneme data among the above at least one phoneme data, the mispronunciation probability of the target phoneme data is determined according to the mispronunciation probability of the single-frame phoneme data corresponding to the target phoneme data.

[0101] Optionally, when the computer device obtains the mispronunciation probability of the target phoneme data, it determines the average value of the mispronunciation probabilities of the single-frame phoneme data corresponding to the target phoneme data as the mispronunciation probability of the target phoneme data. Of course, in the exemplary embodiment, in addition to the method of taking the average value, the mispronunciation probability of the target phoneme data can also be obtained by other methods, such as taking the median, taking the mode, etc.

[0102] Exemplarily, such as Figure 5As shown, the sorted single-frame phoneme data is: W (mispronunciation probability: 0.1), W (mispronunciation probability: 0.4), IH (mispronunciation probability: 0.3), IH (mispronunciation probability: 0.8), IH (mispronunciation probability: 0.8), L (mispronunciation probability: 0.4), L (mispronunciation probability: 0.6), L (mispronunciation probability: 0.8). The phoneme data included in the shadowing audio data obtained after sequential merging processing is: W (mispronunciation probability: 0.25), IH (mispronunciation probability: 0.63), L (mispronunciation probability: 0.6). Subsequently, the phonemes included in the shadowing text are obtained as: W, IH, T. And the mispronunciation probability of the phoneme data is proportional to the pronunciation standard degree. On this basis, the computer device determines that the first mispronounced phoneme of the shadowing audio data is W, the second mispronounced phoneme is L for the mispronounced phoneme of T, and the correctly pronounced phoneme is IH.

[0103] Optionally, in the embodiments of the present application, the above phoneme-level pronunciation error correction method may be executed by a phoneme detection model, and the phoneme detection model includes a feature fusion network, a local feature extraction network, and a phoneme classification and error judgment network. Among them, the feature fusion network is used to fuse the audio features and phoneme features corresponding to each audio frame respectively to obtain the fusion features corresponding to each audio frame; the local feature extraction network is used to obtain the local features corresponding to the target audio frame according to the fusion features corresponding to the target audio frame in each audio frame and the adjacent fusion features of the fusion features corresponding to the target audio frame in time series; the phoneme classification and error judgment network is used to perform phoneme classification and phoneme error judgment on the shadowing audio data according to the local features corresponding to each audio frame respectively to obtain the phoneme recognition results corresponding to each audio frame.

[0104] Next, the training method of the phoneme detection model will be introduced and described in combination with embodiments.

[0105] Please refer to Figure 6 , which shows a flowchart of the training method of the phoneme detection model provided by an embodiment of the present application. This method can be applied to Figure 1 the server 20 of the phoneme-level pronunciation error correction system shown in Figure 1 (not shown in

[0106] Step 601, obtain the training samples of the phoneme detection model.

[0107] The training samples are used to train the phoneme detection model. Among them, the training samples include sample shadowing text and sample shadowing audio data corresponding to the sample shadowing text. In the embodiments of the present application, before training the phoneme detection model, the training samples of the phoneme detection model are obtained.

[0108] Optionally, obtain the follow-up text and the corresponding follow-up audio data in the network environment to obtain the above training samples; or obtain the above training samples by means of recording.

[0109] Optionally, in the embodiments of the present application, each sample follow-up audio data in the above training samples corresponds to a label, and the label is used to indicate the accurate phoneme data included in the sample follow-up audio data.

[0110] Step 602: Obtain the audio features corresponding to each sample audio frame in the sample follow-up audio data, and obtain the phoneme features of the phonemes included in the sample follow-up text.

[0111] In the embodiments of the present application, after the computer device obtains the sample follow-up text and the sample follow-up audio data corresponding to the sample follow-up text, it obtains the audio features corresponding to each sample audio frame in the sample follow-up audio data, and obtains the phoneme features of the phonemes included in the sample follow-up text.

[0112] Optionally, in the embodiments of the present application, after the computer device obtains the above sample follow-up audio data, it performs frame splitting on the sample follow-up audio data to obtain at least one sample audio frame of the sample follow-up audio data, and then performs feature extraction on each sample audio frame to obtain the audio features corresponding to each sample audio frame.

[0113] Optionally, in the embodiments of the present application, after the computer device obtains the above sample follow-up text, it performs feature extraction on the phonemes included in the sample follow-up text to obtain phoneme features, where the phoneme features include the phoneme feature representations corresponding to each phoneme included in the sample follow-up text.

[0114] Step 603: Fuse the audio features corresponding to each sample audio frame with the phoneme features respectively to obtain the fusion features corresponding to each audio frame.

[0115] In the embodiments of the present application, after the computer device obtains the above audio features and phoneme features, it fuses the audio features corresponding to each sample audio frame with the audio features respectively to obtain the fusion features corresponding to each sample audio frame.

[0116] Taking the target sample audio frame in each of the above sample audio frames as an example, optionally, after the computer device obtains the audio features corresponding to the target sample audio frame, according to the phoneme features and the audio features corresponding to the target sample audio frame, it generates the features to be spliced corresponding to the target sample audio frame. Further, it splices the audio features corresponding to the target sample audio frame and the features to be spliced corresponding to the target sample audio frame to obtain the fused features corresponding to the target sample audio frame. Optionally, in the embodiments of the present application, feature fusion of the audio features and phoneme features is performed through an attention mechanism.

[0117] Step 604, according to the fused features respectively corresponding to each sample audio frame, obtain at least one phoneme data included in the sample follow-up audio data, and the mispronunciation probability of each phoneme data.

[0118] In the embodiments of the present application, after the computer device obtains the above-mentioned fused features, according to the fused features respectively corresponding to each sample audio frame, it obtains at least one phoneme data included in the sample follow-up audio data, and the mispronunciation probability of each audio data. It should be noted that the phoneme data included in the sample follow-up audio data obtained at this time is the predicted phoneme data obtained through the phoneme detection model.

[0119] Optionally, when the computer device obtains the phoneme data and the mispronunciation probability, according to the fused features respectively corresponding to each sample audio frame, it performs phoneme classification and phoneme error judgment on the sample follow-up audio data to obtain the phoneme recognition results respectively corresponding to each sample audio frame; further, it performs temporal merging processing on each phoneme recognition result to obtain at least one phoneme data included in the sample follow-up audio data, and the mispronunciation probability of each phoneme data. Among them, the above phoneme recognition results include the single-frame phoneme data corresponding to the sample audio frame, and the mispronunciation probability corresponding to the single-frame phoneme data.

[0120] Optionally, when the computer device obtains the phoneme recognition result, taking the target sample audio frame in each of the above sample audio frames as an example, according to the fused feature corresponding to the target fake sample audio frame and the adjacent fused features in time series of the fused feature corresponding to the target sample audio frame, it obtains the local feature corresponding to the target sample audio frame; further, the computer device performs phoneme classification and phoneme error judgment on the sample follow-up audio data according to the local features respectively corresponding to each sample audio frame to obtain the phoneme recognition results respectively corresponding to each sample audio frame.

[0121] Optionally, after obtaining the phoneme recognition result, the computer device sorts the single-frame phoneme data included in each phoneme recognition result based on the time sequence of each sample audio frame in the sample follow-up audio data to obtain the sorted single-frame phoneme data; further, merges adjacent and identical phoneme data in the sorted single-frame phoneme data into the same phoneme data to obtain at least one phoneme data included in the sample follow-up audio data. Then, for the target phoneme data among the at least one phoneme data, the mispronunciation probability of the target phoneme data is determined according to the mispronunciation probability of the single-frame phoneme data corresponding to the target phoneme data.

[0122] Step 605: Determine the phoneme detection result of the sample follow-up audio data based on the mispronunciation probability of each phoneme data.

[0123] In the embodiment of the present application, after obtaining the mispronunciation probability, the computer device determines the phoneme detection result in the sample follow-up audio data based on the mispronunciation probability of each phoneme data. Among them, the phoneme detection result includes a phoneme classification result and a phoneme misjudgment result. The phoneme classification result is used to represent the phoneme data included in the sample follow-up audio data, and the phoneme misjudgment result is used to represent the mispronunciation probability of the phoneme data included in the sample follow-up audio data.

[0124] Step 606: Calculate the training loss of the phoneme detection model according to the phoneme detection result and the label of the training sample.

[0125] In the embodiment of the present application, after obtaining the above phoneme detection result, the computer device calculates the training loss of the phoneme detection model according to the phoneme detection result and the label of the training sample. Optionally, the computer device obtains the first sub-loss and the second sub-loss of the phoneme detection model according to the phoneme detection result and the label of the training sample; further, determines the training loss of the phoneme detection model according to the first sub-loss and the second sub-loss. Among them, the first sub-loss is used to measure the accuracy of the phoneme classification result in the phoneme detection result, and the second sub-loss is used to measure the accuracy of the phoneme misjudgment result in the phoneme detection result.

[0126] Exemplarily, assume that the phoneme classification result is p phone , after the computer device obtains the phoneme classification result as p phone , according to the phoneme classification result as p phone and the label of the training sample, the obtained first sub-loss L phone is:

[0127]

[0128] Assume that the phoneme misjudgment result is p mis , after the computer device obtains the phoneme misjudgment result as p mis , according to the phoneme misjudgment result as pmis and the second sub-loss L obtained from the labels of the training samples mis is as follows:

[0129]

[0130] The training loss L of the phoneme detection model total is as follows:

[0131] L total = λL phone + (1 - λ)L mis ;

[0132] where m above refers to the number of samples in the training samples, and this number of samples can be the number of sample follow-up text or the number of sample follow-up audio data; c above refers to the number of categories of phoneme data corresponding to the training samples. Optionally, in the embodiments of the present application, the categories of phoneme data include 39 phonemes in the phoneme dictionary and the silent segment, that is, the number of categories of phoneme data is 40; the above refers to the probability that the i-th sample follow-up audio data indicated by the label of the training sample contains the j-th category of phoneme data; the above refers to the probability that the i-th sample follow-up audio data indicated by the phoneme classification result contains the j-th category of phoneme data; c mis refers to the number of categories of phoneme misjudgment results corresponding to the training samples. Optionally, in the embodiments of the present application, the phoneme misjudgment results include two categories: error and non-error, that is, c mis is 2; the above refers to the probability that the i-th sample follow-up audio data indicated by the label of the training sample contains the j-th category of phoneme misjudgment result; the above refers to the probability that the i-th sample follow-up audio data indicated by the phoneme misjudgment result contains the j-th category of phoneme misjudgment result.

[0133] Step 607: Adjust the parameters of the phoneme detection model according to the training loss.

[0134] In the embodiments of the present application, after the computer device obtains the above training loss, it adjusts the parameters of the phoneme detection model according to the training loss. Further, for the phoneme detection model with adjusted parameters, continue to use the above training samples for training and obtain a new training loss, and adjust the parameters of the phoneme detection model according to the new training loss, and then continue to repeat the above steps until the training loss converges, then it is determined that the training of the phoneme detection model is completed.

[0135] In summary, in the technical solution provided by the embodiment of the present application, through the phoneme detection model, according to the phoneme features of the phonemes included in the follow-up sample and the audio features respectively corresponding to each audio frame of the follow-up audio data, the mispronounced phonemes in the follow-up audio data are obtained. On the one hand, in the process of phoneme error judgment for the follow-up audio data, in addition to processing the audio features of the follow-up audio data, the phoneme features of the phonemes included in the follow-up text are also considered, improving the accuracy of phoneme error judgment. On the other hand, the phoneme detection model can fuse the phoneme features of the phonemes included in the follow-up text into the audio features. When phoneme error judgment can be performed through the fused features subsequently, it is not necessary to generate standard pronunciation data based on the phonemes included in the follow-up text. Phoneme error judgment can be performed through the phoneme detection model, simplifying the phoneme error judgment process and improving the processing efficiency of phoneme error judgment.

[0136] Optionally, in the embodiment of the present application, in addition to including the above-mentioned feature fusion network, the above-mentioned local feature extraction network, and the above-mentioned phoneme classification and error judgment network, the phoneme detection model further includes: a pre-trained audio feature extraction network and a pre-trained phoneme feature extraction network. Among them, the audio feature extraction network is used to extract features from each audio frame respectively to obtain the audio features respectively corresponding to each audio frame; the phoneme feature extraction network is used to extract features from the phonemes included in the follow-up text to obtain the above-mentioned phoneme features. Of course, in the exemplary embodiment, the phoneme detection model further includes: a result output module. Among them, the result output module is used to output the mispronounced phonemes in the follow-up text according to the mispronunciation probability of each phoneme data.

[0137] Exemplarily, in combination with reference to Figure 7, after obtaining the follow - up audio data corresponding to the follow - up text, input the phonemes included in the follow - up text and the follow - up audio data into the phoneme detection model 70. Then, through the audio feature extraction network 71 in the phoneme detection model 70, obtain the audio features corresponding to each audio frame of the follow - up audio data respectively, and through the phoneme feature extraction network 72 in the phoneme detection model 70, obtain the phoneme features of the phonemes included in the follow - up text; further, through the feature fusion network 73, fuse the audio features corresponding to each audio frame with the phoneme features respectively to obtain the fusion features corresponding to each audio frame, where the feature fusion network 73 includes an attention mechanism; then, through the local feature extraction network 74, according to the fusion feature corresponding to the target audio frame in each audio frame and the adjacent fusion features of the fusion feature corresponding to the target audio frame in time series, obtain the local feature corresponding to the target audio frame, and further obtain the local features corresponding to each audio frame respectively; through the phoneme classification and error judgment network 75, according to the local features corresponding to each audio frame respectively, perform phoneme classification and phoneme error judgment on the follow - up audio data to obtain at least one phoneme data included in the follow - up audio data and the mispronunciation probability corresponding to each phoneme data respectively; finally, through the result output block 76, according to the mispronunciation probability of each phoneme data, output the mispronounced phonemes in the follow - up text.

[0138] It should be noted that for some details of the phoneme detection model, reference can be made to the content introduced in the above Figure 3 embodiment.

[0139] It should also be noted that the introduction of the present application through the embodiments above is only exemplary and explanatory. New embodiments formed by arbitrarily combining the steps in the above embodiments are also within the protection scope of the present application.

[0140] The following is an embodiment of the apparatus of the present application, which can be used to execute the method embodiment of the present application. For details not disclosed in the embodiment of the apparatus of the present application, please refer to the method embodiment of the present application.

[0141] Please refer to Figure 8 , which shows a block diagram of a phoneme - level pronunciation error correction device provided by an embodiment of the present application. This device has the function of implementing the above - mentioned phoneme - level pronunciation error correction method, and this function can be implemented by hardware or by hardware executing corresponding software. This device can be a computer device or can be set in a computer device. The device 800 may include: an audio acquisition module 810, a feature acquisition module 820, a feature fusion module 830, a probability acquisition module 840, and a phoneme error judgment module 850.

[0142] The audio acquisition module 810 is used to acquire the follow - up audio data corresponding to the follow - up text.

[0143] A feature acquisition module 820, configured to acquire audio features respectively corresponding to each audio frame in the shadowing audio data, and acquire phoneme features of phonemes included in the shadowing text.

[0144] A feature fusion module 830, configured to respectively fuse the audio features respectively corresponding to each audio frame with the phoneme features to obtain fusion features respectively corresponding to each audio frame.

[0145] A probability acquisition module 840, configured to obtain at least one phoneme data included in the shadowing audio data and the mispronunciation probability of each phoneme data according to the fusion features respectively corresponding to each audio frame.

[0146] A phoneme misjudgment module 850, configured to determine mispronounced phonemes in the shadowing audio data based on the mispronunciation probability of each phoneme data.

[0147] In an exemplary embodiment, as Figure 9 shown, the feature fusion module 830 includes: a feature fusion unit 831 and a feature splicing unit 832.

[0148] The feature fusion unit 831 is configured to generate, for a target audio frame in each audio frame, a to-be-spliced feature corresponding to the target audio frame according to the phoneme feature and the audio feature corresponding to the target audio frame.

[0149] The feature splicing unit 832 is configured to splice the audio feature corresponding to the target audio frame and the to-be-spliced feature corresponding to the target audio frame to obtain a fusion feature corresponding to the target audio frame.

[0150] In an exemplary embodiment, as Figure 9 shown, the probability acquisition module 840 includes: a data acquisition unit 841 and a data merging unit 842.

[0151] The data acquisition unit 841 is configured to perform phoneme classification and phoneme misjudgment on the shadowing audio data according to the fusion features respectively corresponding to each audio frame to obtain phoneme recognition results respectively corresponding to each audio frame; wherein, the phoneme recognition results include single-frame phoneme data corresponding to the audio frame and the mispronunciation probability corresponding to the single-frame phoneme data.

[0152] The data merging unit 842 is configured to perform sequential merging processing on each phoneme recognition result to obtain at least one phoneme data included in the shadowing audio data and the mispronunciation probability of each phoneme data.

[0153] In an exemplary embodiment, the data acquisition unit 841 is configured to obtain, for a target audio frame in each of the audio frames, a local feature corresponding to the target audio frame according to the fusion feature corresponding to the target audio frame and the adjacent fusion features of the fusion feature corresponding to the target audio frame in time series; and perform phoneme classification and phoneme error judgment on the shadowing audio data according to the local features respectively corresponding to the audio frames, so as to obtain a phoneme recognition result respectively corresponding to each of the audio frames.

[0154] In an exemplary embodiment, the data merging unit 842 is configured to sort the single-frame phoneme data included in each of the phoneme recognition results based on the time sequence of each of the audio frames in the shadowing audio data, so as to obtain sorted single-frame phoneme data; merge adjacent and identical phoneme data in the sorted single-frame phoneme data into the same phoneme data, so as to obtain at least one phoneme data included in the shadowing audio data; and determine a mispronunciation probability of the target phoneme data according to the mispronunciation probability of the single-frame phoneme data corresponding to the target phoneme data for the at least one phoneme data.

[0155] In an exemplary embodiment, the phoneme error judgment module 850 is configured to determine the phoneme data whose mispronunciation probability meets the condition as the first mispronounced phoneme in the shadowing audio data; wherein the first mispronounced phoneme refers to the phoneme with inaccurate pronunciation in the shadowing audio data.

[0156] In an exemplary embodiment, the phoneme error judgment module 850 is further configured to, if there is phoneme data in the phoneme data included in the shadowing audio data that does not match the phonemes included in the shadowing text, determine the unmatched phoneme data as the second mispronounced phoneme in the shadowing audio data; wherein the second mispronounced phoneme refers to the incorrect phoneme in the shadowing audio data relative to the shadowing text.

[0157] In summary, in the technical solution provided by the embodiment of the present application, through the fusion of the audio features of the audio frames and the phoneme features of the phonemes included in the shadowing text, the fusion features respectively corresponding to the audio frames are obtained, so that the fusion features can represent both the audio features of the audio frames and the phoneme features. Further, according to the fusion features, the phoneme data included in the shadowing audio data and the mispronunciation probability of the phoneme data are obtained. In the process of phoneme error judgment for the shadowing audio data, in addition to processing the audio features of the shadowing audio data, the phoneme features of the phonemes included in the shadowing text are also considered, thereby improving the accuracy of phoneme error judgment; moreover, by fusing the phoneme features of the phonemes included in the shadowing text into the audio features, when phoneme error judgment can be performed through the fusion features subsequently, it is not necessary to generate standard pronunciation data based on the phonemes included in the shadowing text, simplifying the phoneme error judgment process and improving the processing efficiency of phoneme error judgment.

[0158] Please refer to Figure 10 , which shows a block diagram of a training device for a phoneme detection model provided by an embodiment of the present application. The device has the function of implementing the training method of the above-mentioned phoneme detection model, and this function can be implemented by hardware or by hardware executing corresponding software. The device can be a computer device or can be set in a computer device. The device 1000 may include: a sample acquisition module 1010, a feature determination module 1020, a feature processing module 1030, a probability determination module 1040, a result determination module 1050, a loss acquisition module 1060, and a parameter adjustment module 1070.

[0159] The sample acquisition module 1010 is configured to acquire training samples of the phoneme detection model, and the training samples include sample following texts and sample following audio data corresponding to the sample following texts.

[0160] The feature determination module 1020 is configured to acquire audio features respectively corresponding to each sample audio frame in the sample following audio data, and acquire phoneme features of phonemes included in the sample following text.

[0161] The feature processing module 1030 is configured to fuse the audio features respectively corresponding to each sample audio frame with the phoneme features respectively to obtain fused features respectively corresponding to each audio frame.

[0162] The probability determination module 1040 is configured to acquire at least one phoneme data included in the sample following audio data and the misreading probability of each phoneme data according to the fused features respectively corresponding to each sample audio frame.

[0163] The result determination module 1050 is configured to determine the phoneme detection result of the sample following audio data based on the misreading probability of each phoneme data.

[0164] The loss acquisition module 1060 is configured to calculate the training loss of the phoneme detection model according to the phoneme detection result and the label of the training sample.

[0165] The parameter adjustment module 1070 is configured to adjust the parameters of the phoneme detection model according to the training loss.

[0166] In an exemplary embodiment, the phoneme detection model includes a feature fusion network, a local feature extraction network, and a phoneme classification and error judgment network; wherein, the feature fusion network is configured to fuse the audio features respectively corresponding to each audio frame in the shadowing audio data with the phoneme features of the phonemes included in the shadowing text to obtain the fusion features respectively corresponding to each audio frame; the local feature extraction network is configured to obtain the local feature corresponding to the target audio frame according to the fusion feature corresponding to the target audio frame in each audio frame and the adjacent fusion features of the fusion feature corresponding to the target audio frame in time series; the phoneme classification and error judgment network is configured to perform phoneme classification and phoneme error judgment on the shadowing audio data according to the local features respectively corresponding to each audio frame to obtain the phoneme recognition results respectively corresponding to each audio frame; wherein, the phoneme recognition results include the single-frame phoneme data corresponding to the audio frame and the mispronunciation probability corresponding to the single-frame phoneme data.

[0167] In an exemplary embodiment, the phoneme detection model further includes: a pre-trained audio feature extraction network and a pre-trained phoneme feature extraction network; wherein, the audio feature extraction network is configured to perform feature extraction on each audio frame respectively to obtain the audio features respectively corresponding to each audio frame; the phoneme feature extraction network is configured to perform feature extraction on the phonemes included in the shadowing text to obtain the phoneme features; wherein, the phoneme features include the phoneme feature representations respectively corresponding to each phoneme included in the shadowing text.

[0168] In summary, in the technical solution provided by the embodiments of the present application, through the phoneme detection model, according to the phoneme features of the phonemes included in the shadowing sample and the audio features respectively corresponding to each audio frame of the shadowing audio data, the mispronounced phonemes in the shadowing audio data are obtained. On the one hand, in the process of phoneme error judgment for the shadowing audio data, in addition to processing the audio features of the shadowing audio data, the phoneme features of the phonemes included in the shadowing text are also considered, improving the accuracy of phoneme error judgment. On the other hand, the phoneme detection model can fuse the phoneme features of the phonemes included in the shadowing text into the audio features. When performing phoneme error judgment through the fusion features subsequently, it is not necessary to generate standard pronunciation data based on the phonemes included in the shadowing text. Phoneme error judgment can be performed through the phoneme detection model, simplifying the phoneme error judgment process and improving the processing efficiency of phoneme error judgment.

[0169] It should be noted that, when the device provided in the above embodiments realizes its functions, only the division of the above functional modules is used for illustration. In actual applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the device provided in the above embodiments and the method embodiments belong to the same concept. For the specific implementation process, please refer to the method embodiments and will not be elaborated here.

[0170] Please refer to Figure 11 , which shows a structural block diagram of a computer device provided in an embodiment of the present application. This computer device can be used to implement the functions of the above phoneme-level pronunciation correction method or the training method of the phoneme detection model. Specifically:

[0171] The computer device 1100 includes a central processing unit (CPU) 1101, a system memory 1104 including a random access memory (RAM) 1102 and a read only memory (ROM) 1103, and a system bus 1105 connecting the system memory 1104 and the central processing unit 1101. The computer device 1100 also includes a basic input / output system (Input / Output, I / O system) 1106 for transmitting information between various components within the computer, and a mass storage device 1107 for storing an operating system 1113, application programs 1114, and other program modules 1115.

[0172] The basic input / output system 1106 includes a display 1108 for displaying information and input devices 1109 such as a mouse and a keyboard for user input. Both the display 1108 and the input devices 1109 are connected to the central processing unit 1101 through an input / output controller 1110 connected to the system bus 1105. The basic input / output system 1106 may also include an input / output controller 1110 for receiving and processing inputs from multiple other devices such as a keyboard, a mouse, or an electronic stylus. Similarly, the input / output controller 1110 also provides outputs to a display screen, a printer, or other types of output devices.

[0173] The mass storage device 1107 is connected to the central processing unit 1101 through a mass storage controller (not shown) connected to the system bus 1105. The mass storage device 1107 and its associated computer-readable medium provide non-volatile storage for the computer device 1100. That is to say, the mass storage device 1107 may include a computer-readable medium (not shown) such as a hard disk or a CD-ROM (Compact Disc Read-Only Memory) drive.

[0174] Without loss of generality, computer-readable media can include computer storage media and communication media. Computer storage media includes volatile and non-volatile, removable and non-removable media implemented by any method or technology for storing information such as computer-readable instructions, data structures, program modules, or other data. Computer storage media includes RAM, ROM, EPROM (Erasable Programmable Read Only Memory), EEPROM (Electrically Erasable Programmable Read Only Memory), flash memory or other solid-state storage devices, CD-ROM, DVD (Digital Video Disc) or other optical storage, magnetic tape cartridges, tapes, disk storage or other magnetic storage devices. Of course, those skilled in the art will know that computer storage media is not limited to the above several types. The above-mentioned system memory 1104 and mass storage device 1107 can be collectively referred to as memory.

[0175] According to various embodiments of the present application, the computer device 1100 can also run by connecting to a remote computer on the network such as the Internet. That is, the computer device 1100 can be connected to the network 1112 through the network interface unit 1111 connected to the system bus 1105. Or rather, the network interface unit 1111 can also be used to connect to other types of networks or remote computer systems (not shown).

[0176] The memory further includes a computer program, which is stored in the memory and is configured to be executed by one or more processors to implement the above phoneme-level pronunciation correction method or the above training method of the phoneme detection model.

[0177] In an exemplary embodiment, a computer-readable storage medium is further provided. At least one instruction, at least one program, a code set, or an instruction set is stored in the storage medium. When the at least one instruction, the at least one program, the code set, or the instruction set is executed by a processor, the above phoneme-level pronunciation correction method is implemented, or the above training method of the phoneme detection model is implemented.

[0178] Optionally, the computer-readable storage medium may include: ROM (Read Only Memory), RAM (Random Access Memory), SSD (Solid State Drives), or optical discs, etc. Among them, the random access memory may include ReRAM (Resistance Random Access Memory) and DRAM (Dynamic Random Access Memory).

[0179] In an exemplary embodiment, a computer program product or a computer program is further provided. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the above phoneme-level pronunciation correction method, or executes the above training method of the phoneme detection model.

[0180] It should be understood that "a plurality of" mentioned in this article refers to two or more. "And / or" describes the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally represents an "or" relationship between the associated objects before and after. In addition, the step numbers described in this article only exemplarily show a possible execution sequence between steps. In some other embodiments, the above steps may not be executed in the order of the numbers. For example, two steps with different numbers are executed simultaneously, or two steps with different numbers are executed in the reverse order of the illustration. The embodiments of the present application do not limit this.

[0181] The above are only exemplary embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. A phoneme-level pronunciation error correction method, characterized in that, The method includes: Obtaining the follow - up audio data corresponding to the follow - up text; Obtaining the audio features respectively corresponding to each audio frame in the follow - up audio data, and obtaining the phoneme features of the phonemes included in the follow - up text; Fusing the audio features respectively corresponding to each audio frame with the phoneme features respectively to obtain the fusion features respectively corresponding to each audio frame; For a target audio frame among each audio frame, according to the fusion feature corresponding to the target audio frame and the adjacent fusion features of the fusion feature corresponding to the target audio frame in time series, obtaining the local feature corresponding to the target audio frame; According to the local features respectively corresponding to each audio frame, performing phoneme classification and phoneme error judgment on the follow - up audio data to obtain the phoneme recognition results respectively corresponding to each audio frame; Performing time - series merging processing on each phoneme recognition result to obtain at least one phoneme data included in the follow - up audio data and the mispronunciation probability of each phoneme data; Based on the mispronunciation probabilities of each phoneme data, determining the mispronounced phonemes in the follow - up audio data.

2. The method according to claim 1, characterized in that, The step of fusing the audio features respectively corresponding to each audio frame with the phoneme features respectively to obtain the fusion features respectively corresponding to each audio frame includes: For a target audio frame among each audio frame, generating a to - be - spliced feature corresponding to the target audio frame according to the phoneme feature and the audio feature corresponding to the target audio frame; Splicing the audio feature corresponding to the target audio frame and the to - be - spliced feature corresponding to the target audio frame to obtain the fusion feature corresponding to the target audio frame.

3. The method according to claim 1, characterized in that, The step of performing time - series merging processing on each phoneme recognition result to obtain at least one phoneme data included in the follow - up audio data and the mispronunciation probability of each phoneme data includes: Based on the time order of each audio frame in the follow - up audio data, sorting the single - frame phoneme data included in each phoneme recognition result to obtain the sorted single - frame phoneme data; Merging adjacent and identical phoneme data in the sorted single - frame phoneme data into the same phoneme data to obtain at least one phoneme data included in the follow - up audio data; For a target phoneme data among the at least one phoneme data, determining the mispronunciation probability of the target phoneme data according to the mispronunciation probabilities of the single - frame phoneme data corresponding to the target phoneme data.

4. The method according to any one of claims 1 to 3, characterized in that, The step of determining the mispronounced phonemes in the follow - up audio data based on the mispronunciation probabilities of each phoneme data includes: Determining the phoneme data whose mispronunciation probability meets the condition as the first mispronounced phonemes in the follow - up audio data; Wherein, the first mispronounced phonemes refer to the phonemes with inaccurate pronunciation in the follow - up audio data.

5. The method according to claim 4, characterized in that, The method further includes: If there is phoneme data in the phoneme data included in the follow - up audio data that does not match the phonemes included in the follow - up text, determining the unmatched phoneme data as the second mispronounced phonemes in the follow - up audio data; Wherein, the second mispronounced phonemes refer to the wrong phonemes in the follow - up audio data relative to the follow - up text.

6. A training method for a phoneme detection model, characterized in that, The method includes: Obtain training samples of the phoneme detection model, where the training samples include sample shadowing texts and sample shadowing audio data corresponding to the sample shadowing texts; Obtain audio features respectively corresponding to each sample audio frame in the sample shadowing audio data, and obtain phoneme features of phonemes included in the sample shadowing text; Fuse the audio features respectively corresponding to each sample audio frame with the phoneme features respectively, to obtain fused features respectively corresponding to each audio frame; For a target sample audio frame among each sample audio frame, according to the fused feature corresponding to the target sample audio frame and the adjacent fused features in time series of the fused feature corresponding to the target sample audio frame, obtain the local feature corresponding to the target sample audio frame; According to the local features respectively corresponding to each sample audio frame, perform phoneme classification and phoneme error judgment on the sample shadowing audio data, to obtain phoneme recognition results respectively corresponding to each sample audio frame; Perform time series merging processing on each phoneme recognition result, to obtain at least one phoneme data included in the sample shadowing audio data and the mispronunciation probability of each phoneme data; Based on the mispronunciation probability of each phoneme data, determine the phoneme detection result of the sample shadowing audio data; Calculate the training loss of the phoneme detection model according to the phoneme detection result and the label of the training sample; Adjust the parameters of the phoneme detection model according to the training loss.

7. The method according to claim 6, characterized in that, The phoneme detection model includes a feature fusion network, a local feature extraction network, and a phoneme classification and error judgment network; wherein, The feature fusion network is used to fuse the audio features respectively corresponding to each audio frame in the shadowing audio data with the phoneme features of phonemes included in the shadowing text respectively, to obtain fused features respectively corresponding to each audio frame; The local feature extraction network is used to obtain the local feature corresponding to the target audio frame according to the fused feature corresponding to the target audio frame among each audio frame and the adjacent fused features in time series of the fused feature corresponding to the target audio frame; The phoneme classification and error judgment network is used to perform phoneme classification and phoneme error judgment on the shadowing audio data according to the local features respectively corresponding to each audio frame, to obtain phoneme recognition results respectively corresponding to each audio frame; wherein, the phoneme recognition result includes the single-frame phoneme data corresponding to the audio frame and the mispronunciation probability corresponding to the single-frame phoneme data.

8. The method according to claim 7, wherein The phoneme detection model further includes: a pre-trained audio feature extraction network and a pre-trained phoneme feature extraction network; wherein, The audio feature extraction network is used to perform feature extraction on each audio frame respectively, to obtain audio features respectively corresponding to each audio frame; The phoneme feature extraction network is used to perform feature extraction on the phonemes included in the shadowing text, to obtain the phoneme features; wherein, the phoneme features include phoneme feature representations respectively corresponding to each phoneme included in the shadowing text.

9. A phoneme-level pronunciation error correction device, wherein The device includes: An audio acquisition module, used to acquire shadowing audio data corresponding to a shadowing text; A feature acquisition module, configured to acquire the audio features respectively corresponding to each audio frame in the shadowing audio data, and acquire the phoneme features of the phonemes included in the shadowing text; A feature fusion module, configured to fuse the audio features respectively corresponding to each audio frame with the phoneme features respectively, to obtain the fusion features respectively corresponding to each audio frame; A probability acquisition module, for a target audio frame in each audio frame, according to the fusion feature corresponding to the target audio frame, and the adjacent fusion features of the fusion feature corresponding to the target audio frame in time series, to obtain the local feature corresponding to the target audio frame; according to the local features respectively corresponding to each audio frame, perform phoneme classification and phoneme error judgment on the shadowing audio data, to obtain the phoneme recognition results respectively corresponding to each audio frame; perform time series merging processing on each phoneme recognition result, to obtain at least one phoneme data included in the shadowing audio data, and the mispronunciation probability of each phoneme data; A phoneme error judgment module, configured to determine the mispronounced phonemes in the shadowing audio data based on the mispronunciation probabilities of each phoneme data.

10. A training device for a phoneme detection model, wherein The apparatus includes: A sample acquisition module, configured to acquire the training samples of the phoneme detection model, where the training samples include sample shadowing texts and the sample shadowing audio data corresponding to the sample shadowing texts; A feature determination module, configured to acquire the audio features respectively corresponding to each sample audio frame in the sample shadowing audio data, and acquire the phoneme features of the phonemes included in the sample shadowing text; A feature processing module, configured to fuse the audio features respectively corresponding to each sample audio frame with the phoneme features respectively, to obtain the fusion features respectively corresponding to each audio frame; A probability determination module, for a target sample audio frame in each sample audio frame, according to the fusion feature corresponding to the target sample audio frame, and the adjacent fusion features of the fusion feature corresponding to the target sample audio frame in time series, to obtain the local feature corresponding to the target sample audio frame; according to the local features respectively corresponding to each sample audio frame, perform phoneme classification and phoneme error judgment on the sample shadowing audio data, to obtain the phoneme recognition results respectively corresponding to each sample audio frame; perform time series merging processing on each phoneme recognition result, to obtain at least one phoneme data included in the sample shadowing audio data, and the mispronunciation probability of each phoneme data; A result determination module, configured to determine the phoneme detection result of the sample shadowing audio data based on the mispronunciation probabilities of each phoneme data; A loss acquisition module, configured to calculate the training loss of the phoneme detection model according to the phoneme detection result and the label of the training sample; A parameter adjustment module, configured to adjust the parameters of the phoneme detection model according to the training loss.

11. A computer device, wherein The computer device includes a processor and a memory. A computer program is stored in the memory and is loaded and executed by the processor to implement the phoneme-level pronunciation correction method according to any one of claims 1 to 5, or to implement the training method of the phoneme detection model according to any one of claims 6 to 8.

12. A computer-readable storage medium, wherein A computer program is stored in the storage medium and is loaded and executed by a processor to implement the phoneme-level pronunciation correction method according to any one of claims 1 to 5, or to implement the training method of the phoneme detection model according to any one of claims 6 to 8.

13. A computer program product, wherein The computer program product includes a computer program. The computer program is stored in a computer-readable storage medium, and the processor reads and executes the computer program to implement the phoneme-level pronunciation correction method according to any one of claims 1 to 5, or to implement the training method of the phoneme detection model according to any one of claims 6 to 8.

Citation Information

Patent Citations

  • Pronunciation detection method and apparatus

    CN101510423A

  • Standard pronunciation generation method and related device

    CN111930900A

  • Pronunciation detection method and device and computer readable medium

    CN113409768A