English spoken pronunciation correction auxiliary system based on speech recognition
By using a visual action feature extraction method based on dimensional fusion and feature simplification, combined with 3D convolution and multi-level dense layers and the Swin Transformer module, the problem of inaccurate lip movement guidance in the existing system is solved, and more accurate pronunciation correction assistance is achieved.
Patent Information
- Application Number
- CN202510906500.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-02
- Publication Date
- 2025-09-09
- Estimated Expiration
- 2045-07-02
AI Technical Summary
Existing English oral pronunciation correction assistance systems are rather vague in their guidance suggestions, making it difficult to accurately identify pronunciation accuracy and lip movements, and lack precise guidance.
A visual action feature extraction method based on dimensionality fusion and feature simplification is adopted. The three-dimensional video data is converted into two-dimensional images through a visual feature extraction method of dimensionality transposition and reshaping. 3D convolution and multi-level dense layers are combined with the Swin Transformer module to capture the detailed features of lip movements. A multi-head self-attention mechanism is introduced to optimize semantic features, thereby achieving accurate association between lip movements and speech text.
It achieves more precise lip movement guidance, improves the accuracy of pronunciation correction and the meticulousness of guidance, and enhances the understanding of the correlation between lip movements and speech text.
Smart Images

Figure CN120412648B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of intelligent speech recognition, and in particular relates to an English spoken pronunciation correction auxiliary system based on speech recognition. Background Art
[0002] The English pronunciation correction assistance system is a tool based on computer technology and phonetics principles. It combines speech recognition, machine learning algorithms, and computer vision technology to help learners improve their spoken English pronunciation, providing personalized pronunciation guidance and feedback. However, existing pronunciation correction assistance systems offer vague guidance, making it difficult to accurately identify pronunciation accuracy and provide lip movement guidance based on spoken text. Summary of the Invention
[0003] In view of the above situation, in order to overcome the defects of the prior art, the present invention provides an English oral pronunciation correction auxiliary system based on speech recognition. In view of the problem that the prior art guidance suggestions are relatively vague and it is difficult to accurately identify the pronunciation accuracy and guide the lip movements based on the spoken text, the present invention adopts a visual action feature extraction method based on dimension fusion and feature simplification to extract the lip movements of the standard pronunciation corresponding to the spoken text, and extracts the convolution features of the pronunciation video through a visual feature extraction method based on dimension transposition and reshaping. The three-dimensional video data is converted into two-dimensional image data through a dimension transposition operation. While generating a two-dimensional image, the key information in the video is retained, and 3D convolution is used to capture the dynamic movements and subtle changes of the lips, accurately identifying the detailed features of the lip movements of the standard pronunciation, and on this basis, a multi-level dense layer and Swin The joint optimization processing of the Transformer module further extracts image features, increases the number of channels on the basis of dimensionality reduction, captures deeper and more detailed lip visual information, and provides more accurate lip movement guidance for pronunciation correction; the present invention creatively guides a more detailed understanding and recognition of lip movements through semantic features on the basis of a visual action feature extraction method based on dimension fusion and feature simplification, and can accurately select the corresponding lip action label description according to the lip action image recognition and classification results, introduces a multi-head self-attention mechanism to optimize the semantic features, better captures the correlation between lip actions and speech text, and adds corresponding position codes to words in the speech text, so that the model understands the order of spoken text pronunciation actions in the lip action label description, and achieves more accurate lip action guidance.
[0004] The English spoken pronunciation correction auxiliary system based on speech recognition provided by the present invention includes a data acquisition module, a speech recognition module, a pronunciation analysis module and a pronunciation correction module;
[0005] The data acquisition module collects audio of the standard pronunciation of spoken English and collects lip pronunciation video of the standard pronunciation of spoken English;
[0006] The speech recognition module converts the standard pronunciation audio into speech text using a speech recognition model;
[0007] The pronunciation analysis module extracts lip movement features from the pronunciation video using a text-guided cross-modal feature fusion extraction method, performs activation classification on the movement features and adds lip movement description labels, and extracts audio features from the corresponding standard pronunciation audio to obtain a joint analysis result of the standard pronunciation spoken pronunciation, speech text, and lip movement description.
[0008] The pronunciation correction module collects the audio that needs to be corrected and the corresponding speech text and sets a correction threshold, extracts the audio features of the audio that needs to be corrected, and compares the similarity with the audio features of the standard pronunciation. When the similarity is higher than the correction threshold, no correction assistance is performed. If the similarity is lower than the correction threshold, the corresponding lip movement description is provided for correction guidance.
[0009] Furthermore, in the pronunciation analysis module, lip movement features of the pronunciation video are extracted by a text-guided cross-modal feature fusion extraction method, which specifically includes the following steps:
[0010] Step S1: Semantic feature extraction based on multi-head self-attention mechanism optimization is performed on the speech text to obtain semantic features, which specifically includes the following steps:
[0011] Step S11: Contextual text feature extraction based on the self-attention mechanism is performed on the speech text to obtain text features:
[0012] ;
[0013] Where, Represents the query vector, key vector, and value vector obtained by encoding the speech text through the self-attention mechanism. Represents the key dimensions of the preset, Represents the relative position deviation parameter, That is, it represents the text features;
[0014] Step S12: Multi-head self-attention mechanism optimization, multi-head self-attention mechanism optimization is performed on the text features to obtain optimized text features:
[0015] ;
[0016] ;
[0017] Where, Represents the head of the multi-head self-attention mechanism, 、 、 and Represent the preset projection parameter matrix, Represents the splicing operation, This means optimizing text features;
[0018] Step S13: Positional encoding embedding: positional encoding embedding is performed on the optimized text features corresponding to all words in the speech text to obtain semantic features. The positional encoding is calculated as follows:
[0019] ;
[0020] Where, Represents the position of the word in the speech text, Represents the dimension of the optimized text feature corresponding to the word, represents the preset embedding dimension, Represents the voice text The position code corresponding to the word;
[0021] Step S2: extracting visual action features based on dimension fusion and feature simplification, using a visual action feature extraction method based on dimension fusion and feature simplification to extract optimized visual features of the pronunciation video, specifically including the following steps:
[0022] Step S21: Convolutional feature extraction, using a visual feature extraction method based on dimensionality transposition and reshaping to extract convolutional features of the pronunciation video, specifically including the following steps:
[0023] Step S211: Video input, input the pronunciation video and obtain its dimensions (BZ, C, L, H, W), which represent the batch size, number of channels, video length, video height and video width of the pronunciation video respectively;
[0024] Step S212: Dimension transposition, swapping the first dimension and the second dimension in the pronunciation video to obtain a transposed video;
[0025] Step S213: reshape the dimension of the transposed video to obtain a reshaped video with a dimension of (BZ×L, C, H, W);
[0026] Step S214: 3D convolution, performing convolution feature extraction based on 3D convolution on the reshaped video to obtain convolution features;
[0027] Step S22: Feature dimensionality reduction, using the Linear Embedding layer to reduce the dimensionality of the convolutional features to obtain reduced dimensionality features;
[0028] Step S23: Feature optimization, which performs a joint optimization process based on a multi-level dense layer and a Swin Transformer module on the dimensionality reduction features to obtain optimized visual features, specifically including the following steps:
[0029] Step S231: Based on the dense layer processing of multi-layer 2D convolution, three layers of BN-ReLU-2DConv processing are performed on the dimensionality reduction features to obtain dense optimized convolution features, where the convolution kernel sizes of each 2DConv layer are 1×1, 3×3, and 1×1 respectively;
[0030] Step S232: Swin Transformer module processing, using the Swin Transformer module to process the dimensionality reduction features to obtain efficient optimized convolution features;
[0031] Step S233: feature fusion, patch merging the densely optimized convolution features and the efficient optimized convolution features to obtain primary optimized convolution features;
[0032] Step S234: multi-layer optimization, repeating steps S32 to S34 n times to obtain optimized visual features;
[0033] Step S3: Cross-modal feature fusion, scaling the semantic features to align them with the optimized visual features, and then performing feature fusion to obtain lip movement features.
[0034] Furthermore, in the pronunciation analysis module, the extraction of audio features specifically includes the following steps:
[0035] Step Q1: audio preprocessing, denoising, equalization, and downsampling the audio signal;
[0036] Step Q2: Audio framing: Use a Hamming window to divide the audio signal into short time periods, namely audio frames;
[0037] Step Q3: Fourier transform, perform Fourier transform on each audio frame to obtain a spectrum;
[0038] Step Q4: Feature extraction, extracting features based on spectrum information from the spectrum graph to obtain spectrum features;
[0039] Step Q5: Feature dimensionality reduction and normalization: perform dimensionality reduction and normalization on the extracted spectral features to obtain audio features.
[0040] The beneficial results achieved by the present invention using the above scheme are as follows:
[0041] (1) The present invention creatively adopts a visual action feature extraction method based on dimensional fusion and feature simplification to extract the lip movements of the standard pronunciation corresponding to the spoken text, extracts the convolution features of the pronunciation video through a visual feature extraction method based on dimensional transposition and reshaping, converts the three-dimensional video data into two-dimensional image data through the dimensional transposition operation, retains the key information in the video while generating the two-dimensional image, and uses 3D convolution to capture the dynamic movements and subtle changes of the lips, accurately identifies the detailed features of the lip movements of the standard pronunciation, and then performs a joint optimization process based on multi-level dense layers and SwinTransformer modules to further extract image features, increase the number of channels on the basis of dimensionality reduction, capture deeper and more detailed lip visual information, and provide more accurate lip movement guidance for pronunciation correction;
[0042] (2) Based on a visual action feature extraction method based on dimension fusion and feature simplification, the present invention guides a more detailed understanding and recognition of lip movements through semantic features, and can accurately select the corresponding lip action label description according to the lip action image recognition and classification results. It introduces a multi-head self-attention mechanism to optimize the semantic features, better capture the association between lip actions and speech text, and add corresponding position codes to words in the speech text, so that the model can understand the order of spoken text pronunciation actions in the lip action label description, and achieve more accurate lip action guidance. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] Figure 1 A module diagram of the English spoken pronunciation correction auxiliary system based on speech recognition provided by the present invention;
[0044] Figure 2 The figure is a flowchart of a text-guided cross-modal feature fusion extraction method.
[0045] The accompanying drawings are used to provide further understanding of the present invention and constitute a part of the specification. They are used to explain the present invention together with the embodiments of the present invention and do not constitute a limitation of the present invention. DETAILED DESCRIPTION
[0046] The technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments; based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.
[0047] Example 1, see Figure 1, an English spoken pronunciation correction auxiliary system based on speech recognition, including a data acquisition module, a speech recognition module, a pronunciation analysis module and a pronunciation correction module;
[0048] The data acquisition module collects audio of the standard pronunciation of spoken English and collects lip pronunciation video of the standard pronunciation of spoken English;
[0049] The speech recognition module converts the standard pronunciation audio into speech text using a speech recognition model;
[0050] The pronunciation analysis module extracts lip movement features from the pronunciation video using a text-guided cross-modal feature fusion extraction method, performs activation classification on the movement features and adds lip movement description labels, and extracts audio features from the corresponding standard pronunciation audio to obtain a joint analysis result of the standard pronunciation spoken pronunciation, speech text, and lip movement description.
[0051] The pronunciation correction module collects the audio that needs to be corrected and the corresponding speech text and sets a correction threshold, extracts the audio features of the audio that needs to be corrected, and compares the similarity with the audio features of the standard pronunciation. When the similarity is higher than the correction threshold, no correction assistance is performed. If the similarity is lower than the correction threshold, the corresponding lip movement description is provided for correction guidance.
[0052] Example 2, see Figure 2 This embodiment is based on the above embodiment. In the pronunciation analysis module, a text-guided cross-modal feature fusion extraction method is used to extract lip movement features from the pronunciation video. Specifically, the following steps are included:
[0053] Step S1: extracting semantic features based on the multi-head self-attention mechanism optimization, extracting semantic features based on the multi-head self-attention mechanism optimization of the speech text to obtain semantic features;
[0054] Step S2: extracting visual action features based on dimension fusion and feature simplification, using a visual action feature extraction method based on dimension fusion and feature simplification to extract optimized visual features of the pronunciation video;
[0055] Step S3: Cross-modal feature fusion, scaling the semantic features to align them with the optimized visual features, and then performing feature fusion to obtain lip movement features.
[0056] By performing the above operations, the convolution features of the pronunciation video are extracted through a visual feature extraction method based on dimensionality transposition and reshaping, and the three-dimensional video data is converted into two-dimensional image data through the dimensionality transposition operation. While generating a two-dimensional image, the key information in the video is retained, and 3D convolution is used to capture the dynamic movements and subtle changes of the lips, accurately identifying the detailed features of the lip movements of standard pronunciation. On this basis, a joint optimization processing based on multi-level dense layers and Swin Transformer modules is performed to further extract image features, increase the number of channels on the basis of dimensionality reduction, capture deeper and more detailed lip visual information, and provide more accurate lip movement guidance for pronunciation correction.
[0057] Embodiment 3: This embodiment is based on the above embodiment, and step S1 specifically includes the following steps:
[0058] Step S11: Contextual text feature extraction based on the self-attention mechanism is performed on the speech text to obtain text features:
[0059] ;
[0060] Where, Represents the query vector, key vector, and value vector obtained by encoding the speech text through the self-attention mechanism. Represents the key dimensions of the preset, Represents the relative position deviation parameter, That is, it represents the text features;
[0061] Step S12: Multi-head self-attention mechanism optimization, multi-head self-attention mechanism optimization is performed on the text features to obtain optimized text features:
[0062] ;
[0063] ;
[0064] Where, Represents the head of the multi-head self-attention mechanism, 、 、 and Represent the preset projection parameter matrix, Represents the splicing operation, This means optimizing text features;
[0065] Step S13: Positional encoding embedding: positional encoding embedding is performed on the optimized text features corresponding to all words in the speech text to obtain semantic features. The positional encoding is calculated as follows:
[0066] ;
[0067] Where, Represents the position of the word in the speech text, Represents the dimension of the optimized text feature corresponding to the word, represents the preset embedding dimension, Represents the voice text The position code corresponding to the word.
[0068] Embodiment 4: This embodiment is based on the above embodiment, and step S2 specifically includes the following steps:
[0069] Step S21: Convolutional feature extraction, using a visual feature extraction method based on dimension transposition and reshaping to extract convolutional features of the pronunciation video;
[0070] Step S22: Feature dimensionality reduction, using the Linear Embedding layer to reduce the dimensionality of the convolutional features to obtain reduced dimensionality features;
[0071] Step S23: Feature optimization, performing joint optimization processing on the dimensionality reduction features based on the multi-level dense layer and the Swin Transformer module to obtain optimized visual features.
[0072] Embodiment 5: This embodiment is based on the above embodiment, and step S21 specifically includes the following steps:
[0073] Step S211: Video input, input the pronunciation video and obtain its dimensions (BZ, C, L, H, W), which represent the batch size, number of channels, video length, video height and video width of the pronunciation video respectively;
[0074] Step S212: Dimension transposition, swapping the first dimension and the second dimension in the pronunciation video to obtain a transposed video;
[0075] Step S213: reshape the dimension of the transposed video to obtain a reshaped video with a dimension of (BZ×L, C, H, W);
[0076] Step S214: 3D convolution, performing convolution feature extraction based on 3D convolution on the reshaped video to obtain convolution features.
[0077] Example 6: This example is based on the above example, and step S23 specifically includes the following steps:
[0078] Step S231: Based on the dense layer processing of multi-layer 2D convolution, three layers of BN-ReLU-2DConv processing are performed on the dimensionality reduction features to obtain dense optimized convolution features, where the convolution kernel sizes of each 2DConv layer are 1×1, 3×3, and 1×1 respectively;
[0079] Step S232: Swin Transformer module processing, using the Swin Transformer module to process the dimensionality reduction features to obtain efficient optimized convolution features;
[0080] Step S233: feature fusion, patch merging the densely optimized convolution features and the efficient optimized convolution features to obtain primary optimized convolution features;
[0081] Step S234: Multi-layer optimization, repeat steps S32 to S34 three times to obtain optimized visual features.
[0082] By performing the above operations, on the basis of a visual action feature extraction method based on dimensional fusion and feature simplification, a more detailed understanding and recognition of lip movements is guided by semantic features, and the corresponding lip action label description can be accurately selected according to the lip action image recognition and classification results. The multi-head self-attention mechanism is introduced to optimize the semantic features, better capture the association between lip actions and speech text, and add corresponding position codes to the words in the speech text, so that the model can understand the order of the pronunciation actions of the spoken text in the lip action label description, and achieve more accurate lip action guidance.
[0083] Embodiment 7: This embodiment is based on the above embodiment. In the pronunciation analysis module, the extraction of audio features specifically includes the following steps:
[0084] Step Q1: audio preprocessing, denoising, equalization, and downsampling the audio signal;
[0085] Step Q2: Audio framing: Use a Hamming window to divide the audio signal into short time periods, namely audio frames;
[0086] Step Q3: Fourier transform, perform Fourier transform on each audio frame to obtain a spectrum;
[0087] Step Q4: Feature extraction, extracting features based on spectrum information from the spectrum graph to obtain spectrum features;
[0088] Step Q5: Feature dimensionality reduction and normalization: perform dimensionality reduction and normalization on the extracted spectral features to obtain audio features.
[0089] Embodiment 8: This embodiment is based on the above embodiment. A simple example of optimizing the technology of the present invention is as follows:
[0090] (1) Speech recognition and feature extraction
[0091] Speech Recognition Technology:
[0092] Use pre-trained automatic speech recognition (ASR) models (such as Google Speech-to-Text, Wav2Vec2.0);
[0093] Audio feature extraction:
[0094] Extracting acoustic features of speech (such as Mel spectrum, MFCC);
[0095] Analyze the duration, frequency, and energy parameters of phonemes, i.e., spectrogram features;
[0096] (2) Standard pronunciation comparison and error detection (i.e., similarity comparison of the pronunciation correction module)
[0097] Dynamic Time Warping (DTW):
[0098] Compare the time series differences between the user's pronunciation and the standard pronunciation to detect phoneme deviations;
[0099] Phonemic-level scoring:
[0100] Use an acoustic model (such as HMM or DNN-HMM) to calculate the accuracy of each phoneme;
[0101] (3) Correction and feedback generation
[0102] Error location:
[0103] Locate the incorrect phoneme based on its acoustic characteristics (e.g., frequency deviation, duration anomaly);
[0104] Feedback Generation:
[0105] Generate correction suggestions based on error types;
[0106] For example:
[0107] The incorrect phoneme is / i / , and the system prompts: "Raise the tongue position and prolong the pronunciation time;".
[0108] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that includes a list of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus.
[0109] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.
[0110] The present invention and its embodiments are described above. This description is not restrictive. The drawings show only one embodiment of the present invention, and the actual structure is not limited thereto. In short, if a person skilled in the art is inspired by this and, without departing from the purpose of the present invention, designs structures and embodiments similar to this technical solution without inventiveness, they shall fall within the scope of protection of the present invention.
Claims
1. An English spoken pronunciation correction assistance system based on speech recognition, characterized by: It includes data acquisition module, speech recognition module, pronunciation analysis module and pronunciation correction module; The data acquisition module collects audio of the standard pronunciation of spoken English and collects lip pronunciation video of the standard pronunciation of spoken English; The speech recognition module converts the standard pronunciation audio into speech text using a speech recognition model; The pronunciation analysis module extracts lip movement features from the pronunciation video using a text-guided cross-modal feature fusion extraction method, performs activation classification on the lip movement features and adds lip movement description labels, and extracts audio features from the corresponding standard pronunciation audio to obtain a joint analysis result of the standard pronunciation spoken pronunciation, speech text, and lip movement description. The cross-modal feature fusion extraction method based on text guidance specifically includes the following steps: Step S1: Semantic feature extraction based on multi-head self-attention mechanism optimization is performed on the speech text to obtain semantic features, which specifically includes the following steps: Step S11: Contextual text feature extraction based on the self-attention mechanism is performed on the speech text to obtain text features: ; Where, Represents the query vector, key vector, and value vector obtained by encoding the speech text through the self-attention mechanism. Represents the key dimensions of the preset, Represents the relative position deviation parameter, That is, it represents the text features; Step S12: Multi-head self-attention mechanism optimization, multi-head self-attention mechanism optimization is performed on the text features to obtain optimized text features: ; ; Where, Represents the head of the multi-head self-attention mechanism, 、 、 and Represent the preset projection parameter matrix, Represents the splicing operation, This means optimizing text features; Step S13: Positional encoding embedding: positional encoding embedding is performed on the optimized text features corresponding to all words in the speech text to obtain semantic features. The positional encoding is calculated as follows: ; Where, Represents the position of the word in the speech text, Represents the dimension of the optimized text feature corresponding to the word, represents the preset embedding dimension, Represents the voice text The position code corresponding to the word; Step S2: extracting visual action features based on dimension fusion and feature simplification, using a visual action feature extraction method based on dimension fusion and feature simplification to extract optimized visual features of the pronunciation video, specifically including the following steps: Step S21: Convolutional feature extraction, using a visual feature extraction method based on dimensionality transposition and reshaping to extract convolutional features of the pronunciation video, specifically including the following steps: Step S211: Video input, input the pronunciation video and obtain its dimensions (BZ, C, L, H, W), which represent the batch size, number of channels, video length, video height and video width of the pronunciation video respectively; Step S212: Dimension transposition, swapping the first dimension and the second dimension in the pronunciation video to obtain a transposed video; Step S213: reshape the dimension of the transposed video to obtain a reshaped video with a dimension of (BZ×L, C, H, W); Step S214: 3D convolution, performing convolution feature extraction based on 3D convolution on the reshaped video to obtain convolution features; Step S22: Feature dimensionality reduction, using the Linear Embedding layer to reduce the dimensionality of the convolutional features to obtain reduced dimensionality features; Step S23: Feature optimization, which performs a joint optimization process based on a multi-level dense layer and a Swin Transformer module on the dimensionality reduction features to obtain optimized visual features, specifically including the following steps: Step S231: Based on the dense layer processing of multi-layer 2D convolution, three layers of BN-ReLU-2DConv processing are performed on the dimensionality reduction features to obtain dense optimized convolution features, where the convolution kernel sizes of each 2DConv layer are 1×1, 3×3, and 1×1 respectively; Step S232: Swin Transformer module processing, using the Swin Transformer module to process the dimensionality reduction features to obtain efficient optimized convolution features; Step S233: feature fusion, patch merging the densely optimized convolution features and the efficient optimized convolution features to obtain primary optimized convolution features; Step S234: multi-layer optimization, repeating steps S231 to S233 n times to obtain optimized visual features; Step S3: Cross-modal feature fusion, scaling the semantic features to align them with the optimized visual features, and then performing feature fusion to obtain lip movement features; The pronunciation correction module collects the audio that needs to be corrected and the corresponding speech text and sets a correction threshold, extracts the audio features of the audio that needs to be corrected, and compares the similarity with the audio features of the standard pronunciation. When the similarity is higher than the correction threshold, no correction assistance is performed. If the similarity is lower than the correction threshold, the corresponding lip movement description is provided to provide correction guidance.
Citation Information
Patent Citations
Lip language recognition method based on Chinese pronunciation visual features
CN112329581A
English auxiliary pronunciation training method and system based on multi-task learning
CN119993197A