Pronunciation Skill Detection Method, Device, Storage Medium and Electronic Device

Through the artificial intelligence-based pronunciation technique detection method, the phoneme sequence and acoustic feature input detection model are used to solve the problems of subjective judgment and auditory fatigue in artificial hearing detection, and the accuracy of the detection is improved.

CN114170997BActive Publication Date: 2025-06-24IFLYTEK CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202111620731.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-28
Publication Date
2025-06-24
Estimated Expiration
2041-12-28

AI Technical Summary

Technical Problem

In the prior art, when artificial hearing detection is used to evaluate pronunciation skills, there are problems of subjective judgment and auditory fatigue, which affects the accuracy of the detection results.

Method used

Using artificial intelligence-based pronunciation skills detection method, by obtaining the phoneme sequence of the text to be detected and the acoustic characteristics of the text spoken by the speaker, inputting the trained pronunciation skills detection model for detection, to obtain the detection results of whether pronunciation skills are required and whether pronunciation skills are used.

Benefits of technology

It improves the accuracy of pronunciation skills detection and avoids the influence of artificial subjective judgment and auditory fatigue.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114170997B_ABST
    Figure CN114170997B_ABST
Patent Text Reader

Abstract

A method, apparatus, storage medium, and electronic device for detecting pronunciation skills. Among them, the method includes obtaining a text to be detected and converting the text to be detected into a corresponding phoneme sequence; obtaining a to-be-detected audio obtained by a speaker speaking the text to be detected and extracting acoustic features of the to-be-detected audio; inputting the phoneme sequence and the acoustic features into a trained pronunciation skill detection model for pronunciation skill detection processing to obtain a first detection result and a second detection result; wherein, the first detection result is used to characterize whether pronunciation skills need to be used to speak the text to be detected, and the second detection result is used to characterize whether the speaker uses pronunciation skills to speak the text to be detected. This application can improve the accuracy of pronunciation skill detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of speech recognition technology, and particularly to a pronunciation skill detection method, device, storage medium, and electronic device. Background Art

[0002] Currently, for any language, whether it is Chinese or English, spoken language is of utmost importance in mastering these languages. For example, for English learners, spoken pronunciation is often the focus and also the weak point in their English learning process. Whether pronunciation skills such as liaison, aspiration loss, and voicing are accurately used reflects the spoken language ability of English learners. In related technologies, the pronunciation ability of a speaker is usually detected by means of manual listening detection. However, factors such as manual subjective judgment and auditory fatigue will affect the accuracy of the pronunciation skill detection results. Summary of the Invention

[0003] This application provides a pronunciation skill detection method, device, storage medium, and electronic device, which can improve the accuracy of pronunciation skill detection.

[0004] The pronunciation skill detection method provided by this application includes:

[0005] Obtain a text to be detected, and convert the text to be detected into a corresponding phoneme sequence;

[0006] Obtain a to-be-detected audio obtained by a speaker speaking the text to be detected, and extract the acoustic features of the to-be-detected audio;

[0007] Input the phoneme sequence and the acoustic features into a trained pronunciation skill detection model for pronunciation skill detection processing to obtain a first detection result and a second detection result;

[0008] Among them, the first detection result is used to represent whether pronunciation skills need to be used to speak the text to be detected, and the second detection result is used to represent whether the speaker uses pronunciation skills to speak the text to be detected.

[0009] The pronunciation skill detection device provided by this application includes:

[0010] A first acquisition module, configured to obtain a text to be detected and convert the text to be detected into a corresponding phoneme sequence;

[0011] A second acquisition module, configured to obtain a to-be-detected audio obtained by a speaker speaking the text to be detected, and extract the acoustic features of the to-be-detected audio;

[0012] A detection module, configured to input the phoneme sequence and the acoustic features into a trained pronunciation skill detection model for pronunciation skill detection processing to obtain a first detection result and a second detection result;

[0013] Among them, the first detection result is used to characterize whether pronunciation skills need to be adopted to speak the text to be detected, and the second detection result is used to characterize whether the speaker adopts pronunciation skills to speak the text to be detected.

[0014] The storage medium provided by this application stores a computer program, which, when loaded by a processor, executes the steps in the pronunciation skill detection method provided by this application.

[0015] The electronic device provided by this application includes a processor and a memory. The memory stores a computer program, and the processor is used to execute the steps in the pronunciation skill detection method provided by this application by loading the computer program.

[0016] In this application, the text to be detected and the audio to be detected obtained by the speaker speaking the text to be detected are acquired, and the pronunciation of the speaker is detected using the audio to be detected and the text to be detected. Among them, the text to be detected is converted into a corresponding phoneme sequence, and the acoustic features of the audio to be detected are extracted, and then the phoneme sequence and the acoustic features are input into the trained pronunciation skill detection model for pronunciation skill detection processing to obtain the first detection result and the second detection result. The first detection result is used to characterize whether pronunciation skills need to be adopted to speak the text to be detected, and the second detection result is used to characterize whether the speaker adopts pronunciation skills to speak the text to be detected. Compared with the related art, this application uses an artificial intelligence-based pronunciation skill detection method to replace the traditional manual hearing detection, which can avoid subjective judgment and auditory fatigue of humans, thereby improving the accuracy of pronunciation skill detection. Description of the Drawings

[0017] In order to more clearly illustrate the technical solutions in the embodiments of this application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative efforts.

[0018] Figure 1 It is a schematic diagram of the scenario of the pronunciation skill detection system provided by the embodiment of this application.

[0019] Figure 2 It is a schematic flowchart of the pronunciation skill detection method provided by the embodiment of this application.

[0020] Figure 3 It is an example diagram of extracting acoustic features in the embodiment of this application.

[0021] Figure 4 It is a structural block diagram of the pronunciation skill detection model provided by the embodiment of this application.

[0022] Figure 5It is a structural block diagram of a phoneme feature extraction network in a pronunciation skill detection model.

[0023] Figure 6 It is a structural block diagram of a phoneme feature extraction sub-module inside a phoneme feature module in a phoneme feature extraction network.

[0024] Figure 7 It is another structural block diagram of a phoneme feature extraction network in a pronunciation skill detection model.

[0025] Figure 8 It is a structural block diagram of a feature encoding module inside an acoustic feature enhancement network in a pronunciation skill detection model.

[0026] Figure 9 It is a refined structural block diagram of an acoustic feature enhancement module inside an acoustic feature enhancement network in a pronunciation skill detection model.

[0027] Figure 10 It is a structural block diagram of a feature fusion network in a pronunciation skill detection model.

[0028] Figure 11 It is a structural block diagram of a first pronunciation skill detection network in a pronunciation skill detection model.

[0029] Figure 12 It is a structural block diagram of a branch detection network inside a second pronunciation skill detection network in a pronunciation skill detection model.

[0030] Figure 13 It is a structural block diagram of a pronunciation skill detection device provided by an embodiment of the present application.

[0031] Figure 14 It is a structural block diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0032] It should be noted that the principle of the present application is illustrated by way of implementation in a suitable computing environment. The following description is based on the specific embodiments of the present application illustrated, and it should not be regarded as limiting other specific embodiments of the present application not detailed herein. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative efforts fall within the scope of protection of the present application.

[0033] The relational terms such as first and second involved in the following embodiments of the present application are only used to distinguish one object or operation from another object or operation, and do not limit the actual sequential relationship between these objects or operations. In the description of the embodiments of the present application, the meaning of "a plurality of" is two or more unless otherwise specifically defined.

[0034] Artificial Intelligence (AI) is a theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines to enable machines to have the functions of perception, reasoning, and decision-making.

[0035] Artificial intelligence technology is an interdisciplinary subject with a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. The software technology of artificial intelligence mainly includes Machine Learning (ML) technology. Among them, Deep Learning (DL) is a new research direction in machine learning, which is introduced into machine learning to make it closer to the original goal, that is, artificial intelligence. Currently, deep learning is mainly applied in fields such as computer vision and natural language processing.

[0036] Deep learning is to learn the internal laws and representation levels of sample data, and the information obtained during these learning processes is very helpful for the interpretation of data such as text, images, and sounds. Using deep learning technology and the corresponding training data sets, network models with different functions can be trained. For example, a deep learning network for gender classification can be trained based on one training data set, and a deep learning network for image optimization can be trained based on another training data set, etc.

[0037] In order to improve the efficiency of pronunciation skill detection, this application introduces deep learning into pronunciation skill detection, and correspondingly provides a pronunciation skill detection method, a pronunciation skill detection device, a storage medium, and an electronic device. Among them, the pronunciation skill detection method can be executed by an electronic device.

[0038] Please refer to Figure 1 , this application also provides a pronunciation skill detection system, such as Figure 1As shown in the figure, the pronunciation skill detection system includes an electronic device 100. For example, the electronic device can obtain the text to be detected for pronunciation skill detection and convert the text to be detected into a corresponding phoneme sequence. When the electronic device is also equipped with a microphone, it can collect audio during the speaker's utterance of the text to be detected, so as to obtain the audio to be detected obtained by the speaker's utterance of the text to be detected, and extract the acoustic features of the audio to be detected. After that, the obtained phoneme sequence and acoustic features are further input into the trained pronunciation skill detection model for pronunciation skill detection processing to obtain a first detection result and a second detection result. Among them, the first detection result is used to represent whether the text to be detected needs to be pronounced using pronunciation skills, and the second detection result is used to represent whether the speaker pronounces the text to be detected using pronunciation skills.

[0039] The electronic device 100 can be any device equipped with a processor and having processing capabilities, such as mobile electronic devices equipped with a processor like smartphones, tablets, personal digital assistants, laptops, etc., or fixed electronic devices equipped with a processor like desktop computers, TVs, servers, etc.

[0040] In addition, as Figure 1 shown, the pronunciation skill detection system may further include a storage device 200 for storing data, including but not limited to the original data, intermediate data, and result data obtained during the pronunciation skill detection process. For example, the electronic device 100 can store the text to be detected, the audio to be detected, the phoneme sequence obtained by converting the text to be detected, the acoustic features extracted from the audio to be detected, and the first detection result and the second detection result output by the pronunciation skill detection model into the storage device 200.

[0041] It should be noted that Figure 1 the scenario schematic diagram of the pronunciation skill detection system shown is only an example. The pronunciation skill detection system and scenario described in the embodiments of the present application are for more clearly illustrating the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided by the embodiments of the present application. Those skilled in the art know that with the evolution of the pronunciation skill detection system and the emergence of new business scenarios, the technical solutions provided by the embodiments of the present application are equally applicable to similar technical problems.

[0042] Please refer to Figure 2 , Figure 2 which is the flowchart of the pronunciation skill detection method provided by the embodiments of the present application. As Figure 2 shown, the process of the pronunciation skill detection method provided by the embodiments of the present application can be as follows:

[0043] In S310, obtain the text to be detected and convert the text to be detected into a corresponding phoneme sequence.

[0044] Among them, the text to be detected refers to the text used for pronunciation skill detection. The pronunciation skill detection includes detecting whether pronunciation skills need to be adopted to pronounce the text to be detected, and detecting whether the speaker pronounces the text to be detected with pronunciation skills.

[0045] It should be noted that different languages such as Chinese and English have their own pronunciation skills. Taking English as an example, there are pronunciation skills such as liaison, loss of plosion, and aspiration.

[0046] Among them, liaison means that the initial and final phonemes of two adjacent words are pronounced together naturally without a pause in the middle.

[0047] Loss of plosion means that when two plosive sounds (such as p, d, t, k, g) are adjacent, the previous plosive sound only forms an obstruction by making the pronunciation mouth shape according to its pronunciation position, but does not explode. After a slight pause, the following consonant is pronounced. The previous plosive sound is called loss of plosion, such as goo(d)bye.

[0048] Aspiration means that when the sound before a voiceless consonant is / s / , the voiceless consonant has its corresponding voiced consonant, and there is also a vowel after the voiceless consonant. At this time, the voiceless consonant is pronounced as its corresponding voiced consonant. Taking speak as an example, before the voiceless consonant / p / there is the sound / s / , the corresponding voiced consonant of / p / is / b / , and there is also a vowel / i: / after / p / . At this time, the original / spi:k / is pronounced as / sbi:k / .

[0049] As above, the pronunciation skill detection method provided by this application can be used for pronunciation skill detection of any language. Correspondingly, according to the requirements of actual pronunciation skill detection, the text to be detected can be a text in any language. After the electronic device obtains the text to be detected, it further converts the obtained text to be detected into a corresponding phoneme sequence. For example, the electronic device can convert the obtained text to be detected into a corresponding phoneme sequence according to a pronunciation dictionary.

[0050] In an optional embodiment, converting the text to be detected into a corresponding phoneme sequence includes:

[0051] Removing the non - pronounced text units in the text to be detected to obtain a new text to be detected;

[0052] Converting each text unit in the new text to be detected into a corresponding phoneme unit to obtain a phoneme sequence.

[0053] It can be understood that for any text, when pronouncing the text, not all text units in it need to be pronounced. For example, for punctuation marks in the text, they do not need to be pronounced.

[0054] Therefore, in order to exclude the interference of unpronounced text units and improve the accuracy of pronunciation skill detection, when converting the text to be detected into a corresponding phoneme sequence, the electronic device first removes the unpronounced text units (such as punctuation marks, emojis, etc.) in the text to be detected to obtain a new text to be detected, and then converts each text unit in the new text to be detected into a corresponding phoneme sequence according to the pronunciation dictionary.

[0055] For example, when performing pronunciation skill detection of English, the electronic device obtains the text to be detected in English as "Please turn on the light.", and the text unit - punctuation mark "." in this text to be detected is an unpronounced text unit. After removing this unpronounced text unit ".", the new text to be detected obtained is "Please turn on thelight", and further converts each text unit in this new text to be detected into a corresponding phoneme unit according to the pronunciation dictionary to obtain a phoneme sequence.

[0056] In addition, in order to more clearly represent the phoneme sequence, the electronic device can also add a start flag and an end flag before and after the phoneme sequence respectively, and represent the start of the phoneme sequence through the start flag and the end of the phoneme sequence through the end flag. There is no specific limitation on what the start flag and the end flag are specifically configured as, and those skilled in the art can configure them according to actual needs.

[0057] For example, the start flag can be configured as " <bos>”, and the configuration end flag is " <eos>”, the electronic device changes to after adding a start flag and an end flag to the above phoneme sequence

[0058] In S320, the to-be-detected audio obtained when the speaker speaks the to-be-detected text is acquired, and the acoustic features of the to-be-detected audio are extracted.

[0059] In this embodiment, in addition to converting the to-be-detected text into a corresponding phoneme sequence, the electronic device also acquires the to-be-detected audio obtained when the speaker speaks the to-be-detected text. Here, no specific limitation is imposed on the data format of the to-be-detected audio, and it can be configured by those skilled in the art according to actual detection needs.

[0060] Among them, the speaker can be a real person or a virtual person.

[0061] For example, when the speaker is a real person, the electronic device can collect the voice of the to-be-detected text spoken by the real person through a configured audio collection device (which can be a built-in audio collection device or an external audio collection device), and use the collected audio as the to-be-detected audio; in addition, the electronic device can also obtain from other electronic devices the to-be-detected audio that has been collected by other electronic devices and spoken by a real person for the to-be-detected text. Correspondingly, using the to-be-detected audio obtained at this time, the electronic device can apply the pronunciation skill detection method provided in this application to detect the pronunciation ability of the real person.

[0062] Also for example, when the speaker is a virtual person, such as an artificial intelligence-based voice synthesis software, the electronic device can directly input the to-be-detected text into the voice synthesis software, and the voice synthesis software performs voice synthesis and outputs the synthesized audio, and uses this audio as the to-be-detected audio. Correspondingly, using the to-be-detected audio obtained at this time, the electronic device can be used to detect the voice synthesis ability of the voice synthesis software by the pronunciation skill detection method provided in this application.

[0063] As described above, after the electronic device acquires the to-be-detected audio obtained when the speaker speaks the to-be-detected text, it also extracts the acoustic features of the to-be-detected audio. Among them, the acoustic features refer to the physical quantities representing the acoustic characteristics of speech, and are also the general term for the acoustic manifestations of various elements of sound, such as the energy concentration area, formant frequency, formant intensity, and bandwidth representing timbre, as well as the duration, fundamental frequency, average speech power, etc. representing the prosodic characteristics of speech.

[0064] In an optional embodiment, in order to further improve the accuracy of pronunciation skill detection, extracting the acoustic features of the to-be-detected audio includes:

[0065] extracting the Filterbank features, fundamental frequency features, and energy features of the to-be-detected audio;

[0066] Fuse the Filterbank features, fundamental frequency features, and energy features to obtain acoustic features.

[0067] In this embodiment, the Filterbank features, fundamental frequency features, and energy features are used as acoustic features related to pronunciation skills. Correspondingly, when extracting acoustic features for pronunciation skill detection, the electronic device extracts the Filterbank features, fundamental frequency features, and energy features of the audio to be detected. Among them, the dimension of the extracted Filterbank features is not specifically limited here and can be configured by those skilled in the art according to actual needs. For example, in this embodiment, the electronic device can extract 40-dimensional Filterbank features of the audio to be detected.

[0068] As described above, after extracting the Filterbank features, fundamental frequency features, and energy features of the audio to be detected, the electronic device further fuses the Filterbank features, fundamental frequency features, and energy features according to the configured fusion strategy to obtain a fusion feature, and uses this fusion feature as the acoustic feature for pronunciation skill detection. The configuration of the fusion strategy here is not specifically limited and can be configured by those skilled in the art according to actual needs.

[0069] For example, please refer to Figure 3 , the fusion strategy configured in this embodiment is to splice the Filterbank features, fundamental frequency features, and energy features in the time dimension to obtain the acoustic features for pronunciation skill detection.

[0070] In S330, the phoneme sequence and the acoustic features are input into the trained pronunciation skill detection model for pronunciation skill detection processing to obtain a first detection result and a second detection result.

[0071] It should be noted that for different languages, the corresponding pronunciation skill detection models are pre-trained in this application. For example, for Chinese, a pronunciation skill detection model for detecting pronunciation skills in Chinese is pre-trained, and for English, a pronunciation skill detection model for detecting pronunciation skills in English is pre-trained. The structure and training method of the pronunciation skill detection model are not specifically limited here and can be selected by those skilled in the art according to actual needs.

[0072] Among them, the pronunciation skill detection model is configured to take the acoustic features of the audio to be detected from the speaker speaking the text to be detected, and the phoneme sequence from the text to be detected as inputs, and correspondingly output the detection result for characterizing whether pronunciation skills need to be used to speak the text to be detected, and the detection result for characterizing whether the speaker uses pronunciation skills to speak the text to be detected.

[0073] Correspondingly, in this embodiment, after obtaining the above phoneme sequence and acoustic features, the electronic device inputs the obtained phoneme sequence and acoustic features into a trained pronunciation skill detection model that matches the language of the text to be detected for pronunciation skill detection processing, and obtains a first detection result and a second detection result output by the pronunciation skill detection model. Among them, the first detection result is used to indicate whether pronunciation skills need to be used to pronounce the text to be detected, and the second detection result is used to indicate whether the speaker uses pronunciation skills to pronounce the text to be detected.

[0074] Among them, taking the above text to be detected "Please turn on the light” as an example, according to expert knowledge, the phonemes "n” and "ɑ” are a connected phoneme combination, and turn and on need to be pronounced in a connected way. For the phoneme sequence and acoustic features corresponding to this text to be detected, the first detection result output by the pronunciation skill detection model will indicate that the pronunciation skill of "connected pronunciation” needs to be used to output the text to be detected, depending on whether the speaker uses the pronunciation skill of "connected pronunciation” to pronounce the text to be detected, and the pronunciation skill detection model will output the corresponding second detection result.

[0075] In addition, the phoneme sequence can be input as the original phoneme sequence, or after digital encoding, the phoneme sequence is converted into a digital form of the phoneme sequence, that is, the corresponding phoneme is represented by a number in the converted phoneme sequence. Correspondingly, if the phoneme sequence is input in digital form, then when training the pronunciation skill detection model, it is necessary to train with digital-form phoneme sequence samples.

[0076] In an optional embodiment, the pronunciation skill detection model includes a phoneme feature extraction network, an acoustic feature enhancement network, a feature fusion network, a first pronunciation skill detection network, and a second pronunciation skill detection network. Inputting the phoneme sequence and acoustic features into the trained pronunciation skill detection model for pronunciation skill detection processing to obtain the first detection result and the second detection result includes:

[0077] Input the phoneme sequence into the phoneme feature extraction network for feature extraction processing to obtain a phoneme feature matrix;

[0078] Input the phoneme feature matrix into the first pronunciation skill detection network for pronunciation skill detection processing to obtain the first detection result;

[0079] If the first detection result indicates that pronunciation skills need to be used to pronounce the text to be detected, then input the acoustic features into the acoustic feature enhancement network for feature enhancement processing to obtain an enhanced acoustic feature matrix;

[0080] Input the enhanced acoustic feature matrix and the phoneme feature matrix into the feature fusion network for feature fusion processing to obtain a fusion feature matrix;

[0081] Input the fused feature matrix into the second pronunciation skill detection network for pronunciation skill detection processing to obtain a second detection result.

[0082] Please refer to Figure 4 , the pronunciation skill detection model provided in this embodiment consists of five major parts, namely a phoneme feature extraction network, an acoustic feature enhancement network, a feature fusion network, a first pronunciation skill detection network, and a second pronunciation skill detection network.

[0083] Among them, the phoneme feature extraction network is configured to extract features from the input phoneme sequence to obtain a phoneme feature matrix reflecting the relationship between phonemes in the phoneme sequence.

[0084] The acoustic feature enhancement network is configured to perform feature enhancement processing on the input acoustic features to enhance the features more relevant to pronunciation skills and obtain an enhanced acoustic feature matrix.

[0085] The feature fusion network is configured to perform feature fusion on the input phoneme feature matrix and the enhanced acoustic feature matrix without performing interaction information between phonemes and acoustic features to obtain a fused feature matrix.

[0086] The first pronunciation skill detection network is configured to perform pronunciation skill detection processing on the input phoneme feature matrix and output a first detection result for characterizing whether pronunciation skills need to be used to say the text to be detected.

[0087] The second pronunciation skill detection network is configured to perform pronunciation skill detection processing on the input fused feature matrix and output a second detection result for characterizing whether the speaker uses pronunciation skills to say the text to be detected.

[0088] Correspondingly, in this embodiment, when inputting the phoneme sequence and acoustic features into the trained pronunciation skill detection model for pronunciation skill detection processing, the electronic device can perform feature extraction processing on the phoneme sequence through the phoneme feature extraction network to obtain a phoneme feature matrix, and then input the phoneme feature matrix into the first pronunciation skill detection network for pronunciation skill detection processing to obtain a first detection result.

[0089] At the same time, the electronic device inputs the acoustic features into the acoustic feature enhancement network for feature enhancement processing to obtain an enhanced acoustic feature matrix, inputs the enhanced acoustic feature and the phoneme feature matrix into the feature fusion network for feature fusion processing to obtain a fused feature matrix, and inputs the fused feature matrix obtained by fusion into the second pronunciation skill detection network for pronunciation skill detection processing to obtain a second detection result.

[0090] In addition, the electronic device can also determine whether to output the second detection result according to the first detection result. Among them, after the electronic device obtains the first detection result and the second detection result obtained by performing pronunciation skill detection processing using the pronunciation skill detection model, it determines whether it is necessary to pronounce the text to be detected using pronunciation skills according to the first detection result. If it is determined that it is necessary to pronounce the text to be detected using pronunciation skills, the electronic device outputs both the first detection result and the second detection result at the same time. At this time, the first detection result represents that it is necessary to pronounce the text to be detected using pronunciation skills, and the second detection result represents whether the speaker pronounces the text to be detected using pronunciation skills. If it is determined that it is not necessary to pronounce the text to be detected using pronunciation skills, the electronic device can discard the second detection result and only output the first detection result.

[0091] In other embodiments, if the first detection result indicates that it is necessary to pronounce the text to be detected using pronunciation skills, the electronic device inputs the acoustic features into an acoustic feature enhancement network for feature enhancement processing to obtain an enhanced acoustic feature matrix. Then, the electronic device further inputs the enhanced acoustic features and the phoneme feature matrix into a feature fusion network for feature fusion processing to obtain a fused feature matrix. Finally, the electronic device inputs the fused feature matrix obtained by fusion into a second pronunciation skill detection network for pronunciation skill detection processing to obtain the second detection result.

[0092] In addition, if the first detection result indicates that it is not necessary to pronounce the text to be detected using pronunciation skills, there is no need to perform further pronunciation skill detection. At this time, the electronic device no longer uses the acoustic features for pronunciation skill detection and can only output the first detection result.

[0093] In an optional embodiment, the phoneme feature extraction network includes a phoneme embedding module and a phoneme feature extraction module. Inputting the phoneme sequence into the phoneme feature extraction network for feature extraction processing to obtain a phoneme feature matrix includes:

[0094] Inputting the phoneme sequence into the phoneme embedding module for embedding processing to obtain a phoneme vector matrix;

[0095] Inputting the phoneme vector matrix into the phoneme feature extraction module for feature extraction processing to obtain a phoneme feature matrix.

[0096] Please refer to Figure 5 , in this embodiment, the phoneme feature extraction network consists of two parts, namely a phoneme embedding module and a phoneme feature extraction module. Among them, the phoneme embedding module is configured to perform embedding processing on the input phoneme sequence to vectorize it and obtain a phoneme vector matrix; the phoneme feature extraction module is configured to perform feature extraction on the input phoneme vector matrix to obtain a phoneme feature matrix reflecting the mutual relationship between phonemes in the phoneme sequence.

[0097] Correspondingly, in this embodiment, when inputting a phoneme sequence into a phoneme feature extraction network for feature extraction processing, the electronic device first inputs the phoneme sequence into a phoneme embedding module for embedding processing to obtain a phoneme vector matrix, and then inputs the phoneme vector matrix into a phoneme feature extraction module for feature extraction processing to obtain a phoneme feature matrix.

[0098] Among them, the phoneme feature extraction module includes at least 1 phoneme feature extraction sub-module. Inputting the phoneme vector matrix into the phoneme feature extraction module for feature extraction processing to obtain a phoneme feature matrix includes:

[0099] When there is 1 phoneme feature extraction sub-module, input the phoneme vector matrix into the phoneme feature extraction sub-module for feature extraction processing to obtain a phoneme feature matrix; or,

[0100] When there are N phoneme feature extraction sub-modules, input the phoneme vector matrix into the N phoneme feature extraction sub-modules in sequence for feature extraction processing to obtain a phoneme feature matrix, where N is an integer greater than 1. Here, the value of N is not specifically limited and can be configured by those skilled in the art according to actual needs. For example, N can be configured to 2.

[0101] In this embodiment, the audio feature extraction module can be composed of 1 phoneme feature extraction sub-module, or can be composed of N phoneme feature extraction sub-modules connected in sequence. Among them, when the audio feature extraction module is composed of N phoneme feature extraction sub-modules, each audio extraction sub-module performs the same feature extraction processing. The following takes the feature extraction processing process of 1 audio feature extraction sub-module as an example for illustration.

[0102] Please refer to Figure 6 , the audio feature extraction sub-module is composed of 3 sub-layers, namely the first matrix conversion layer, the first multi-head attention layer, and the first matrix fusion layer. Among them,

[0103] The first matrix conversion layer is configured to perform matrix conversion processing on the input matrix, and convert the input matrix into a query matrix, a key matrix, and a value matrix respectively;

[0104] The first multi-head attention layer is configured to perform attention enhancement processing on the input query matrix, key matrix, and value matrix to obtain an attention enhancement matrix;

[0105] The first matrix fusion layer is configured to perform matrix fusion processing on the input matrix of the first matrix conversion layer and the output matrix of the first multi-head attention layer to obtain a fusion matrix.

[0106] Correspondingly, when there is 1 phoneme feature extraction sub-module, the electronic device can extract a phoneme feature matrix in the following manner:

[0107] Input the phoneme vector matrix into the first matrix conversion layer for matrix conversion processing to obtain a query matrix, a key matrix, and a value matrix, denoted as the first query matrix, the first key matrix, and the first value matrix respectively;

[0108] Input the first query matrix, the first key matrix, and the first value matrix into the first multi-head attention layer for attention enhancement processing to obtain an attention enhancement matrix, denoted as the first attention enhancement matrix;

[0109] Input the first attention enhancement matrix and the phoneme vector matrix into the first matrix fusion layer for matrix fusion processing to obtain a fusion matrix, denoted as the phoneme feature matrix.

[0110] It can be understood that when there are N phoneme feature extraction sub-modules, only need to input the phoneme vector matrix into the first phoneme feature extraction sub-module among the N sequentially connected phoneme feature extraction sub-modules, and the N phoneme feature sub-modules perform feature extraction processing in the above manner in sequence, and use the fusion feature output by the Nth phoneme feature sub-module as the phoneme feature matrix.

[0111] Among them, there is no specific limitation on how the first matrix fusion layer performs matrix fusion in this embodiment, and it can be configured by those skilled in the art according to actual needs.

[0112] For example, the first matrix fusion layer may include two sub-layers, namely an addition layer and a layer normalization layer. When performing matrix fusion, first add the two input matrices through the addition layer to obtain a sum value matrix, and then perform layer normalization processing on the sum value matrix through the layer normalization layer to obtain a fusion matrix.

[0113] In an optional embodiment, please refer to Figure 7 , the phoneme feature extraction network further includes a first position encoding module and a second matrix fusion layer. Before inputting the phoneme vector matrix into the phoneme feature extraction module for feature extraction processing to obtain the phoneme feature matrix, it further includes:

[0114] Input the phoneme vector matrix into the first position encoding module for position encoding processing to obtain a first position encoding matrix;

[0115] Input the first position encoding matrix and the phoneme vector matrix into the second matrix fusion layer for matrix fusion processing to obtain a phoneme position fusion matrix;

[0116] Inputting the phoneme vector matrix into the phoneme feature extraction module for feature extraction processing to obtain the phoneme feature matrix includes:

[0117] Input the phoneme position fusion matrix into the phoneme feature extraction module for feature extraction processing to obtain the phoneme feature matrix.

[0118] In this embodiment, in order to further improve the accuracy of pronunciation skill detection, instead of directly inputting the original phoneme vector matrix into the phoneme feature extraction module for feature extraction, the phoneme vector matrix is first subjected to positional encoding, and then the phoneme vector matrix carrying positional information is input into the phoneme feature extraction module for feature extraction.

[0119] Specifically, the electronic device first inputs the phoneme vector matrix into the first positional encoding module for positional encoding processing to obtain a positional encoding matrix, denoted as the first positional encoding matrix. This first positional encoding matrix represents the positional information of each matrix unit in the phoneme vector matrix, which can be relative positional information or absolute positional information.

[0120] After obtaining the above-mentioned first positional encoding matrix, the electronic device then inputs the first positional encoding matrix and the phoneme vector matrix into the second matrix fusion layer for matrix fusion processing to obtain a fusion matrix, denoted as the phoneme-position fusion matrix. Subsequently, the electronic device further inputs the phoneme-position fusion matrix carrying positional information into the phoneme feature extraction module for feature extraction processing to obtain a phoneme feature matrix. Regarding how the phoneme feature extraction module performs feature extraction processing, please refer to the relevant descriptions in the above embodiments for details and will not be elaborated here.

[0121] In addition, it should be noted that there is no specific limitation on the matrix fusion method of the second matrix fusion layer in this embodiment, which can be configured by those skilled in the art according to actual needs. For example, the second matrix fusion layer is configured to perform an addition operation on the two input matrices and output the sum matrix as the fusion matrix.

[0122] In an alternative embodiment, the acoustic feature enhancement network includes a feature encoding module and at least one acoustic feature enhancement module. When the acoustic features are input into the acoustic feature enhancement network for feature enhancement processing to obtain an enhanced acoustic feature matrix, the process includes:

[0123] Input the acoustic features into the feature encoding module for feature encoding processing to obtain an acoustic feature matrix;

[0124] When there is one acoustic feature enhancement module, input the acoustic feature matrix into the acoustic feature enhancement module for feature enhancement processing to obtain an enhanced acoustic feature matrix; or,

[0125] When there are M acoustic feature enhancement modules (M is an integer greater than 1), input the acoustic feature matrix into the M acoustic feature enhancement modules in sequence for feature enhancement processing to obtain an enhanced acoustic feature matrix. There is no specific limitation on the value of M here, which can be configured by those skilled in the art according to actual needs. For example, M can be configured to 4.

[0126] It should be noted that the Filterbank features, fundamental frequency features, and energy features obtained in the above embodiments are all presented in the form of feature maps. Correspondingly, the acoustic features obtained by fusing the Filterbank features, fundamental frequency features, and energy features are also presented in the form of feature maps.

[0127] In order to effectively perform feature enhancement processing on the acoustic features, in this embodiment, the acoustic feature enhancement network consists of 1 feature encoding module and at least 1 acoustic feature enhancement module. Among them, the feature encoding module is configured to perform encoding processing on the acoustic features, compress the feature dimensions, and obtain the corresponding acoustic feature matrix; the acoustic feature enhancement module is configured to perform feature enhancement processing on the input acoustic feature matrix, enhance the features related to pronunciation skills therein, and obtain the enhanced acoustic feature matrix.

[0128] Correspondingly, when inputting the acoustic features into the acoustic feature enhancement network for feature enhancement processing, the electronic device first inputs the acoustic features into the feature encoding module for feature encoding processing to obtain the acoustic feature matrix.

[0129] Please refer to Figure 8 , the feature encoding module consists of 4 sub-layers, including the first convolutional layer, the first pooling layer, the second convolutional layer, and the second pooling layer. Among them,

[0130] The first convolutional layer is configured to perform convolutional processing on the input feature map to obtain the corresponding convolutional result;

[0131] The first pooling layer is configured to perform pooling processing on the convolutional result output by the first convolutional layer to obtain the corresponding pooling result;

[0132] The second convolutional layer is configured to perform convolutional processing on the pooling result output by the first pooling layer to obtain the corresponding convolutional result;

[0133] The second pooling layer is configured to perform pooling processing on the convolutional result output by the second convolutional layer to obtain the feature matrix of the corresponding feature map.

[0134] Correspondingly, the electronic device can encode the input acoustic features into the feature encoding module in the following manner:

[0135] Input the acoustic features into the first convolutional layer for convolutional processing to obtain the convolutional result, denoted as the first convolutional result;

[0136] Input the first convolutional result into the first pooling layer for pooling processing to obtain the pooling result, denoted as the first pooling result;

[0137] Input the first pooling result into the second convolutional layer for convolutional processing to obtain the convolutional result, denoted as the second convolutional result;

[0138] The second convolution result is input into the second pooling layer for pooling processing to obtain an acoustic feature matrix.

[0139] It should be noted that in this embodiment, there are no specific restrictions on the convolution kernel size, stride, and padding size of the first convolution layer and the second convolution layer, nor on the pooling kernel size and stride of the first pooling layer and the second pooling layer. Those skilled in the art can configure them according to actual needs.

[0140] For example, in this embodiment, the convolution kernel size of the first convolution layer is configured as [3, 3], the stride is [1, 1], and the padding size is [1, 1]; the convolution kernel size of the second convolution layer is configured as [3, 3], the stride is [1, 1], and the padding size is [1, 1]; the pooling type of the first pooling layer is configured as max pooling, the pooling kernel size is [2, 2], and the stride is [1, 1]; the pooling type of the second pooling layer is configured as max pooling, the pooling kernel size is [2, 2], and the stride is [1, 1].

[0141] Further, when there is 1 acoustic feature enhancement module, the electronic device directly inputs the acoustic feature matrix into this acoustic feature enhancement module for feature enhancement processing to obtain an enhanced acoustic feature matrix.

[0142] The following takes the feature enhancement processing process of 1 acoustic feature enhancement module as an example for illustration.

[0143] Please refer to Figure 9 , the acoustic feature enhancement module consists of 6 sub - layers, namely the second matrix transformation layer, the second multi - head attention layer, the third matrix fusion layer, the third convolution layer, the transposed convolution layer, and the fourth matrix fusion layer. Among them,

[0144] The second matrix transformation layer is configured to perform matrix transformation processing on the input matrix, and transform the input matrix into a query matrix, a key matrix, and a value matrix respectively;

[0145] The second multi - head attention layer is configured to perform attention enhancement processing on the input query matrix, key matrix, and value matrix to obtain an attention - enhanced matrix;

[0146] The third matrix fusion layer is configured to perform matrix fusion processing on the input matrix of the second matrix transformation layer and the output matrix of the second multi - head attention layer to obtain a fusion matrix;

[0147] The third convolution layer is configured to perform convolution processing on the input fusion matrix to obtain a convolution result;

[0148] The transposed convolution layer is configured to perform transposed convolution processing on the input convolution result to obtain a transposed convolution result in matrix form;

[0149] The fourth matrix fusion layer is configured to perform matrix fusion processing on the fusion matrix output by the third matrix fusion layer and the transposed convolution result output by the transposed convolution layer to obtain a fusion matrix.

[0150] Correspondingly, when there is 1 acoustic feature enhancement module, the electronic device can enhance the obtained enhanced acoustic feature matrix in the following manner:

[0151] Input the acoustic feature matrix into the second matrix conversion layer for matrix conversion processing to obtain a query matrix, a key matrix, and a value matrix, denoted as the second query matrix, the second key matrix, and the second value matrix respectively;

[0152] Input the second query matrix, the second key matrix, and the second value matrix into the second multi-head attention layer for attention enhancement processing to obtain an attention enhancement matrix, denoted as the second attention enhancement matrix;

[0153] Input the second attention enhancement matrix and the acoustic feature matrix into the third matrix fusion layer for matrix fusion processing to obtain a fusion matrix, denoted as the acoustic fusion matrix;

[0154] Input the acoustic fusion matrix into the third convolution layer for convolution processing to obtain a third convolution result;

[0155] Input the third convolution result into the transposed convolution layer for transposed convolution processing to obtain a transposed convolution result in matrix form;

[0156] Input the acoustic fusion matrix and the transposed convolution result into the fourth matrix fusion layer for matrix fusion processing to obtain a fusion matrix, denoted as the enhanced acoustic feature matrix.

[0157] Among them, regarding how the third matrix fusion layer and the fourth matrix fusion layer perform matrix fusion, no specific limitation is made in this embodiment, and it can be configured by those skilled in the art according to actual needs.

[0158] For example, the third matrix fusion layer and the fourth matrix fusion layer have the same structure, both including two sub-layers, namely an addition layer and a layer normalization layer. When performing matrix fusion, first add the two input matrices through the addition layer to obtain a sum value matrix, and then perform layer normalization processing on the sum value matrix through the layer normalization layer to obtain a fusion matrix.

[0159] It should be noted that when there are M acoustic feature enhancement modules, each acoustic feature enhancement module performs the same feature enhancement processing. Only need to input the acoustic feature matrix into the first acoustic feature enhancement module among the M sequentially connected acoustic feature enhancement modules, and the M acoustic feature enhancement modules perform feature enhancement processing in the above manner in sequence, and use the fusion feature output by the Mth acoustic feature enhancement module as the enhanced acoustic feature matrix.

[0160] In addition, in this embodiment, there are no specific restrictions on the configuration of the convolution kernel size, stride size, and padding size in the above third convolution layer and transposed convolution layer, and those skilled in the art can take values according to actual needs.

[0161] It can be understood that by enhancing the convolution process and the transposed convolution process during the enhancement of the acoustic features in this embodiment, the acoustic features in the form of feature maps can be enhanced more effectively, and finally the features more relevant to the pronunciation skills can be extracted.

[0162] In an optional embodiment, the acoustic feature enhancement network further includes a second position encoding module and a fifth matrix fusion layer. Before inputting the acoustic feature matrix into the acoustic feature enhancement module for feature enhancement processing to obtain an enhanced acoustic feature matrix, it further includes:

[0163] Inputting the acoustic feature matrix into the second position encoding module for position encoding processing to obtain a second position encoding matrix;

[0164] Inputting the second position encoding matrix and the acoustic feature matrix into the fifth matrix fusion layer for matrix fusion processing to obtain an acoustic position fusion matrix;

[0165] Inputting the acoustic feature matrix into the acoustic feature enhancement module for feature enhancement processing to obtain an enhanced acoustic feature matrix includes:

[0166] Inputting the acoustic position fusion matrix into the acoustic feature enhancement module for feature enhancement processing to obtain an enhanced acoustic feature matrix.

[0167] In this embodiment, in order to further improve the accuracy of pronunciation skill detection, the original acoustic feature matrix is not input into the acoustic feature enhancement module for feature enhancement. Instead, after position encoding, the acoustic feature matrix carrying position information is input into the acoustic feature enhancement module for feature enhancement.

[0168] Among them, the electronic device first inputs the acoustic feature matrix into the second position encoding module for position encoding processing to obtain a position encoding matrix, denoted as the second position encoding matrix. This second position encoding matrix represents the position information of each matrix unit in the acoustic feature matrix, which can be relative position information or absolute position information.

[0169] After obtaining the second position encoding matrix as described above, the electronic device inputs the second position encoding matrix and the acoustic feature matrix into the fifth matrix fusion layer for matrix fusion processing to obtain a fusion matrix, denoted as the acoustic position fusion matrix. After that, the electronic device further inputs the acoustic position fusion matrix carrying position information into the acoustic feature enhancement module for feature enhancement processing to obtain an enhanced acoustic feature matrix. For how the acoustic feature enhancement module performs feature enhancement processing, please refer to the relevant descriptions in the above embodiments and will not be elaborated here.

[0170] In addition, it should also be noted that the matrix fusion method of the fifth matrix fusion layer in this embodiment is not specifically limited and can be configured by those skilled in the art according to actual needs. For example, the fifth matrix fusion layer is configured to perform an addition process on the two input matrices and output the sum value matrix as the fusion matrix.

[0171] In an alternative embodiment, please refer to Figure 10 , the feature fusion network includes a third matrix conversion layer, a fourth matrix conversion layer, a third multi-head attention layer, a sixth matrix fusion layer, a feed-forward network layer, and a seventh matrix fusion layer, where

[0172] The third matrix conversion layer is configured to perform matrix conversion processing on the input matrix to obtain a key matrix and a value matrix;

[0173] The fourth matrix conversion layer is configured to perform matrix conversion processing on the input matrix to obtain a query matrix;

[0174] The third multi-head attention layer is configured to perform attention enhancement processing on the key matrix and value matrix output by the third matrix conversion layer and the query matrix output by the fourth matrix conversion layer to obtain an attention enhancement matrix;

[0175] The sixth matrix fusion layer is configured to perform matrix fusion processing on the attention enhancement matrix output by the third multi-head attention layer to obtain a fusion matrix;

[0176] The feed-forward network layer is configured to perform feed-forward calculation processing on the fusion matrix output by the sixth matrix fusion layer to obtain a feed-forward matrix;

[0177] The seventh matrix fusion layer is configured to perform matrix fusion processing on the feed-forward matrix output by the feed-forward network layer and the fusion matrix output by the sixth matrix fusion layer to obtain a fusion feature matrix.

[0178] Correspondingly, the electronic device can input the enhanced acoustic feature matrix and the phoneme feature matrix into the feature fusion network for feature fusion processing in the following manner:

[0179] Input the enhanced acoustic feature matrix into the third matrix conversion layer for matrix conversion processing to obtain a key matrix and a value matrix, denoted as the third key matrix and the third value matrix respectively;

[0180] Input the phoneme feature matrix into the fourth matrix conversion layer for matrix conversion processing to obtain a query matrix, denoted as the third query matrix;

[0181] Input the third query matrix, the third key matrix, and the third value matrix into the third multi-head attention layer for attention enhancement processing to obtain an attention enhancement matrix, denoted as the third attention enhancement matrix;

[0182] Input the third attention enhancement matrix and the third query matrix into the sixth matrix fusion layer for matrix fusion processing to obtain a fusion matrix, denoted as the acoustic-phoneme fusion matrix;

[0183] Input the acoustic-phoneme fusion matrix into the feed-forward network layer for feed-forward calculation processing to obtain a feed-forward matrix;

[0184] Input the feed-forward matrix and the acoustic-phoneme fusion matrix into the seventh matrix fusion layer for matrix fusion processing to obtain a fusion matrix, denoted as the fusion feature matrix.

[0185] Wherein, regarding how the sixth matrix fusion layer and the seventh matrix fusion layer perform matrix fusion, no specific limitation is made in this embodiment, and it can be configured by those skilled in the art according to actual needs.

[0186] For example, the sixth matrix fusion layer and the seventh matrix fusion layer have the same structure, both including two sub-layers, namely an addition layer and a layer normalization layer. When performing matrix fusion, first add the two input matrices through the addition layer to obtain a sum value matrix, and then perform layer normalization processing on the sum value matrix through the layer normalization layer to obtain a fusion matrix.

[0187] In an optional embodiment, please refer to Figure 11 , the first pronunciation skill detection network includes a first fully connected layer and a first classification function layer. Input the phoneme feature matrix into the first pronunciation skill detection network for pronunciation skill detection processing to obtain a first detection result, including:

[0188] Input the phoneme feature matrix into the first fully connected layer for fully connected processing to obtain a first fully connected result;

[0189] Input the first fully connected result into the first classification function layer for classification processing to obtain a first detection result.

[0190] It should be noted that since this embodiment is for detecting multiple types of pronunciation skills, correspondingly, the first classification function layer can adopt any multi-classification function.

[0191] Taking the Softmax function as an example, the dimension of the output vector of the Softmax function matches the number of pronunciation skills to be detected. For example, taking the English language as an example, the pronunciation skills to be detected include liaison, aspiration loss, and voicing. Then, the output vector of the Softmax function includes elements of four dimensions. Among them, one element is used to represent whether the pronunciation skill of "liaison" needs to be used to pronounce the text to be detected, one element is used to represent whether the pronunciation skill of "aspiration loss" needs to be used to pronounce the text to be detected, one element is used to represent whether the pronunciation skill of "voicing" needs to be used to pronounce the text to be detected, and one element is used to represent that no pronunciation skill needs to be used to pronounce the text to be detected.

[0192] Correspondingly, input the first fully connected result into the Softmax function to obtain a 4D output vector of the Softmax function. Take this 4D output vector as the first detection result. According to this first detection result, it can be determined whether a pronunciation skill needs to be used to pronounce the text to be detected, and when a pronunciation skill needs to be used to pronounce the text to be detected, which specific pronunciation skill needs to be used to pronounce the text to be detected.

[0193] In an optional embodiment, the second pronunciation skill detection network includes L branch detection networks. Each branch detection network corresponds to a different pronunciation skill. L is an integer greater than 1. Input the fused feature matrix into the second pronunciation skill detection network for pronunciation skill detection processing to obtain a second detection result, including:

[0194] Input the fused feature matrix into each branch detection network for pronunciation skill detection to obtain the branch pronunciation skill detection results of each branch detection network. The branch pronunciation skill detection results of each branch detection network represent whether the speaker uses the pronunciation skill corresponding to each branch detection network to pronounce the text to be detected;

[0195] Obtain the second detection result according to the branch pronunciation skill detection results of each branch detection network.

[0196] Among them, each branch detection network corresponds to a pronunciation skill and is configured to detect whether the speaker uses its corresponding pronunciation skill to pronounce the text to be detected. Correspondingly, the number of branch pronunciation skill detection results obtained is equal to the number of branch detection networks.

[0197] For example, taking the English language as an example, when the pronunciation skills to be detected include liaison, aspiration loss, and sonorization, L takes the value of 3. That is, the second pronunciation skill detection network will include 3 branch detection networks. Among them, 1 branch detection network corresponds to the pronunciation skill "liaison", 1 branch detection network corresponds to the pronunciation skill "aspiration loss", and 1 branch detection network corresponds to the pronunciation skill "sonorization". Correspondingly, these 3 branch detection networks will each output 1 branch pronunciation skill detection result, and a total of 3 branch pronunciation skill detection results are combined into the second detection result. At this time, the second detection result characterizes whether the speaker uses pronunciation skills to speak the text to be detected, and when the speaker uses pronunciation skills to speak the text to be detected, which specific pronunciation skills are used to speak the text to be detected.

[0198] It should be noted that the structures of each branch detection network are the same. Taking one branch detection network as an example for illustration, please refer to Figure 12 , the branch detection network includes a second fully connected layer and a second classification function layer. The fused feature matrix is input into each branch detection network for pronunciation skill detection to obtain the branch pronunciation skill detection results of each branch detection network, including:

[0199] The fused feature matrix is input into the second fully connected layer for fully connected processing to obtain the second fully connected result;

[0200] The second fully connected result is input into the second classification function layer for classification processing to obtain the branch pronunciation skill detection result.

[0201] It should be noted that the second classification function layer can adopt any binary classification function.

[0202] Taking the sigmoid function as an example, the output value of the sigmoid function is in the range of [0, 1]. After model training, the output of the sigmoid function can characterize the probability that the speaker uses the pronunciation skill corresponding to its branch detection network to speak the text to be detected. For example, when the output value of the sigmoid function reaches a preset threshold (which can be an empirical value taken by those skilled in the art according to actual needs), it can be determined that the speaker uses the pronunciation skill corresponding to its branch detection network to speak the text to be detected.

[0203] Correspondingly, the second fully connected result is input into the sigmoid function to obtain the output value of the sigmoid function, and this output value is used as the branch pronunciation skill detection result. According to this branch pronunciation skill detection result, it can be determined whether the speaker uses the pronunciation skill corresponding to its branch detection network to speak the text to be detected.

[0204] In an optional embodiment, before obtaining the text to be detected and converting the text to be detected into the corresponding phoneme sequence, it further includes:

[0205] Obtain multiple categories of first sample texts that are known to require different pronunciation techniques, and convert each category of first sample texts into corresponding positive sample phoneme sequences;

[0206] Obtain the first sample audio of each category of first sample texts spoken by the sample user using different pronunciation techniques, and extract the positive sample acoustic features of the first sample audio of each category of first sample texts;

[0207] Obtain second sample texts that are known not to require pronunciation techniques, and convert the second sample texts into corresponding negative sample phoneme sequences;

[0208] Obtain the second sample audio of the sample user speaking the second sample texts, and extract the negative sample acoustic features of the second sample audio;

[0209] Perform model training based on each category of positive sample phoneme sequences, each category of positive sample acoustic features, negative sample phoneme sequences, and negative sample acoustic features to obtain a pronunciation technique detection model.

[0210] In this embodiment, acoustic feature samples and phoneme sequence samples are not artificially constructed. Instead, starting from a data-driven idea, the model is allowed to learn different pronunciation techniques from a large amount of data. The following takes a specific language as an example for illustration.

[0211] For this language, the electronic device respectively obtains multiple categories of first sample texts that are known to require different pronunciation techniques. Here, there is no specific limit on the number of first sample texts for each acquired pronunciation technique, and it can be configured by those skilled in the art according to actual needs.

[0212] For each category of first sample texts obtained for each pronunciation technique (hereinafter simply referred to as each category of first sample texts), the electronic device respectively converts each category of first sample texts into corresponding phoneme sequences, denoted as positive sample phoneme sequences. Among them, for how to convert the first sample texts into phoneme sequences, it can be implemented by referring to the method of converting the text to be detected into phoneme sequences in the above embodiments, and will not be elaborated here.

[0213] The electronic device also obtains the audio of each category of first sample texts spoken by the sample user using different pronunciation techniques, denoted as the first sample audio, and extracts the acoustic features of the first sample audio of each category of first sample texts, denoted as positive sample acoustic features. Among them, the sample user can be a real person with pronunciation techniques or a virtual person with pronunciation techniques. Correspondingly, for how to obtain the first sample audio of each category of first sample texts and how to extract the positive sample acoustic features, it can be implemented by referring to the method of obtaining the audio to be detected and extracting the acoustic features of the audio to be detected in the above embodiments, and will not be elaborated here.

[0214] In addition, the electronic device also obtains text that is known not to require pronunciation skills to be spoken, denoted as the second sample text, and converts the second sample text into a corresponding phoneme sequence, denoted as the negative sample phoneme sequence. Among them, for how to convert the second sample text into a phoneme sequence, it can be implemented accordingly with reference to the method of converting the text to be detected into a phoneme sequence in the above embodiments, which will not be elaborated here.

[0215] The electronic device also obtains the audio of the sample user speaking the second sample text, denoted as the second sample audio, and extracts the acoustic features of the second sample audio, denoted as the negative sample acoustic features. Among them, for how to obtain the second sample audio of the second sample text and how to extract the negative sample acoustic features, it can be implemented accordingly with reference to the method of obtaining the audio to be detected and extracting the acoustic features of the audio to be detected in the above embodiments, which will not be elaborated here.

[0216] It should be noted that the number of sample users in this embodiment is not specifically limited and can be configured by those skilled in the art according to actual needs. For example, in this embodiment, 500 sample users are used to obtain the above positive sample phoneme sequence, positive sample acoustic features, negative sample phoneme sequence, and negative sample acoustic features.

[0217] After obtaining the above positive sample phoneme sequence, positive sample acoustic features, negative sample phoneme sequence, and negative sample acoustic features, the electronic device trains the model according to each type of positive sample phoneme sequence, each type of positive sample acoustic features, negative sample phoneme sequence, and negative sample acoustic features. Until the preset stop condition is met, a pronunciation skill detection model is obtained. The preset stop condition can be configured as the number of iterations of the model during training reaches a preset number, or the model converges.

[0218] Please refer to Figure 13 , to better execute the pronunciation skill detection method provided by this application, this application further provides a pronunciation skill detection device 400, as Figure 13 shown. The pronunciation skill detection device 400 includes:

[0219] A first acquisition module 410, configured to acquire the text to be detected and convert the text to be detected into a corresponding phoneme sequence;

[0220] A second acquisition module 420, configured to acquire the audio to be detected obtained by the speaker speaking the text to be detected and extract the acoustic features of the audio to be detected;

[0221] A detection module 430, configured to input the phoneme sequence and the acoustic features into the trained pronunciation skill detection model for pronunciation skill detection processing, and obtain a first detection result and a second detection result;

[0222] Among them, the first detection result is used to characterize whether pronunciation skills need to be adopted to pronounce the text to be detected, and the second detection result is used to characterize whether the speaker pronounces the text to be detected by using pronunciation skills.

[0223] In an optional embodiment, the pronunciation skill detection model includes a phoneme feature extraction network, an acoustic feature enhancement network, a feature fusion network, a first pronunciation skill detection network, and a second pronunciation skill detection network. The detection module 430 is configured to:

[0224] Input the phoneme sequence into the phoneme feature extraction network for feature extraction processing to obtain a phoneme feature matrix;

[0225] Input the phoneme feature matrix into the first pronunciation skill detection network for pronunciation skill detection processing to obtain a first detection result;

[0226] If the first detection result indicates that pronunciation skills need to be adopted to pronounce the text to be detected, then input the acoustic features into the acoustic feature enhancement network for feature enhancement processing to obtain an enhanced acoustic feature matrix;

[0227] Input the enhanced acoustic feature matrix and the phoneme feature matrix into the feature fusion network for feature fusion processing to obtain a fused feature matrix;

[0228] Input the fused feature matrix into the second pronunciation skill detection network for pronunciation skill detection processing to obtain a second detection result.

[0229] In an optional embodiment, the phoneme feature extraction network includes a phoneme embedding module and a phoneme feature extraction module. The detection module 430 is configured to:

[0230] Input the phoneme sequence into the phoneme embedding module for embedding processing to obtain a phoneme vector matrix;

[0231] Input the phoneme vector matrix into the phoneme feature extraction module for feature extraction processing to obtain a phoneme feature matrix.

[0232] In an optional embodiment, the phoneme feature extraction module includes at least 1 phoneme feature extraction sub-module. The detection module 430 is configured to:

[0233] When the number of phoneme feature extraction sub-modules is 1, input the phoneme vector matrix into the phoneme feature extraction sub-module for feature extraction processing to obtain a phoneme feature matrix; or,

[0234] When the number of phoneme feature extraction sub-modules is N, input the phoneme vector matrix into N phoneme feature extraction sub-modules in sequence for feature extraction processing to obtain a phoneme feature matrix, where N is an integer greater than 1.

[0235] In an alternative embodiment, the phoneme feature extraction sub-module includes a first matrix conversion layer, a first multi-head attention layer, and a first matrix fusion layer. The detection module 430 is configured to:

[0236] Input the phoneme vector matrix into the first matrix conversion layer for matrix conversion processing to obtain a first query matrix, a first key matrix, and a first value matrix;

[0237] Input the first query matrix, the first key matrix, and the first value matrix into the first multi-head attention layer for attention enhancement processing to obtain a first attention enhancement matrix;

[0238] Input the first attention enhancement matrix and the phoneme vector matrix into the first matrix fusion layer for matrix fusion processing to obtain a phoneme feature matrix.

[0239] In an alternative embodiment, the phoneme feature extraction network further includes a first position encoding module and a second matrix fusion layer. Before inputting the phoneme vector matrix into the phoneme feature extraction module for feature extraction processing to obtain a phoneme feature matrix, the detection module 430 is further configured to:

[0240] Input the phoneme vector matrix into the first position encoding module for position encoding processing to obtain a first position encoding matrix;

[0241] Input the first position encoding matrix and the phoneme vector matrix into the second matrix fusion layer for matrix fusion processing to obtain a phoneme position fusion matrix;

[0242] When inputting the phoneme vector matrix into the phoneme feature extraction module for feature extraction processing to obtain a phoneme feature matrix, the detection module 430 is configured to input the phoneme position fusion matrix into the phoneme feature extraction module for feature extraction processing to obtain a phoneme feature matrix.

[0243] In an alternative embodiment, the acoustic feature enhancement network includes a feature encoding module and at least one acoustic feature enhancement module. The detection module 430 is configured to:

[0244] Input the acoustic feature into the feature encoding module for feature encoding processing to obtain an acoustic feature matrix;

[0245] When there is one acoustic feature enhancement module, input the acoustic feature matrix into the acoustic feature enhancement module for feature enhancement processing to obtain an enhanced acoustic feature matrix; or,

[0246] When there are M acoustic feature enhancement modules, input the acoustic feature matrix into the M acoustic feature enhancement modules for sequential feature enhancement processing to obtain an enhanced acoustic feature matrix, where M is an integer greater than 1.

[0247] In an alternative embodiment, the feature encoding module includes a first convolutional layer, a first pooling layer, a second convolutional layer, and a second pooling layer. The detection module 430 is configured to:

[0248] Input the acoustic features into the first convolutional layer for convolutional processing to obtain a first convolutional result;

[0249] Input the first convolutional result into the first pooling layer for pooling processing to obtain a first pooling result;

[0250] Input the first pooling result into the second convolutional layer for convolutional processing to obtain a second convolutional result;

[0251] Input the second convolutional result into the second pooling layer for pooling processing to obtain an acoustic feature matrix.

[0252] In an alternative embodiment, the acoustic feature enhancement module includes a second matrix transformation layer, a second multi-head attention layer, a third matrix fusion layer, a third convolutional layer, a transposed convolutional layer, and a fourth matrix fusion layer. The detection module 430 is configured to:

[0253] Input the acoustic feature matrix into the second matrix transformation layer for matrix transformation processing to obtain a second query matrix, a second key matrix, and a second value matrix;

[0254] Input the second query matrix, the second key matrix, and the second value matrix into the second multi-head attention layer for attention enhancement processing to obtain a second attention enhancement matrix;

[0255] Input the second attention enhancement matrix and the acoustic feature matrix into the third matrix fusion layer for matrix fusion processing to obtain an acoustic fusion matrix;

[0256] Input the acoustic fusion matrix into the third convolutional layer for convolutional processing to obtain a third convolutional result;

[0257] Input the third convolutional result into the transposed convolutional layer for transposed convolutional processing to obtain a transposed convolutional result;

[0258] Input the acoustic fusion matrix and the transposed convolutional result into the fourth matrix fusion layer for matrix fusion processing to obtain an enhanced acoustic feature matrix.

[0259] In an alternative embodiment, the acoustic feature enhancement network further includes a second positional encoding module and a fifth matrix fusion layer. Before inputting the acoustic feature matrix into the acoustic feature enhancement module for feature enhancement processing to obtain an enhanced acoustic feature matrix, the detection module 430 is further configured to:

[0260] Input the acoustic feature matrix into the second positional encoding module for positional encoding processing to obtain a second positional encoding matrix;

[0261] Input the second position encoding matrix and the acoustic feature matrix into the fifth matrix fusion layer for matrix fusion processing to obtain an acoustic position fusion matrix;

[0262] When inputting the acoustic feature matrix into the acoustic feature enhancement module for feature enhancement processing to obtain an enhanced acoustic feature matrix, the detection module 430 is used for:

[0263] Input the acoustic position fusion matrix into the acoustic feature enhancement module for feature enhancement processing to obtain an enhanced acoustic feature matrix.

[0264] In an optional embodiment, the feature fusion network includes a third matrix conversion layer, a fourth matrix conversion layer, a third multi-head attention layer, a sixth matrix fusion layer, a feed-forward network layer, and a seventh matrix fusion layer. The detection module 430 is used for:

[0265] Input the enhanced acoustic feature matrix into the third matrix conversion layer for matrix conversion processing to obtain a third key matrix and a third value matrix;

[0266] Input the phoneme feature matrix into the fourth matrix conversion layer for matrix conversion processing to obtain a third query matrix;

[0267] Input the third query matrix, the third key matrix, and the third value matrix into the third multi-head attention layer for attention enhancement processing to obtain a third attention enhancement matrix;

[0268] Input the third attention enhancement matrix and the third query matrix into the sixth matrix fusion layer for matrix fusion processing to obtain an acoustic phoneme fusion matrix;

[0269] Input the acoustic phoneme fusion matrix into the feed-forward network layer for feed-forward calculation processing to obtain a feed-forward matrix;

[0270] Input the feed-forward matrix and the acoustic phoneme fusion matrix into the seventh matrix fusion layer for matrix fusion processing to obtain a fusion feature matrix.

[0271] In an optional embodiment, the first pronunciation skill detection network includes a first fully connected layer and a first classification function layer. The detection module 430 is used for:

[0272] Input the phoneme feature matrix into the first fully connected layer for fully connected processing to obtain a first fully connected result;

[0273] Input the first fully connected result into the first classification function layer for classification processing to obtain a first detection result.

[0274] In an optional embodiment, the second pronunciation skill detection network includes L branch detection networks. Each branch detection network corresponds to a different pronunciation skill. L is an integer greater than 1. The detection module 430 is used for:

[0275] Input the fused feature matrix into each branch detection network for pronunciation skill detection to obtain the branch pronunciation skill detection results of each branch detection network. The branch pronunciation skill detection results of each branch detection network indicate whether the speaker uses the pronunciation skills corresponding to each branch detection network to speak the text to be detected;

[0276] Obtain the second detection result according to the branch pronunciation skill detection results of each branch detection network.

[0277] In an optional embodiment, the branch detection network includes a second fully connected layer and a second classification function layer. The detection module 430 is used for:

[0278] Input the fused feature matrix into the second fully connected layer for fully connected processing to obtain the second fully connected result;

[0279] Input the second fully connected result into the second classification function layer for classification processing to obtain the branch pronunciation skill detection result.

[0280] In an optional embodiment, the pronunciation skill detection device provided in this application further includes a training module, which is used for:

[0281] Obtain multiple categories of first sample texts that are known to need to be spoken with different pronunciation skills, and convert each category of first sample texts into corresponding positive sample phoneme sequences;

[0282] Obtain the first sample audio of each category of first sample texts spoken by the sample user with different pronunciation skills, and extract the positive sample acoustic features of the first sample audio of each category of first sample texts;

[0283] Obtain a second sample text that is known not to need to be spoken with pronunciation skills, and convert the second sample text into a corresponding negative sample phoneme sequence;

[0284] Obtain the second sample audio of the sample user speaking the second sample text, and extract the negative sample acoustic features of the second sample audio;

[0285] Perform model training according to each category of positive sample phoneme sequences, each category of positive sample acoustic features, negative sample phoneme sequences, and negative sample acoustic features to obtain a pronunciation skill detection model.

[0286] In an optional embodiment, the second acquisition module 420 is used for:

[0287] Extract the Filterbank feature, fundamental frequency feature, and energy feature of the audio to be detected;

[0288] Fuse the Filterbank feature, fundamental frequency feature, and energy feature to obtain an acoustic feature.

[0289] In an optional embodiment, the first acquisition module 410 is used for:

[0290] Remove the unpronounced text units from the text to be detected to obtain a new text to be detected;

[0291] Convert each text unit in the new text to be detected into a corresponding phoneme unit to obtain a phoneme sequence.

[0292] It should be noted that the pronunciation skill detection device 400 provided in the embodiments of the present application belongs to the same concept as the pronunciation skill detection method in the above embodiments. For the specific implementation process, please refer to the above related embodiments, which will not be elaborated here.

[0293] The embodiments of the present application further provide an electronic device, including a memory and a processor. The processor is configured to execute the steps in the pronunciation skill detection method provided in this embodiment by calling a computer program stored in the memory.

[0294] Please refer to Figure 14 , Figure 14 , which is a schematic structural diagram of the electronic device 100 provided in the embodiments of the present application.

[0295] The electronic device 100 may include components such as a network interface 110, a memory 120, a processor 130, and a screen component. Those skilled in the art can understand that Figure 14 the structure of the electronic device 100 shown in

[0296] does not limit the electronic device 100, and it may include more or fewer components than shown in the figure, or combine some components, or have different component arrangements.

[0297] The network interface 110 can be used for network connections between devices.

[0298] The memory 120 can be used to store computer programs and data. The computer program stored in the memory 120 contains executable code. The computer program can be divided into various functional modules. The processor 130 executes various functional applications and data processing by running the computer program stored in the memory 120.

[0299] In an embodiment of the present application, the processor 130 in the electronic device 100 will load the executable code corresponding to one or more computer programs into the memory 120 according to the following instructions, and the processor 130 will execute the steps in the pronunciation skill detection method provided by the present application, such as:

[0300] Obtain the text to be detected, and convert the text to be detected into a corresponding phoneme sequence;

[0301] Obtain the audio to be detected obtained by the speaker speaking the text to be detected, and extract the acoustic features of the audio to be detected;

[0302] Input the phoneme sequence and the acoustic features into the trained pronunciation skill detection model for pronunciation skill detection processing to obtain a first detection result and a second detection result;

[0303] Among them, the first detection result is used to represent whether pronunciation skills need to be used to speak the text to be detected, and the second detection result is used to represent whether the speaker uses pronunciation skills to speak the text to be detected.

[0304] It should be noted that the electronic device 100 provided in the embodiment of the present application and the pronunciation skill detection method in the above embodiment belong to the same concept. The specific implementation process is detailed in the above related embodiments and will not be repeated here.

[0305] The present application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program stored thereon is executed on the processor of the electronic device provided in the embodiment of the present application, the processor of the electronic device is enabled to execute the steps in any of the above pronunciation skill detection methods suitable for the electronic device. Among them, the storage medium may be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM), etc.

[0306] The above has introduced in detail a pronunciation skill detection method, device, storage medium, and electronic device provided by the present application. Specific examples are used in this article to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application; at the same time, for those skilled in the art, according to the idea of the present application, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present application.< / eos> < / bos>

Claims

1. A pronunciation skill detection method, characterized in that Including: Obtain the text to be detected, and convert the text to be detected into a corresponding phoneme sequence; Obtain the audio to be detected obtained by the speaker speaking the text to be detected, and extract the acoustic features of the audio to be detected; Input the phoneme sequence and the acoustic features into a trained pronunciation skill detection model for pronunciation skill detection processing to obtain a first detection result and a second detection result; Wherein, the first detection result is used to characterize whether pronunciation skills need to be used to speak the text to be detected, and the second detection result is used to characterize whether the speaker uses pronunciation skills to speak the text to be detected; The pronunciation skill detection model includes a phoneme feature extraction network and a first pronunciation skill detection network. The first pronunciation skill detection network includes a first fully connected layer and a first classification function layer. The step of inputting the phoneme sequence and the acoustic features into the trained pronunciation skill detection model for pronunciation skill detection processing to obtain a first detection result includes: Input the phoneme sequence into the phoneme feature extraction network for feature extraction processing to obtain a phoneme feature matrix; Input the phoneme feature matrix into the first fully connected layer for fully connected processing to obtain a first fully connected result; Input the first fully connected result into the first classification function layer for classification processing to obtain the first detection result.

2. The pronunciation skill detection method according to claim 1, characterized in that The pronunciation skill detection model further includes an acoustic feature enhancement network, a feature fusion network and a second pronunciation skill detection network. The step of inputting the phoneme sequence and the acoustic features into the trained pronunciation skill detection model for pronunciation skill detection processing to obtain a first detection result and a second detection result further includes: Input the acoustic features into the acoustic feature enhancement network for feature enhancement processing to obtain an enhanced acoustic feature matrix; Input the enhanced acoustic feature matrix and the phoneme feature matrix into the feature fusion network for feature fusion processing to obtain a fusion feature matrix; Input the fusion feature matrix into the second pronunciation skill detection network for pronunciation skill detection processing to obtain the second detection result.

3. The pronunciation skill detection method according to claim 1, wherein The phoneme feature extraction network includes a phoneme embedding module and a phoneme feature extraction module. The step of inputting the phoneme sequence into the phoneme feature extraction network for feature extraction processing to obtain a phoneme feature matrix includes: Input the phoneme sequence into the phoneme embedding module for embedding processing to obtain a phoneme vector matrix; Input the phoneme vector matrix into the phoneme feature extraction module for feature extraction processing to obtain the phoneme feature matrix.

4. The pronunciation skill detection method according to claim 3, characterized in that The phoneme feature extraction module includes at least 1 phoneme feature extraction sub-module. The step of inputting the phoneme vector matrix into the phoneme feature extraction module for feature extraction processing to obtain the phoneme feature matrix includes: When there is 1 phoneme feature extraction sub-module, input the phoneme vector matrix into the phoneme feature extraction sub-module for feature extraction processing to obtain the phoneme feature matrix; or, When there are N phoneme feature extraction sub-modules, the phoneme vector matrix is input into the N phoneme feature extraction sub-modules in sequence for feature extraction processing to obtain the phoneme feature matrix, where N is an integer greater than 1.

5. The pronunciation skill detection method according to claim 4, characterized in that The phoneme feature extraction sub-module includes a first matrix conversion layer, a first multi-head attention layer, and a first matrix fusion layer. The process of inputting the phoneme vector matrix into the phoneme feature extraction sub-module for feature extraction processing to obtain the phoneme feature matrix includes: Input the phoneme vector matrix into the first matrix conversion layer for matrix conversion processing to obtain a first query matrix, a first key matrix, and a first value matrix; Input the first query matrix, the first key matrix, and the first value matrix into the first multi-head attention layer for attention enhancement processing to obtain a first attention enhancement matrix; Input the first attention enhancement matrix and the phoneme vector matrix into the first matrix fusion layer for matrix fusion processing to obtain the phoneme feature matrix.

6. The pronunciation skill detection method according to claim 3, characterized in that, The phoneme feature extraction network further includes a first position encoding module and a second matrix fusion layer. Before inputting the phoneme vector matrix into the phoneme feature extraction module for feature extraction processing to obtain the phoneme feature matrix, it further includes: Input the phoneme vector matrix into the first position encoding module for position encoding processing to obtain a first position encoding matrix; Input the first position encoding matrix and the phoneme vector matrix into the second matrix fusion layer for matrix fusion processing to obtain a phoneme position fusion matrix; The process of inputting the phoneme vector matrix into the phoneme feature extraction module for feature extraction processing to obtain the phoneme feature matrix includes: Input the phoneme position fusion matrix into the phoneme feature extraction module for feature extraction processing to obtain the phoneme feature matrix.

7. The pronunciation skill detection method according to claim 2, wherein The acoustic feature enhancement network includes a feature encoding module and at least 1 acoustic feature enhancement module. The process of inputting the acoustic feature into the acoustic feature enhancement network for feature enhancement processing to obtain the enhanced acoustic feature matrix includes: Input the acoustic feature into the feature encoding module for feature encoding processing to obtain an acoustic feature matrix; When there is 1 acoustic feature enhancement module, input the acoustic feature matrix into the acoustic feature enhancement module for feature enhancement processing to obtain the enhanced acoustic feature matrix; or, When there are M acoustic feature enhancement modules, input the acoustic feature matrix into the M acoustic feature enhancement modules in sequence for feature enhancement processing to obtain the enhanced acoustic feature matrix, where M is an integer greater than 1.

8. The pronunciation skill detection method according to claim 7, wherein, The feature encoding module includes a first convolutional layer, a first pooling layer, a second convolutional layer, and a second pooling layer. The process of inputting the acoustic feature into the feature encoding module for feature encoding processing to obtain the acoustic feature matrix includes: Input the acoustic feature into the first convolutional layer for convolutional processing to obtain a first convolutional result; Input the first convolutional result into the first pooling layer for pooling processing to obtain a first pooling result; Input the first pooling result into the second convolutional layer for convolutional processing to obtain a second convolutional result; Input the second convolution result into the second pooling layer for pooling processing to obtain the acoustic feature matrix.

9. The pronunciation skill detection method according to claim 7, characterized in that, The acoustic feature enhancement module includes a second matrix transformation layer, a second multi-head attention layer, a third matrix fusion layer, a third convolution layer, a transposed convolution layer, and a fourth matrix fusion layer. Inputting the acoustic feature matrix into the acoustic feature enhancement module for feature enhancement processing to obtain the enhanced acoustic feature matrix includes: Input the acoustic feature matrix into the second matrix transformation layer for matrix transformation processing to obtain a second query matrix, a second key matrix, and a second value matrix; Input the second query matrix, the second key matrix, and the second value matrix into the second multi-head attention layer for attention enhancement processing to obtain a second attention enhancement matrix; Input the second attention enhancement matrix and the acoustic feature matrix into the third matrix fusion layer for matrix fusion processing to obtain an acoustic fusion matrix; Input the acoustic fusion matrix into the third convolution layer for convolution processing to obtain a third convolution result; Input the third convolution result into the transposed convolution layer for transposed convolution processing to obtain a transposed convolution result; Input the acoustic fusion matrix and the transposed convolution result into the fourth matrix fusion layer for matrix fusion processing to obtain the enhanced acoustic feature matrix.

10. The pronunciation skill detection method according to claim 7, characterized in that The acoustic feature enhancement network further includes a second position encoding module and a fifth matrix fusion layer. Before inputting the acoustic feature matrix into the acoustic feature enhancement module for feature enhancement processing to obtain the enhanced acoustic feature matrix, it further includes: Input the acoustic feature matrix into the second position encoding module for position encoding processing to obtain a second position encoding matrix; Input the second position encoding matrix and the acoustic feature matrix into the fifth matrix fusion layer for matrix fusion processing to obtain an acoustic position fusion matrix; Inputting the acoustic feature matrix into the acoustic feature enhancement module for feature enhancement processing to obtain the enhanced acoustic feature matrix includes: Input the acoustic position fusion matrix into the acoustic feature enhancement module for feature enhancement processing to obtain the enhanced acoustic feature matrix.

11. The pronunciation skill detection method according to claim 2, characterized in that, The feature fusion network includes a third matrix transformation layer, a fourth matrix transformation layer, a third multi-head attention layer, a sixth matrix fusion layer, a feed-forward network layer, and a seventh matrix fusion layer. Inputting the enhanced acoustic feature matrix and the phoneme feature matrix into the feature fusion network for feature fusion processing to obtain a fusion feature matrix includes: Input the enhanced acoustic feature matrix into the third matrix transformation layer for matrix transformation processing to obtain a third key matrix and a third value matrix; Input the phoneme feature matrix into the fourth matrix transformation layer for matrix transformation processing to obtain a third query matrix; Input the third query matrix, the third key matrix, and the third value matrix into the third multi-head attention layer for attention enhancement processing to obtain a third attention enhancement matrix; Input the third attention enhancement matrix and the third query matrix into the sixth matrix fusion layer for matrix fusion processing to obtain an acoustic phoneme fusion matrix; Input the acoustic phoneme fusion matrix into the feedforward network layer for feedforward calculation processing to obtain a feedforward matrix; Input the feedforward matrix and the acoustic phoneme fusion matrix into the seventh matrix fusion layer for matrix fusion processing to obtain the fusion feature matrix.

12. The pronunciation skill detection method according to claim 2, characterized in that, The second pronunciation skill detection network includes L branch detection networks, and each branch detection network corresponds to a different pronunciation skill. L is an integer greater than 1. The step of inputting the fusion feature matrix into the second pronunciation skill detection network for pronunciation skill detection processing to obtain the second detection result includes: Input the fusion feature matrix into each of the branch detection networks for pronunciation skill detection to obtain the branch pronunciation skill detection result of each branch detection network. The branch pronunciation skill detection result of each branch detection network indicates whether the speaker uses the pronunciation skill corresponding to each branch detection network to say the text to be detected; Obtain the second detection result according to the branch pronunciation skill detection result of each branch detection network.

13. The pronunciation skill detection method according to claim 12, characterized in that, Each branch detection network includes a second fully connected layer and a second classification function layer. The step of inputting the fusion feature matrix into each branch detection network for pronunciation skill detection to obtain the branch pronunciation skill detection result of each branch detection network includes: Input the fusion feature matrix into the second fully connected layer for fully connected processing to obtain a second fully connected result; Input the second fully connected result into the second classification function layer for classification processing to obtain the branch pronunciation skill detection result.

14. The pronunciation skill detection method according to claim 12, wherein Before obtaining the text to be detected and converting the text to be detected into a corresponding phoneme sequence, it further includes: Obtain multiple categories of first sample texts that are known to need to be spoken with different pronunciation skills, and convert each category of the first sample texts into corresponding positive sample phoneme sequences; Obtain the first sample audio of each category of the first sample texts spoken by the sample user with different pronunciation skills, and extract the positive sample acoustic features of the first sample audio of each category of the first sample texts; Obtain a second sample text that is known not to need to be spoken with a pronunciation skill, and convert the second sample text into a corresponding negative sample phoneme sequence; Obtain the second sample audio of the sample user speaking the second sample text, and extract the negative sample acoustic features of the second sample audio; Perform model training according to each category of the positive sample phoneme sequences, each category of the positive sample acoustic features, the negative sample phoneme sequence, and the negative sample acoustic features to obtain the pronunciation skill detection model.

15. The pronunciation skill detection method according to any one of claims 1-14, characterized in that, The step of extracting the acoustic features of the audio to be detected includes: Extract the Filterbank features, fundamental frequency features, and energy features of the audio to be detected; Fuse the Filterbank features, the fundamental frequency features, and the energy features to obtain the acoustic features.

16. The pronunciation skill detection method according to any one of claims 1-14, characterized in that The step of converting the text to be detected into a corresponding phoneme sequence includes: Remove the text units that do not pronounce in the text to be detected to obtain a new text to be detected; Convert each text unit in the new text to be detected into a corresponding phoneme unit to obtain the phoneme sequence.

17. A pronunciation skill detection device, characterized in that, It includes: A first acquisition module, configured to acquire a text to be detected and convert the text to be detected into a corresponding phoneme sequence; A second acquisition module, configured to acquire a to-be-detected audio obtained by a speaker speaking the text to be detected, and extract acoustic features of the to-be-detected audio; A detection module, configured to input the phoneme sequence and the acoustic features into a trained pronunciation skill detection model for pronunciation skill detection processing, to obtain a first detection result and a second detection result; wherein, the first detection result is used to represent whether pronunciation skills are required to speak the text to be detected, and the second detection result is used to represent whether the speaker uses pronunciation skills to speak the text to be detected; The pronunciation skill detection model includes a phoneme feature extraction network and a first pronunciation skill detection network, the first pronunciation skill detection network includes a first fully-connected layer and a first classification function layer, and the detection module is further configured to: input the phoneme sequence into the phoneme feature extraction network for feature extraction processing to obtain a phoneme feature matrix; input the phoneme feature matrix into the first fully-connected layer for fully-connected processing to obtain a first fully-connected result; input the first fully-connected result into the first classification function layer for classification processing to obtain the first detection result.

18. A storage medium having a computer program stored thereon, characterized in that, When the computer program is loaded by a processor, it executes the steps in the pronunciation skill detection method according to any one of claims 1-16.

19. An electronic device, comprising a processor and a memory, the memory storing a computer program, characterized in that, The processor, by loading the computer program, is configured to execute the steps in the pronunciation skill detection method according to any one of claims 1 to 16.

Citation Information

Patent Citations

  • Voice synthesis method and device

    CN111968618A

  • Voice evaluation method and device

    CN112349300A

  • Vowel weak reading detection method and device

    CN113066510A

  • Spoken language pronunciation evaluation method and device, medium and equipment

    CN113345467A