Speech recognition method and device, storage medium and electronic device
By combining single-phoneme error detection and phoneme sound change recognition with a joint recognition model, the problem of high phoneme error rate in existing technologies is solved, and higher pronunciation recognition accuracy is achieved, especially in accurate judgment when considering sound changes.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- TENCENT TECHNOLOGY (SHENZHEN) CO LTD
- Filing Date
- 2022-09-08
- Publication Date
- 2026-07-21
AI Technical Summary
Existing pronunciation recognition technologies are insufficient in accuracy and fail to effectively account for sound changes, resulting in a high error rate in phoneme identification.
A joint recognition method is adopted, which combines single phoneme error detection and phoneme sound change recognition. The trained joint recognition model performs joint recognition of phoneme features, comprehensively considers the context information of the phoneme, and judges whether the phoneme has sound change, thereby improving the accuracy of pronunciation recognition.
It reduces the misjudgment rate of phoneme misreading, improves the accuracy of pronunciation recognition, and can more accurately identify whether a phoneme has been misread or has undergone sound change.
Smart Images

Figure CN116959441B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular to speech recognition methods, devices, storage media and electronic devices. Background Technology
[0002] Current technologies can automatically evaluate a speaker's pronunciation using software. This evaluation can automatically identify whether there are any errors in the pronunciation of each phoneme in the speaker's audio and provide feedback to the speaker. However, the accuracy of current pronunciation evaluations is not high, meaning that the pronunciation recognition results are not accurate. Summary of the Invention
[0003] To address at least one of the aforementioned technical problems, embodiments of this application provide a pronunciation recognition method, apparatus, storage medium, and electronic device.
[0004] On one hand, embodiments of this application provide a pronunciation recognition method, the method comprising:
[0005] Obtain the target audio;
[0006] Phoneme features are extracted from the target audio to obtain the phoneme features corresponding to each phoneme in the target audio.
[0007] The features of each phoneme are jointly identified to obtain the single phoneme misreading identification result and the sound change identification result corresponding to each phoneme. The joint identification represents an identification method that combines single phoneme misreading identification and phoneme sound change identification.
[0008] For each phoneme, the pronunciation recognition result of the phoneme is obtained based on the single phoneme misreading recognition result and the sound change recognition result of the phoneme. The pronunciation recognition result indicates whether the phoneme is misread in the target audio.
[0009] On the other hand, embodiments of this application provide a pronunciation recognition device, the device comprising:
[0010] The target audio acquisition module is used to acquire the target audio.
[0011] The phoneme feature extraction module is used to extract phoneme features from the target audio and obtain the phoneme features corresponding to each phoneme in the target audio.
[0012] The joint recognition module is used to jointly recognize the features of each phoneme to obtain the single phoneme misreading recognition result and the sound change recognition result corresponding to each phoneme. The joint recognition represents a recognition method that jointly performs single phoneme misreading recognition and phoneme sound change recognition.
[0013] The pronunciation recognition module is used to obtain the pronunciation recognition result of each phoneme based on the single phoneme misreading recognition result and the sound change recognition result corresponding to the phoneme. The pronunciation recognition result indicates whether the phoneme is misread in the target audio.
[0014] On the other hand, embodiments of this application provide a computer-readable storage medium storing at least one instruction or at least one program, wherein the at least one instruction or at least one program is loaded and executed by a processor to implement the above-described pronunciation recognition method.
[0015] On the other hand, embodiments of this application provide an electronic device, including at least one processor and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the at least one processor implements the above-described pronunciation recognition method by executing the instructions stored in the memory.
[0016] On the other hand, embodiments of this application provide a computer program product, including a computer program or instructions, which, when executed by a processor, implement the aforementioned pronunciation recognition method.
[0017] This application proposes a pronunciation recognition method. During execution, this method considers sound change phenomena and can accurately identify whether a phoneme has been mispronounced as a single phoneme, whether the phoneme has undergone sound change, and to reasonably classify the phoneme, thereby accurately determining whether the phoneme has been mispronounced in the target audio. For a given phoneme, the presence of sound change phenomena can be determined based on the preceding phonemes. In other words, this application embodiment can comprehensively consider three aspects: the pronunciation status of the single phoneme, the preceding phonemes, and the classification to which the phoneme belongs, to determine whether the phoneme has been mispronounced, thereby reducing the misjudgment rate of phoneme mispronunciation and improving the accuracy of pronunciation recognition. Attached Figure Description
[0018] To more clearly illustrate the technical solutions and advantages in the embodiments of this application or related technologies, the drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0019] Figure 1 This is a schematic diagram of a feasible implementation framework for the pronunciation recognition method provided in the embodiments of this specification;
[0020] Figure 2This is a schematic flowchart of a pronunciation recognition method provided in an embodiment of this application;
[0021] Figure 3 This is a schematic diagram of the client audio input interface provided in an embodiment of this application;
[0022] Figure 4 This is a schematic diagram of the training method of the joint recognition model provided in the embodiments of this application;
[0023] Figure 5 This is a schematic diagram of the loss of voice provided in an embodiment of this application;
[0024] Figure 6 This is a schematic diagram of continuous reading provided in an embodiment of this application;
[0025] Figure 7 This is a schematic diagram of the joint identification model provided in the embodiments of this application;
[0026] Figure 8 This is a schematic diagram of the pronunciation recognition result display method provided in the embodiments of this application;
[0027] Figure 9 This is a schematic diagram of an actual use scenario provided in the embodiments of this application;
[0028] Figure 10 This is a block diagram of the pronunciation recognition device provided in the embodiments of this application;
[0029] Figure 11 This is a schematic diagram of the hardware structure of a device for implementing the method provided in the embodiments of this application. Detailed Implementation
[0030] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the protection scope of the embodiments of this application.
[0031] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of the embodiments of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that embodiments of the present application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or server that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to these processes, methods, products, or devices.
[0032] To make the objectives, technical solutions, and advantages disclosed in the embodiments of this application clearer, the embodiments of this application will be further described in detail below with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are merely illustrative of the embodiments of this application and are not intended to limit the embodiments of this application.
[0033] Hereinafter, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this embodiment, unless otherwise stated, "multiple" means two or more. To facilitate understanding of the above-described technical solutions and their resulting technical effects in the embodiments of this application, the embodiments of this application first explain the relevant technical terms:
[0034] Cloud technology refers to a hosting technology that unifies hardware, software, and network resources within a wide area network (WAN) or local area network (LAN) to achieve data computation, storage, processing, and sharing. It encompasses network technologies, information technologies, integration technologies, management platform technologies, and application technologies based on cloud computing business models. These technologies can form resource pools, allowing for flexible and convenient on-demand use. Cloud computing technology will become a crucial support. Backend services of technical network systems require substantial computing and storage resources, such as video websites, image websites, and many portal websites. With the rapid development and application of the internet industry, every item may have its own identification mark in the future, requiring transmission to backend systems for logical processing. Data at different levels will be processed separately, and various industry data will require robust system support, which can only be achieved through cloud computing.
[0035] Intelligent Traffic Systems (ITS), also known as Intelligent Transportation Systems, effectively integrate advanced science and technology (information technology, computer technology, data communication technology, sensor technology, electronic control technology, automatic control theory, operations research, artificial intelligence, etc.) into transportation, service control, and vehicle manufacturing. This strengthens the connection between vehicles, roads, and users, thereby forming a comprehensive transportation system that ensures safety, improves efficiency, enhances the environment, and saves energy.
[0036] Intelligent Vehicle Infrastructure Cooperative Systems (IVICS) are a development direction of Intelligent Transportation Systems (ITS). IVICS utilizes advanced wireless communication and next-generation Internet technologies to implement comprehensive, real-time dynamic information interaction between vehicles and infrastructure. Based on the collection and fusion of dynamic traffic information across all times and spaces, it conducts active vehicle safety control and cooperative road management, fully realizing effective collaboration between people, vehicles, and roads. This ensures traffic safety, improves traffic efficiency, and ultimately forms a safe, efficient, and environmentally friendly road traffic system.
[0037] Artificial intelligence (AI) is the theory, methods, technology, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that attempts to understand the essence of intelligence and produce a new kind of intelligent machine that can react in a way similar to human intelligence. AI studies the design principles and implementation methods of various intelligent machines, enabling them to possess the functions of perception, reasoning, and decision-making.
[0038] Artificial intelligence (AI) is a comprehensive discipline encompassing a wide range of fields, including both hardware and software technologies. Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies primarily include computer vision, speech processing, natural language processing, and machine learning / deep learning.
[0039] Machine learning (ML) is a multidisciplinary field involving probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers can simulate or implement human learning behavior to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to endow computers with intelligence; its applications span all areas of artificial intelligence. Machine learning and deep learning typically include techniques such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and instructional learning.
[0040] Deep learning: The concept of deep learning originated from research on artificial neural networks. A multilayer perceptron with multiple hidden layers is a type of deep learning architecture. Deep learning discovers distributed feature representations of data by combining low-level features to form more abstract high-level representations of attribute categories or features.
[0041] Convolutional Neural Networks (CNNs) are a class of feedforward neural networks that incorporate convolutional computations and have a deep structure. They are one of the representative algorithms of deep learning. CNNs have representation learning capabilities and can classify input information in a translation-invariant manner according to their hierarchical structure; therefore, they are also known as "translation-invariant artificial neural networks."
[0042] ASR stands for Automatic Speech Recognition, a technology that converts human speech into text. Speech recognition is a multidisciplinary field, closely linked to numerous disciplines such as acoustics, phonetics, linguistics, digital signal processing theory, information theory, and computer science.
[0043] GMM (Gaussian Mixture Model) is a clustering algorithm that uses a Gaussian distribution as its parameter model and employs the Expectation Maximization (EM) algorithm for training. GMM can also be abbreviated as MOG. It precisely quantifies a phenomenon using a Gaussian probability density function (normal distribution curve), decomposing a single phenomenon into several models based on this function. GMMs have achieved good results in fields such as numerical approximation, speech recognition, image classification, image denoising, image reconstruction, fault diagnosis, video analysis, email filtering, and density estimation.
[0044] A Hidden Markov Model (HMM) is a statistical model used to describe a Markov process with hidden, unknown parameters. The challenge lies in determining these hidden parameters from the observable parameters. These parameters are then used for further analysis, such as pattern recognition. It is a statistical Markov model that models the system being modeled as a Markov process with unobserved (hidden) states.
[0045] Mel-frequency cepstral coefficients (MFCC) feature: A commonly used feature, also known as MFCC feature. MFCC feature retains semantically relevant content while filtering out irrelevant information such as background noise. A key characteristic of MFCC is that it uses a set of key coefficients to create the Mel-frequency cepstral spectrum, making its cepstral spectrum more closely resemble the nonlinear auditory system of humans.
[0046] Fbank: Fbank is one of the speech feature extraction methods. Due to its unique cepstral-based extraction method, it is more in line with the principles of human hearing, and is therefore the most common and effective speech feature extraction algorithm. The Fbank feature extraction method is equivalent to MFCC without the final discrete cosine transform (lossy transform). Compared with MFCC features, Fbank features retain more original speech data.
[0047] BLSTM: Bidirectional Long Short-Term Memory, a type of network used for sequence modeling.
[0048] Precision: Precision represents the proportion of samples that are identified as positive, among those that are actually positive. In short, given the following parameters:
[0049] T: True, F: False, P: Positive, N: Negative
[0050] Then combine:
[0051] TP: True positive; TN: True negative; FP: False positive; FN: False negative.
[0052] In simple terms, precision can be expressed as TP / (TP+FP);
[0053] Recall: Recall is the proportion of positive samples that are correctly identified as positive. Simply put, recall can be expressed as TP / (TP+FN).
[0054] F1 score: The F-Measure is a statistic, also known as the F-Score, which is the weighted harmonic average of precision and recall. It is often used to evaluate the performance of classification models.
[0055] A phone is the smallest phonetic unit divided according to the natural properties of speech. Analyzed based on the pronunciation actions in a syllable, one action constitutes one phone. Phones are divided into two major categories: vowels and consonants. For example, the Chinese syllable "ā" has only one phone, "ài" has two phones, and "dài" has three phones, etc. A phone is the smallest unit or the smallest phonetic segment that constitutes a syllable, and it is the smallest linear phonetic unit divided from the perspective of phonetic quality. A phone is a concrete physical phenomenon. The phonetic symbols of the International Phonetic Alphabet (developed by the International Phonetic Association to uniformly represent the speech sounds of various countries. Also known as the "International Phonetic Alphabet" or the "Universal Phonetic Alphabet") correspond one by one to the phones of all human languages.
[0056] CTC network: Connectionist Temporal Classification, which can be understood as a neural network-based temporal classification. Among them, Classification is relatively easy to understand, representing a classification problem; Temporal can be understood as a temporal problem, and Connectionist can be understood as the connections in a neural network. The CTC network can be used to train the acoustic model of speech recognition. The training of the acoustic model of speech recognition belongs to supervised learning, and it is necessary to know the corresponding label for each frame to perform effective training. During the data preparation stage of training, forced alignment of the speech must be performed. The introduction of the CTC network relaxes this one-to-one correspondence requirement, and only one input sequence and one output sequence are required for training. There are two advantages: no need for data alignment and one-to-one annotation; the CTC directly outputs the probability of sequence prediction without external post-processing.
[0057] Phonetic change refers to the speech changes that occur during pronunciation, such as assimilation, voicing, weak reading, elision, etc. In the current related technologies, speech changes are not taken into account in the process of pronunciation evaluation. That is to say, in the current related technologies, the determination of the phone level of a speaker mainly focuses on phone misjudgment. Phone misjudgment generally refers to judging whether the single pronunciation of a phone is accurate.
[0058] Currently, phoneme misclassification methods in related technologies can be broadly categorized into two types: ASR-based methods and feature-based classification methods. Speech recognition-based methods primarily compare the phoneme sequence predicted by the model with the actual sequence corresponding to the pronounced segment. If a discrepancy is found, the pronounced phoneme is considered incorrect. Typical methods use CTC networks to predict the pronounced phoneme sequence. These methods utilize the network to extract acoustic features and then use the CTC network to predict the phoneme sequence. Another common method is feature-based methods that rely on external phoneme alignment information to determine whether a phoneme has been misclassified. With the development of deep networks, many methods have begun to automatically extract features based on deep networks, such as using LSTM (Long Short-Term Memory) networks and CNN (Convolutional Neural Networks). Based on the features extracted by these networks, and relying on additional phoneme alignment modules, these phoneme alignment information are combined to determine whether each phoneme is pronounced incorrectly.
[0059] As mentioned above, related technologies primarily focus on the pronunciation of the phoneme itself when determining phoneme errors, without considering sound changes. For example, in cases of elision, the phoneme error determination module might assume the phoneme was not pronounced and therefore mispronounced. However, this is actually a sound change and the pronunciation is correct, leading to a discrepancy between the module's final result and the actual situation, resulting in a misjudgment. In view of this, this application proposes a phoneme recognition method that considers sound changes during execution. This method can accurately identify whether a phoneme has been mispronounced as a single phoneme, whether a sound change has occurred, and classify the phoneme appropriately, thereby accurately determining whether the phoneme has been mispronounced in the target audio. For a given phoneme, whether or not a sound change occurs can be determined based on the preceding phonemes. In other words, the embodiments of this application can comprehensively consider three aspects: the pronunciation status of the phoneme, the preceding phonemes, and the classification to which the phoneme belongs, to determine whether the phoneme has been mispronounced, thereby reducing the misjudgment rate of phoneme mispronunciation and improving the accuracy of pronunciation recognition.
[0060] The embodiments of this application can be applied to public cloud, private cloud, or hybrid cloud scenarios. For example, the model used for pronunciation recognition or the pronunciation recognition results generated in this application can be stored in the aforementioned public cloud, private cloud, or hybrid cloud. A private cloud is a cloud infrastructure and hardware / software resources created within a firewall, allowing various departments within an organization or enterprise to share resources within the data center. A public cloud typically refers to a cloud provided by a third-party provider to users. Public clouds are generally accessible via the Internet and may be free or inexpensive. The core attribute of a public cloud is shared resource services. Many instances of this type of cloud exist, providing services across today's open public networks. A hybrid cloud combines public and private clouds and is a major model and development direction of cloud computing in recent years. Private clouds are primarily geared towards enterprise users. For security reasons, enterprises prefer to store data in private clouds, but at the same time, they also want to access the computing resources of public clouds. In this context, hybrid clouds are being adopted more and more. They combine and match public and private clouds to achieve the best results. This personalized solution achieves both cost-effectiveness and security.
[0061] Please see Figure 1 , Figure 1 This is a schematic diagram illustrating a feasible implementation framework of the pronunciation recognition method provided in the embodiments of this specification, such as... Figure 1 As shown, the implementation framework may include at least a client 10 and a pronunciation recognition server 20, which communicate via a network 30. The client 10 can acquire target audio and transmit it to the pronunciation recognition server 20. The pronunciation recognition server 20 can obtain the target audio acquired by the client 10, extract phoneme features from the target audio to obtain the phoneme features corresponding to each phoneme in the target audio, perform joint recognition on each phoneme feature to obtain the single-phoneme misreading recognition result and the sound change recognition result corresponding to each phoneme, wherein the joint recognition represents a recognition method that jointly executes single-phoneme misreading recognition and phoneme sound change recognition; for each phoneme, based on the single-phoneme misreading recognition result and the sound change recognition result corresponding to the phoneme, the pronunciation recognition result of the phoneme is obtained, which represents whether the phoneme has been misread in the target audio. The pronunciation recognition result is transmitted to the client 10, which can then present the pronunciation recognition result.
[0062] The joint recognition model for pronunciation recognition can be set up on the pronunciation recognition server 20, or it can be made accessible to the pronunciation recognition server 20. The framework described above in this embodiment of the invention can provide pronunciation recognition capabilities required for applications in various scenarios, including but not limited to cloud technology, cloud gaming, cloud rendering, artificial intelligence, smart transportation, assisted driving, video media, smart communities, and instant messaging. Each component in this framework can be a terminal device or a server. Terminal devices include, but are not limited to, mobile phones, computers, intelligent voice interaction devices, smart home appliances, and in-vehicle terminals.
[0063] The following describes a pronunciation recognition method according to an embodiment of this application. Figure 2 This document illustrates a flowchart of a pronunciation recognition method provided in an embodiment of this application. This pronunciation recognition method can be executed based on the pronunciation recognition server described above. The embodiments of this application provide the method operation steps as described in the embodiments or flowcharts, but based on conventional or non-inventive labor, more or fewer operation steps may be included. The order of steps listed in the embodiments is merely one possible execution order among many steps and does not represent the only execution order. In actual systems, terminal devices, or server products, the method can be executed sequentially according to the embodiments or accompanying drawings, or in parallel (e.g., in a parallel processor or multi-threaded processing environment). The above method may include:
[0064] S101. Obtain the target audio.
[0065] In this application embodiment, the target audio can be understood as any audio that needs to be recognized for pronunciation. This application does not limit the method of obtaining the target audio. For example, it can come from a preset audio library or from a client that communicates with the pronunciation recognition server. The client can collect the target audio entered by the user and send the target audio to the aforementioned pronunciation recognition server.
[0066] This application does not limit the interface for the client to acquire the target audio, as long as audio recording can be implemented. Please refer to... Figure 3 It shows a schematic diagram of the client's audio input interface. Figure 3 The left image shows a "Start Reading Aloud" control. Users can click this control to trigger the client's audio input module and begin reading the English text in the left image. After the control is clicked, a "End Reading Aloud" control appears. Clicking this control stops audio input and extracts the audio content recorded in the audio input module, sending it as the target audio to the pronunciation recognition server.
[0067] S102. Extract phoneme features from the target audio to obtain the phoneme features corresponding to each phoneme in the target audio.
[0068] This application does not limit the method for extracting phoneme features. For example, a pre-trained acoustic model can be used to extract phoneme features. In one embodiment, a first acoustic model can be obtained through pre-training. This first acoustic model can be used to extract acoustic features from the target audio. A second acoustic model can also be obtained through pre-training. This second acoustic model can be used to extract alignment information from the target audio. This alignment information establishes the association between phonemes and frames in the target audio. That is, the alignment information expresses which frames in the target video express a certain phoneme. Of course, the alignment information can also establish the association between phonemes and time periods in the target audio. For example, the alignment information expresses which time period in the target video expresses a certain phoneme, and the frames included in that time period naturally correspond to that phoneme. This application does not limit the pre-training process and specific structure of the first and second acoustic models. They can be designed and trained independently according to actual conditions. These two models can also refer to related technologies, and this application does not limit them.
[0069] Based on the first acoustic model and the second acoustic model, it is possible to determine which frames each phoneme in the target audio corresponds to, and based on the acoustic features of these frames, the phoneme features corresponding to each phoneme can be obtained.
[0070] Of course, in some other embodiments, a third acoustic model can also be pre-trained, which can independently extract the phoneme features corresponding to each phoneme in the target audio.
[0071] The embodiments of this application do not limit the specific principles and structures of the first acoustic model, the second acoustic model, and the third acoustic model. For example, they can all use the ASR model.
[0072] S103. Jointly identify the features of each phoneme to obtain the single phoneme misreading identification results and the sound change identification results corresponding to each phoneme respectively. The joint identification represents the identification method that jointly performs single phoneme misreading identification and phoneme sound change identification.
[0073] In the embodiments of this application, single phoneme error recognition can be understood as a single phoneme error recognition operation. That is, it only focuses on the pronunciation of the phoneme itself, without considering the phoneme's pronunciation changes. If the actual pronunciation of the phoneme is consistent with the pronunciation that the phoneme should have, then the single phoneme misreading recognition result obtained based on the single phoneme error recognition indicates that there is no misreading; otherwise, it indicates that there is a misreading.
[0074] In the embodiments of this application, phoneme sound change recognition can be understood as a sound change recognition operation. For a certain phoneme, if it is read after a certain specific phoneme, the phoneme may undergo sound change, such as light reading, omission, or voicing. The sound change recognition operation can be understood as a recognition operation that takes into account the preceding information of the audio.
[0075] This application embodiment can train a joint recognition model to perform a joint recognition task, which includes at least a single phoneme error recognition task and a phoneme sound change recognition task. Thus, the joint recognition model can directly output the single phoneme misreading recognition result and the sound change recognition result corresponding to each phoneme based on the features of each phoneme.
[0076] S104. For each phoneme, based on the single phoneme misreading recognition result and the sound change recognition result corresponding to the phoneme, the pronunciation recognition result of the phoneme is obtained. The pronunciation recognition result indicates whether the phoneme is misread in the target audio.
[0077] The pronunciation recognition result needs to consider not only whether a single phoneme is mispronounced, but also whether the phoneme has undergone a sound change based on the preceding context. If a sound change occurs, it means that the pronunciation of that phoneme in the target audio is correct, even if this pronunciation differs from that of a single phoneme. The pronunciation recognition result may not be consistent with the single phoneme mispronunciation recognition result. This reflects the improvement of the embodiments of this application compared to related technologies that only focus on the correctness of single phoneme pronunciation. It avoids the situation where related technologies misjudge sound changes as pronunciation errors, thus improving the accuracy of the pronunciation recognition result.
[0078] Specifically, if neither the sound change recognition result nor the monophone mispronunciation recognition result is present, then the pronunciation recognition result is not mispronounced. Conversely, if neither the sound change recognition result nor the monophone mispronunciation recognition result is present, then the pronunciation recognition result is mispronounced. Finally, if the sound change recognition result does occur, then the pronunciation recognition result is not mispronounced.
[0079] In one embodiment, steps S102-S103 can be implemented by a joint recognition model, which includes a phoneme feature extraction model and a joint recognition network. The phoneme feature extraction model and the joint recognition network are used to implement steps S102 and S103, respectively.
[0080] This application does not limit the specific structure of the phoneme feature extraction model. In one embodiment, the phoneme feature extraction model includes an acoustic feature model and an alignment model. The acoustic feature model can be understood as the first acoustic model described above, and the alignment model can be understood as the second acoustic model described above.
[0081] The first acoustic model can be obtained through pre-training. In one implementation, Fbank features of each frame can be extracted based on sample audio, and these audio features are input into the first acoustic model. The first acoustic model can be composed of multiple nonlinear networks. It ultimately outputs the posterior probability of each frame. These posterior probabilities are then subjected to Bayesian transformation to obtain the output probability of a Hidden Markov Model (HMM). Based on this output probability and the sample audio, the loss is calculated, thereby optimizing the parameters of the first acoustic model. The trained first acoustic model is used as the acoustic feature model in the aforementioned phoneme feature extraction model. The second acoustic model can also be obtained through pre-training using ASR technology. Based on this second acoustic model, it is possible to determine which frames each phoneme occupies in the target video. For example, in one implementation, the second acoustic model can output the start and end times corresponding to each phoneme in the audio, and these start and end times can be used to determine which frames each phoneme occupies.
[0082] Please refer to Figure 4 The diagram illustrates the training process of the aforementioned joint recognition model. The training method for the joint recognition model includes:
[0083] S201. Obtain multiple sample audios, where each phoneme in each of the above sample audios has a corresponding phoneme category label, a single phoneme misreading label, and a sound change label.
[0084] Related technologies define phoneme categories, and phoneme category labels for each phoneme in sample audio can be generated by referring to phoneme-related standards in related technologies. Single-phoneme mispronunciation labels are used to characterize whether the corresponding phoneme, as a single phoneme, is mispronounced without considering preceding information. This application focuses on describing sound change labels.
[0085] Since the current definition of sound changes is mainly based on some rules, this application's embodiments first need to organize the sound change rules and label the sample audio with sound change tags. Sound changes are mainly divided into three cases: elision, linking, and voicing. The final pronunciation of these sound changes still exists in the original phoneme table and does not generate pronunciations outside the phoneme table.
[0086] For cases of aphonia, please refer to [the relevant documentation / reference]. Figure 5The diagram illustrates the phenomenon of elision. Taking the first row of boxes as an example, if there are two consecutive phonemes, and both of them are plosives, then the second phoneme can be silent, i.e., elision occurs. Plosives mainly include [p], [b], [t], [d], [k], and [g].
[0087] For cases involving connected speech, please refer to [the relevant documentation]. Figure 6 The diagram illustrates the connection between sounds. Taking the first box as an example, if there are two consecutive phonemes, and these two phonemes belong to a consonant and a vowel respectively, then these two phonemes are connected.
[0088] In cases of voicing, the main change is that voiceless consonants become voiced consonants.
[0089] This application embodiment considers these three sound change scenarios and labels each phoneme with a sound change tag. Phonemes that meet the relevant rules are labeled with a sound change tag of 1, while the remaining phonemes that do not meet the conditions are labeled with a sound change tag of 0. For single phoneme mispronunciation tags, if the labeled phoneme category tag does not match the correct pronunciation, the single phoneme mispronunciation tag is 1; otherwise, it is 0.
[0090] S202. Input the above sample audio into the above phoneme feature extraction model to obtain the sample phoneme features corresponding to each phoneme in the above sample audio.
[0091] Specifically, the sample audio can be input into the acoustic feature model to obtain the acoustic features of each frame in the sample audio. The sample audio can then be input into the alignment model to obtain the start and end frames corresponding to each phoneme in the sample audio. For each phoneme, based on the start frame, the end frame, and the acoustic features of each frame, the sample phoneme features are obtained. That is, the acoustic features of all frames within the closed interval formed by the start and end frames of the phoneme are fused to obtain the sample phoneme features. This application does not limit the fusion method; for example, it can be achieved through direct averaging, weighted averaging, convolution, pooling, etc.
[0092] In simple terms, in a specific implementation, step S202 involves inputting audio into a pre-trained acoustic feature model to obtain the acoustic representation corresponding to each frame. The audio is then input into an alignment model to obtain the start and end frames corresponding to each phoneme. Based on the start and end frames of each phoneme, the multi-frame representations corresponding to each phoneme are averaged to obtain the acoustic feature representation corresponding to each phoneme. This operation accurately represents the features of each phoneme, minimizing the loss of phoneme feature information and improving the accuracy of sample factor features.
[0093] S203. Input the phoneme features of each of the above samples into the above joint recognition network to obtain phoneme classification prediction results, single phoneme misreading prediction results, and sound change prediction results.
[0094] In fact, this joint recognition network performs three tasks: classification, phoneme mispronunciation detection, and phoneme sound change detection. These three tasks output phoneme classification prediction results, single phoneme mispronunciation prediction results, and sound change prediction results, respectively. In the application after training, the phoneme mispronunciation detection and phoneme sound change detection tasks were mainly used.
[0095] All three tasks mentioned above share the same joint identification network. The structure of the joint identification network is not limited in this application embodiment. In one specific implementation, the joint identification network is transformed through three fully connected layers, where each fully connected layer is composed of multiple nonlinear networks.
[0096] S204. Determine the model loss based on the above phoneme category labels, the above monophone misreading labels, the above sound change labels, the above phoneme classification prediction results, the above monophone misreading prediction results, and the above sound change prediction results.
[0097] Specifically, the model loss is determined based on the aforementioned phoneme category labels, monophone mispronunciation labels, sound change labels, phoneme classification prediction results, monophone mispronunciation prediction results, and sound change prediction results, including:
[0098] S2041. Based on the difference between the above phoneme category labels and the above phoneme classification prediction results, the phoneme classification loss is obtained.
[0099] This application does not limit the specific expression of phoneme classification loss. In a specific implementation, the following formula can be referred to:
[0100] Where i represents the sequence number of a phoneme in a sample audio file, phone i Represents the i-th phoneme. The phoneme category label specifically indicates that the i-th phone belongs to the j-th phoneme category. It is either 0 or 1. If the phoneme belongs to the j-th phoneme category, the value is 1; otherwise, it is 0. Let c be the probability that the i-th phone belongs to the j-th phoneme class. Specifically, it represents the probability that the i-th phone is predicted to belong to the j-th phoneme class. It is 0 if it does not belong to any phoneme class at all, or 1 if it belongs to any phoneme class at all. Otherwise, it is a decimal between 0 and 1. Here, c is the number of all phonemes in the phoneme table.
[0101] S2042. Based on the difference between the above monophone misreading labels and the above monophone misreading prediction results, the monophone misreading loss is obtained.
[0102] This application does not limit the specific expression of monophone misreading loss. In a specific implementation, the following formula two can be referred to:
[0103] Where i represents the sequence number of a phoneme in a sample audio file. This represents the mispronunciation label for the i-th phone and the j-th phoneme class. Specifically, it indicates whether the pronunciation of the i-th phone matches the pronunciation of the j-th phoneme class. It is 0 or 1, with a value of 1 if it matches and 0 otherwise. Let c be the probability that the i-th phone is predicted to belong to the j-th phoneme class. Specifically, it represents the probability that the i-th phone is predicted to have a pronunciation belonging to the j-th phoneme class. A value of 0 indicates no phoneme belonging to any of the classes, while a value of 1 indicates no phoneme belonging to any of the classes. mis There are two categories: wrong and right.
[0104] S2043. Based on the difference between the above sound change labels and the above sound change prediction results, the phoneme sound change loss is obtained.
[0105] This application does not limit the specific way of expressing phoneme sound change loss. In a specific implementation, the following formula three can be referred to:
[0106] Where i represents the sequence number of a phoneme in a sample audio file. This represents the phoneme sound change label for the i-th phone and the j-th phoneme class. Specifically, it indicates whether the pronunciation of the i-th phone has changed relative to the pronunciation of the j-th phoneme class. It is 0 or 1, with a value of 1 if it has changed and 0 otherwise. Let c be the probability that the i-th phone and j-th phoneme class are predicted to have a sound change. Specifically, it represents the probability that the i-th phone is predicted to have a sound change relative to the j-th phoneme class. 0 represents no change, and 1 represents a change. var There are two categories: sound changes and non-sound changes.
[0107] S2044. By integrating the above phoneme classification loss, the above monophone misreading loss, and the above phoneme sound change loss, the above model loss is obtained.
[0108] The model loss integrates the losses from three tasks, comprehensively characterizing the loss performance of each task. Optimizing model parameters based on this integrated loss enables the trained joint recognition model to accurately perform all three tasks, ensuring optimal performance. This application does not limit the fusion method; for example, a weighted summation method can be used to fuse the three losses, with no fixed weights that can be set according to actual conditions.
[0109] S205. Adjust the parameters in the joint recognition network according to the model loss described above until training is complete.
[0110] The embodiments of this application can optimize model parameters based on gradient descent or least squares methods and other related techniques. The specific optimization method is not limited, and the training conditions can also be limited according to the actual situation. For example, the training can be controlled by setting a loss threshold or a parameter tuning number threshold. Of course, the loss threshold or parameter tuning number threshold can be set according to the actual situation, which will not be elaborated further.
[0111] Please refer to Figure 7 The diagram illustrates the joint recognition model. During the training phase, at least one target frame corresponding to the sample phoneme is obtained based on the starting frame and ending frame corresponding to the sample phoneme. The acoustic features corresponding to the at least one target frame are then fused to obtain the sample phoneme features. These sample phoneme features are then transmitted to a joint recognition network consisting of three fully connected layers, which outputs the execution results of the three tasks. Based on these execution results and the labels, the parameters of the joint recognition model can be adjusted.
[0112] In the application phase, the target audio is input into the acoustic feature model. After multi-layer network transformation, the audio representation of each frame is obtained. Using start and end frame information, the acoustic representation of each phoneme across multiple frames is determined. These multi-frame acoustic representations are then fused to obtain the phoneme features for each phoneme. A multi-layer network transformation is then applied to each phoneme feature to obtain the phoneme category recognition result, single-phoneme mispronunciation recognition result, and sound change recognition result. When the sound change recognition result is 1, and the single-phoneme mispronunciation recognition result is 1 or 0, the pronunciation recognition result indicates that the pronunciation is correct, but there is a sound change. When the sound change recognition result is 0, the single-phoneme mispronunciation recognition result is used as the final pronunciation recognition result.
[0113] In some embodiments, the pronunciation recognition results corresponding to each phoneme in the target audio can also be displayed. Specifically, when a first target phoneme exists, the first target phoneme is displayed using a first display method, and the pronunciation recognition result corresponding to the first target phoneme indicates that a pronunciation change has occurred; when a second target phoneme exists, the second target phoneme is displayed using a second display method, and the pronunciation recognition result corresponding to the second target phoneme indicates that a misreading has occurred; the first display method and the second display method are different. Of course, this application does not limit the display method; for example, the display method may include font, font size, font color, whether it is bold, whether it is black, whether it is underlined, etc.
[0114] Please refer to Figure 8 The diagram illustrates how pronunciation recognition results are displayed. In the example, the phoneme "t" in "the" is mispronounced, and the phoneme "t" in "fact" undergoes a sound change. These two instances are displayed differently to help users clearly identify whether a mispronunciation or sound change has occurred. This improves the accuracy of pronunciation recognition results, making the granularity of pronunciation recognition finer and increasing user engagement.
[0115] Please refer to Figure 9 This diagram illustrates a practical usage scenario of an embodiment of this application. The user opens a target application, which performs phoneme recognition based on the method provided in this embodiment. The target application can display text for repetition. When the recording control of the target application is triggered, the user can repeat the text. After the repetition is completed, the target application sends the recorded audio and the repetition text to the server corresponding to the target application. The server, through communication with the pronunciation recognition server in this embodiment, can obtain the pronunciation recognition results of the phonemes. These phoneme pronunciation recognition results can include at least two of the following: pronunciation recognition results, single-phoneme mispronunciation recognition results, and sound change recognition results.
[0116] In the practical application scenario of this application embodiment, 1000 bilingual audio data entries were used to evaluate the technical effect of this application embodiment. The ground truth data included phoneme misreading annotation results and sound change annotation results. The F1 score was used to measure the phoneme misreading recognition effect and the sound change recognition effect. A comparison was made between the phoneme misreading recognition method that does not consider sound changes and the phoneme misreading recognition method that considers sound changes in this application embodiment. The comparison results are shown in the table below. It can be seen that the implementation of this application embodiment has higher accuracy in phoneme misreading recognition.
[0117]
[0118] This application proposes a phoneme recognition method. During execution, this method considers sound change phenomena and can accurately identify whether a phoneme has been mispronounced as a single phoneme, whether a sound change has occurred, and to classify the phoneme appropriately. This allows for accurate determination of whether the phoneme has been mispronounced in the target audio. For a given phoneme, the presence of a sound change phenomenon can be determined based on its preceding phonemes. In other words, this application can comprehensively consider three aspects: the pronunciation status of the single phoneme, its preceding phonemes, and its classification, to determine whether a phoneme has been mispronounced, thereby reducing the misjudgment rate and improving the accuracy of pronunciation recognition.
[0119] Please refer to Figure 10 The diagram shows a block diagram of a pronunciation recognition device according to this embodiment, the device comprising:
[0120] Target audio acquisition module 101 is used to acquire target audio;
[0121] The phoneme feature extraction module 102 is used to extract phoneme features from the target audio and obtain the phoneme features corresponding to each phoneme in the target audio.
[0122] The joint recognition module 103 is used to jointly recognize the features of each phoneme to obtain the single phoneme misreading recognition result and the sound change recognition result corresponding to each phoneme. The joint recognition represents a recognition method that jointly performs single phoneme misreading recognition and phoneme sound change recognition.
[0123] The pronunciation recognition module 104 is used to obtain the pronunciation recognition result of each phoneme based on the single phoneme misreading recognition result and the sound change recognition result corresponding to the phoneme. The pronunciation recognition result indicates whether the phoneme is misread in the target audio.
[0124] In one embodiment, the pronunciation recognition module described above is configured to perform the following operations:
[0125] If no sound change is found in the above sound change recognition result representation, and no monophone misreading is found in the above monophone misreading result representation, then the above pronunciation recognition result representation has not been misread.
[0126] If no sound change is found in the above sound change recognition result representation, and a monophone misreading is found in the above monophone misreading recognition result representation, then the above pronunciation recognition result representation is misread.
[0127] In cases where a sound change occurs in the above sound change recognition result representation, the above pronunciation recognition result representation is not misread.
[0128] In one embodiment, the apparatus further includes a training module for training a joint recognition model implementation, the joint recognition model including a phoneme feature extraction model and a joint recognition network, and the training module for performing the following operations:
[0129] Multiple sample audios are acquired, and each phoneme in each of the above sample audios has a corresponding phoneme category label, a single phoneme misreading label, and a sound change label.
[0130] The above sample audio is input into the above phoneme feature extraction model to obtain the sample phoneme features corresponding to each phoneme in the above sample audio.
[0131] The phoneme features of each of the above samples are input into the above joint recognition network to obtain phoneme classification prediction results, single phoneme misreading prediction results and sound change prediction results.
[0132] Based on the above phoneme category labels, the above monophone mispronunciation labels, the above sound change labels, the above phoneme classification prediction results, the above monophone mispronunciation prediction results, and the above sound change prediction results, determine the model loss;
[0133] Adjust the parameters in the joint recognition network based on the model loss described above until training is complete.
[0134] In one embodiment, the training module described above is used to perform the following operations:
[0135] Based on the difference between the above phoneme category labels and the above phoneme classification prediction results, the phoneme classification loss is obtained;
[0136] Based on the difference between the above monophone misreading labels and the above monophone misreading prediction results, the monophone misreading loss is obtained;
[0137] Based on the difference between the above sound change labels and the above sound change prediction results, the phoneme sound change loss is obtained;
[0138] By combining the above phoneme classification loss, the above monophone misreading loss, and the above phoneme sound change loss, the above model loss is obtained.
[0139] In one embodiment, the training module described above is used to perform the following operations:
[0140] The above sample audio is input into the above acoustic feature model to obtain the acoustic features of each frame in the above sample audio.
[0141] Input the above sample audio into the above alignment model to obtain the start frame and end frame corresponding to each phoneme in the above sample audio.
[0142] For each phoneme, the sample phoneme features corresponding to the phoneme are obtained based on the start frame corresponding to the phoneme, the end frame corresponding to the phoneme, and the acoustic features of each frame.
[0143] In one embodiment, the training module described above is used to perform the following operations:
[0144] Based on the start frame and the end frame corresponding to the above phonemes, at least one target frame corresponding to the above phonemes is obtained.
[0145] The acoustic features corresponding to at least one of the target frames are fused to obtain the sample phoneme features corresponding to the phonemes.
[0146] In one embodiment, the above-described apparatus includes a display module, which is configured to perform the following operations:
[0147] This displays the pronunciation recognition results for each phoneme in the target audio.
[0148] In one embodiment, the display module is configured to perform the following operations:
[0149] In the presence of a first target phoneme, the first target phoneme is displayed in a first display mode, and the sound change recognition result corresponding to the first target phoneme indicates the occurrence of a sound change.
[0150] In the presence of a second target phoneme, the second target phoneme is displayed in a second display mode, and the pronunciation recognition result corresponding to the second target phoneme indicates a misreading.
[0151] The first display method described above is different from the second display method described above.
[0152] The apparatus portion of the embodiments in this application is based on the same inventive concept as the method embodiment, and will not be described in detail here.
[0153] Furthermore, Figure 11 A schematic diagram of a hardware structure for implementing the method provided in the embodiments of this application is shown. This device can participate in or include the apparatus or system provided in the embodiments of this application. Figure 11As shown, device 10 may include one or more processors 102 (shown as 102a, 102b, ..., 102n in the figure) 102 (processor 102 may include, but is not limited to, a microprocessor MCU or a programmable logic device FPGA, etc.), a memory 104 for storing data, and a transmission device 106 for communication functions. In addition, it may also include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of the I / O interface), a network interface, a power supply, and / or a camera. Those skilled in the art will understand that... Figure 11 The structure shown is for illustrative purposes only and does not limit the structure of the electronic device described above. For example, device 10 may also include a... Figure 11 The more or fewer components shown, or having the same Figure 11 The different configurations shown.
[0154] It should be noted that the aforementioned one or more processors 102 and / or other data processing circuits are generally referred to herein as "data processing circuits". These data processing circuits may be embodied, in whole or in part, in software, hardware, firmware, or any other combination thereof. Furthermore, the data processing circuits may be a single, independent processing module, or may be integrated, in whole or in part, into any other element within the device 10 (or mobile device). As involved in the embodiments of this application, the data processing circuits serve as a processor control mechanism (e.g., selection of a variable resistor termination path connected to an interface).
[0155] The memory 104 can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the method described in the embodiments of this application. The processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, thereby realizing the above-described pronunciation recognition method. The memory 104 may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include memory remotely located relative to the processor 102, and these remote memories can be connected to the device 10 via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.
[0156] The transmission device 106 is used to receive or send data via a network. Specific examples of the network described above may include a wireless network provided by the communication provider of device 10. In one example, the transmission device 106 includes a Network Interface Controller (NIC), which can connect to other network devices via a base station to communicate with the Internet. In another example, the transmission device 106 may be a radio frequency (RF) module used for wireless communication with the Internet.
[0157] The display may be, for example, a touchscreen liquid crystal display (LCD) that allows a user to interact with the user interface of device 10 (or a mobile device).
[0158] It should be noted that the order of the embodiments described above is merely for descriptive purposes and does not represent the superiority or inferiority of the embodiments. Furthermore, the above description focuses on specific embodiments of this application. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps described in the claims can be performed in a different order than that shown in the embodiments and still achieve the desired result. Additionally, the processes depicted in the drawings do not necessarily require a specific or sequential order to achieve the desired result. In some implementations, multitasking and parallel processing are also possible or may be advantageous.
[0159] The embodiments in this application are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the device and server embodiments are basically similar to the method embodiments, so the descriptions are relatively simple; relevant parts can be referred to the descriptions of the method embodiments.
[0160] Those skilled in the art will understand that all or part of the steps of the above embodiments can be implemented by hardware or by a program instructing related hardware. The program can be stored in a computer-readable storage medium, such as a read-only memory, a disk, or an optical disk.
[0161] The instructions in the aforementioned storage medium can execute a pronunciation recognition method, the method comprising:
[0162] Obtain the target audio;
[0163] Phoneme features are extracted from the target audio to obtain the phoneme features corresponding to each phoneme in the target audio.
[0164] The features of each phoneme are jointly identified to obtain the single phoneme misreading identification results and the sound change identification results corresponding to each phoneme. The joint identification represents the identification method that combines single phoneme misreading identification and phoneme sound change identification.
[0165] For each phoneme, the pronunciation recognition result of the phoneme is obtained based on the single phoneme misreading recognition result and the sound change recognition result of the phoneme. The pronunciation recognition result indicates whether the phoneme is misread in the target audio.
[0166] In one embodiment, obtaining the pronunciation recognition result of the phoneme based on the single phoneme mispronunciation recognition result and the sound change recognition result of the phoneme includes:
[0167] If no sound change is found in the above sound change recognition result representation, and no monophone misreading is found in the above monophone misreading result representation, then the above pronunciation recognition result representation has not been misread.
[0168] If no sound change is found in the above sound change recognition result representation, and a monophone misreading is found in the above monophone misreading recognition result representation, then the above pronunciation recognition result representation is misread.
[0169] In cases where a sound change occurs in the above sound change recognition result representation, the above pronunciation recognition result representation is not misread.
[0170] In one embodiment, the above method is implemented through a joint recognition model, which includes a phoneme feature extraction model and a joint recognition network. The training method of the joint recognition model includes:
[0171] Multiple sample audios are acquired, and each phoneme in each of the above sample audios has a corresponding phoneme category label, a single phoneme misreading label, and a sound change label.
[0172] The above sample audio is input into the above phoneme feature extraction model to obtain the sample phoneme features corresponding to each phoneme in the above sample audio.
[0173] The phoneme features of each of the above samples are input into the above joint recognition network to obtain phoneme classification prediction results, single phoneme misreading prediction results and sound change prediction results.
[0174] Based on the above phoneme category labels, the above monophone mispronunciation labels, the above sound change labels, the above phoneme classification prediction results, the above monophone mispronunciation prediction results, and the above sound change prediction results, determine the model loss;
[0175] Adjust the parameters in the joint recognition network based on the model loss described above until training is complete.
[0176] In one embodiment, determining the model loss based on the phoneme category label, the monophone mispronunciation label, the sound change label, the phoneme classification prediction result, the monophone mispronunciation prediction result, and the sound change prediction result includes:
[0177] Based on the difference between the above phoneme category labels and the above phoneme classification prediction results, the phoneme classification loss is obtained;
[0178] Based on the difference between the above monophone misreading labels and the above monophone misreading prediction results, the monophone misreading loss is obtained;
[0179] Based on the difference between the above sound change labels and the above sound change prediction results, the phoneme sound change loss is obtained;
[0180] By combining the above phoneme classification loss, the above monophone misreading loss, and the above phoneme sound change loss, the above model loss is obtained.
[0181] In one embodiment, the phoneme feature extraction model includes an acoustic feature model and an alignment model. The process of inputting the sample audio into the phoneme feature extraction model to obtain the sample phoneme features corresponding to each phoneme in the sample audio includes:
[0182] The above sample audio is input into the above acoustic feature model to obtain the acoustic features of each frame in the above sample audio.
[0183] Input the above sample audio into the above alignment model to obtain the start frame and end frame corresponding to each phoneme in the above sample audio.
[0184] For each phoneme, the sample phoneme features corresponding to the phoneme are obtained based on the start frame corresponding to the phoneme, the end frame corresponding to the phoneme, and the acoustic features of each frame.
[0185] In one embodiment, obtaining the sample phoneme features corresponding to the phoneme based on the start frame corresponding to the phoneme, the end frame corresponding to the phoneme, and the acoustic features of each frame includes:
[0186] Based on the start frame and the end frame corresponding to the above phonemes, at least one target frame corresponding to the above phonemes is obtained.
[0187] The acoustic features corresponding to at least one of the target frames are fused to obtain the sample phoneme features corresponding to the phonemes.
[0188] In one embodiment, the above method further includes:
[0189] This displays the pronunciation recognition results for each phoneme in the target audio.
[0190] In one embodiment, displaying the pronunciation recognition results corresponding to each phoneme in the target audio includes:
[0191] In the presence of a first target phoneme, the first target phoneme is displayed in a first display mode, and the sound change recognition result corresponding to the first target phoneme indicates the occurrence of a sound change.
[0192] In the presence of a second target phoneme, the second target phoneme is displayed in a second display mode, and the pronunciation recognition result corresponding to the second target phoneme indicates a misreading.
[0193] The first display method described above is different from the second display method described above.
[0194] The above description is only a preferred embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present application should be included within the protection scope of the present application.
Claims
1. A pronunciation recognition method, characterized in that, The method includes: Obtain the target audio; Phoneme features are extracted from the target audio to obtain the phoneme features corresponding to each phoneme in the target audio. The features of each phoneme are jointly identified to obtain the single phoneme misreading identification result and the sound change identification result corresponding to each phoneme. The joint identification represents an identification method that combines single phoneme misreading identification and phoneme sound change identification. For each phoneme, the pronunciation recognition result of the phoneme is obtained based on the single phoneme misreading recognition result and the sound change recognition result of the phoneme. The pronunciation recognition result indicates whether the phoneme is misread in the target audio. The step of obtaining the pronunciation recognition result of the phoneme based on the single phoneme mispronunciation recognition result and the sound change recognition result of the phoneme includes: If no sound change is found in the sound change recognition result representation, and no single phoneme misreading is found in the single phoneme misreading recognition result representation, then the pronunciation recognition result representation has not been misread. If no sound change is found in the sound change recognition result representation, and a monophone misreading is found in the monophone misreading recognition result representation, then the pronunciation recognition result representation is misread. In cases where a sound change occurs in the sound change recognition result representation, the pronunciation recognition result representation is not misread.
2. The method according to claim 1, characterized in that, The method is implemented through a joint recognition model, which includes a phoneme feature extraction model and a joint recognition network. The training method of the joint recognition model includes: Multiple sample audios are acquired, and each phoneme in each sample audio has a corresponding phoneme category label, a single phoneme misreading label, and a sound change label; The sample audio is input into the phoneme feature extraction model to obtain the sample phoneme features corresponding to each phoneme in the sample audio. The phoneme features of each sample are input into the joint recognition network to obtain phoneme classification prediction results, single phoneme misreading prediction results, and sound change prediction results. The model loss is determined based on the phoneme category label, the monophone mispronunciation label, the sound change label, the phoneme classification prediction result, the monophone mispronunciation prediction result, and the sound change prediction result; The parameters in the joint recognition network are adjusted based on the model loss until training is complete.
3. The method according to claim 2, characterized in that, The step of determining the model loss based on the phoneme category label, the monophone mispronunciation label, the sound change label, the phoneme classification prediction result, the monophone mispronunciation prediction result, and the sound change prediction result includes: The phoneme classification loss is obtained based on the difference between the phoneme category label and the phoneme classification prediction result; The monophone misreading loss is obtained based on the difference between the monophone misreading label and the monophone misreading prediction result. Based on the difference between the sound change label and the sound change prediction result, the phoneme sound change loss is obtained; The model loss is obtained by fusing the phoneme classification loss, the single phoneme misreading loss, and the phoneme sound change loss.
4. The method according to claim 3, characterized in that, The phoneme feature extraction model includes an acoustic feature model and an alignment model. The step of inputting the sample audio into the phoneme feature extraction model to obtain the sample phoneme features corresponding to each phoneme in the sample audio includes: The sample audio is input into the acoustic feature model to obtain the acoustic features of each frame in the sample audio; The sample audio is input into the alignment model to obtain the start frame and end frame corresponding to each phoneme in the sample audio. For each phoneme, the sample phoneme features corresponding to the phoneme are obtained based on the start frame corresponding to the phoneme, the end frame corresponding to the phoneme, and the acoustic features of each frame.
5. The method according to claim 4, characterized in that, The step of obtaining the sample phoneme features corresponding to the phoneme based on the start frame corresponding to the phoneme, the end frame corresponding to the phoneme, and the acoustic features of each frame includes: Based on the start frame and the end frame corresponding to the phoneme, at least one target frame corresponding to the phoneme is obtained; The acoustic features corresponding to the at least one target frame are fused to obtain the sample phoneme features corresponding to the phoneme.
6. The method according to any one of claims 1 to 5, characterized in that, The method further includes: The display shows the pronunciation recognition results corresponding to each phoneme in the target audio.
7. The method according to claim 6, characterized in that, The display of pronunciation recognition results for each phoneme in the target audio includes: In the presence of a first target phoneme, the first target phoneme is displayed in a first display mode, and the sound change recognition result corresponding to the first target phoneme indicates the occurrence of a sound change. In the presence of a second target phoneme, the second target phoneme is displayed in a second display mode, and the pronunciation recognition result corresponding to the second target phoneme indicates a misreading. The first display method is different from the second display method.
8. A pronunciation recognition device, characterized in that, The device includes: The target audio acquisition module is used to acquire the target audio. The phoneme feature extraction module is used to extract phoneme features from the target audio and obtain the phoneme features corresponding to each phoneme in the target audio. The joint recognition module is used to jointly recognize the features of each phoneme to obtain the single phoneme misreading recognition result and the sound change recognition result corresponding to each phoneme. The joint recognition represents a recognition method that jointly performs single phoneme misreading recognition and phoneme sound change recognition. The pronunciation recognition module is used to obtain the pronunciation recognition result of each phoneme based on the single phoneme misreading recognition result and the sound change recognition result corresponding to the phoneme. The pronunciation recognition result indicates whether the phoneme is misread in the target audio. The step of obtaining the pronunciation recognition result of the phoneme based on the single phoneme mispronunciation recognition result and the sound change recognition result of the phoneme includes: If no sound change is found in the sound change recognition result representation, and no single phoneme misreading is found in the single phoneme misreading recognition result representation, then the pronunciation recognition result representation has not been misread. If no sound change is found in the sound change recognition result representation, and a monophone misreading is found in the monophone misreading recognition result representation, then the pronunciation recognition result representation is misread. In cases where a sound change occurs in the sound change recognition result representation, the pronunciation recognition result representation is not misread.
9. The apparatus according to claim 8, characterized in that, The device further includes a training module for training a joint recognition model implementation, the joint recognition model including a phoneme feature extraction model and a joint recognition network, and the training module is used to perform the following operations: Multiple sample audios are acquired, and each phoneme in each sample audio has a corresponding phoneme category label, a single phoneme misreading label, and a sound change label; The sample audio is input into the phoneme feature extraction model to obtain the sample phoneme features corresponding to each phoneme in the sample audio. The phoneme features of each sample are input into the joint recognition network to obtain phoneme classification prediction results, single phoneme misreading prediction results, and sound change prediction results. The model loss is determined based on the phoneme category label, the monophone mispronunciation label, the sound change label, the phoneme classification prediction result, the monophone mispronunciation prediction result, and the sound change prediction result; The parameters in the joint recognition network are adjusted based on the model loss until training is complete.
10. The apparatus according to claim 9, characterized in that, The training module is used to perform the following operations: The phoneme classification loss is obtained based on the difference between the phoneme category label and the phoneme classification prediction result; The monophone misreading loss is obtained based on the difference between the monophone misreading label and the monophone misreading prediction result. Based on the difference between the sound change label and the sound change prediction result, the phoneme sound change loss is obtained; The model loss is obtained by fusing the phoneme classification loss, the single phoneme misreading loss, and the phoneme sound change loss.
11. The apparatus according to claim 10, characterized in that, The phoneme feature extraction model includes an acoustic feature model and an alignment model, and the training module is used to perform the following operations: The sample audio is input into the acoustic feature model to obtain the acoustic features of each frame in the sample audio; The sample audio is input into the alignment model to obtain the start frame and end frame corresponding to each phoneme in the sample audio. For each phoneme, the sample phoneme features corresponding to the phoneme are obtained based on the start frame corresponding to the phoneme, the end frame corresponding to the phoneme, and the acoustic features of each frame.
12. The apparatus according to claim 11, characterized in that, The training module is used to perform the following operations: Based on the start frame and the end frame corresponding to the phoneme, at least one target frame corresponding to the phoneme is obtained; The acoustic features corresponding to the at least one target frame are fused to obtain the sample phoneme features corresponding to the phoneme.
13. The apparatus according to any one of claims 8 to 12, characterized in that, The device includes a display module, which is configured to perform the following operations: The display shows the pronunciation recognition results corresponding to each phoneme in the target audio.
14. The apparatus according to claim 13, characterized in that, The display module is used to perform the following operations: In the presence of a first target phoneme, the first target phoneme is displayed in a first display mode, and the sound change recognition result corresponding to the first target phoneme indicates the occurrence of a sound change. In the presence of a second target phoneme, the second target phoneme is displayed in a second display mode, and the pronunciation recognition result corresponding to the second target phoneme indicates a misreading. The first display method is different from the second display method.
15. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores at least one instruction or at least one program, which is loaded and executed by a processor to implement a pronunciation recognition method as described in any one of claims 1 to 7.
16. An electronic device, characterized in that, The method includes at least one processor and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the at least one processor implements a pronunciation recognition method as described in any one of claims 1 to 7 by executing the instructions stored in the memory.
17. A computer program product comprising a computer program or instructions, characterized in that, When the computer program or instructions are executed by the processor, they implement a pronunciation recognition method according to any one of claims 1 to 7.