Pronunciation Detection Method, Device, and Computer-Readable Medium
By extracting the audio frame characteristics of speech audio and generating posterior probability scores, the problem of inadequate pronunciation detection results in the prior art is solved, and higher detection accuracy and learning efficiency are achieved.
Patent Information
- Application Number
- CN202011119857.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-10-19
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2040-10-19
AI Technical Summary
Existing language learning software fails to fully consider the user's verbal habits and level in pronunciation detection, resulting in insufficient objective and inaccurate detection, which affects learning efficiency and enthusiasm.
By extracting audio frame features from the voice audio to be detected, posterior probability is generated based on these features and phoneme matching degree in the preset language, and the probability score of the phoneme is generated using neural network regression to improve the accuracy of pronunciation detection.
More precise pronunciation detection results are achieved, reducing the impact of first language pronunciation habits on second language detection results, and improving the objectivity of pronunciation detection and learners' practice efficiency.
Smart Images

Figure CN113409768B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technologies, and in particular, to a pronunciation detection method, apparatus, and computer-readable medium. Background Art
[0002] In many language learning software applications for education, the user's spoken voice is obtained for recognition to determine the user's pronunciation level, and corresponding teaching is executed when the pronunciation is incorrect or inaccurate. However, in many cases, the recognition methods in related technologies only identify the pronunciation of the current voice based on the phonemes of the voice, without considering information such as the user's language habits and level, resulting in the problem that the pronunciation detection results are not objective and accurate enough, which may in turn affect the learning efficiency and enthusiasm of learners. Summary of the Invention
[0003] Embodiments of this application provide a pronunciation detection method, apparatus, and computer-readable medium, which can at least to a certain extent obtain accurate pronunciation detection results and improve the accuracy of pronunciation detection and the practice efficiency of the pronunciation speaker.
[0004] Other features and advantages of this application will become apparent through the following detailed description, or be partially learned through the practice of this application.
[0005] According to one aspect of the embodiments of this application, a pronunciation detection method is provided, including: extracting audio frame features from a speech audio to be detected; generating a first posterior probability based on the matching degree between the audio frame features and first language phonemes in a preset first language, and generating a second posterior probability based on the matching degree between the audio frame features and second language phonemes in a preset second language; performing neural network regression processing on the first posterior probability and the second posterior probability to generate a probability score that the phonemes in the speech audio correspond to the second language phonemes.
[0006] According to one aspect of the embodiments of this application, a pronunciation detection apparatus is provided, including: an extraction unit configured to extract audio frame features from a speech audio to be detected; a probability unit configured to generate a first posterior probability based on the matching degree between the audio frame features and first language phonemes in a preset first language, and generate a second posterior probability based on the matching degree between the audio frame features and second language phonemes in a preset second language; a score unit configured to perform neural network regression processing on the first posterior probability and the second posterior probability to generate a probability score that the phonemes in the speech audio correspond to the second language phonemes.
[0007] In some embodiments of the present application, based on the foregoing solution, the extraction unit includes: an enhancement unit for performing signal enhancement processing on the speech audio to generate enhanced speech; a framing unit for performing framing processing on the enhanced speech based on a set frame length to generate a speech sequence; a windowing unit for performing windowing processing on the speech sequence based on a set window length to generate a windowed speech sequence; a transformation unit for performing Fourier transform on the windowed speech sequence to generate a frequency-domain speech signal; and a filtering unit for performing filtering processing on the frequency-domain speech signal to generate the audio frame features.
[0008] In some embodiments of the present application, based on the foregoing solution, the enhancement unit is configured to: obtain a first signal corresponding to a first moment in the speech audio and a second signal corresponding to a second moment before the first moment; calculate a weighted signal corresponding to the second signal based on a set signal coefficient and the second signal; generate an enhanced signal corresponding to the first moment based on the difference between the first signal and the weighted signal; and combine the enhanced signals corresponding to each moment in the speech audio to obtain the enhanced speech.
[0009] In some embodiments of the present application, based on the foregoing solution, the probability unit includes: a first model unit for inputting the audio frame features into a first acoustic model trained based on a first language sample and outputting a first posterior probability corresponding to the matching degree between the audio frame features and the first language phonemes; a first moment unit for identifying the start and end moments of the phonemes based on the waveforms corresponding to the phonemes in the speech audio; a first feature unit for determining the audio frame features included in the phonemes based on the start and end moments of the phonemes and the time frame information corresponding to the audio frame features; and a first probability unit for calculating the mean value of the first posterior probabilities corresponding to the audio frame features included in the phonemes to generate a first posterior probability of the phonemes corresponding to the first language phonemes.
[0010] In some embodiments of the present application, based on the foregoing solution, the pronunciation detection device is further configured to: obtain a first speech sample generated based on a first language and a first speech text corresponding to the first speech sample, and obtain a second speech sample generated based on a second language and a second speech text corresponding to the second speech sample; construct an acoustic model for identifying phonemes included in the audio based on a time-delay neural network; input the first speech sample into the acoustic model, and adjust the parameters of the acoustic model based on a first loss function obtained from the first phonemes output and the first speech text to obtain the first acoustic model; and input the second speech sample into the acoustic model, and adjust the parameters of the acoustic model based on a second loss function obtained from the second phonemes output and the second speech text to obtain a second acoustic model.
[0011] In some embodiments of the present application, based on the foregoing solution, the probability unit includes: a second model unit, configured to input the audio frame features into a second acoustic model trained based on second language samples, and output a second posterior probability corresponding to the matching degree between the audio frame features and the second language phonemes; a second time unit, configured to identify the start and end times corresponding to the phonemes based on the waveform of the speech audio; a second feature unit, configured to determine the audio frame features included in the phonemes based on the start and end times corresponding to the phonemes and the time frame information corresponding to the audio frame features; and a second probability unit, configured to calculate the mean of the second posterior probabilities corresponding to the audio frame features in the phonemes based on the start and end times corresponding to the phonemes, and determine the second posterior probability of the phonemes corresponding to the second language phonemes.
[0012] In some embodiments of the present application, based on the foregoing solution, the scoring unit is configured to: splice the first posterior probability and the second posterior probability to obtain a probability feature; perform neural network regression processing on the probability feature to generate a probability score of the phonemes in the speech audio corresponding to the second language phonemes.
[0013] In some embodiments of the present application, based on the foregoing solution, the display unit includes: a confidence unit, configured to determine the confidence between the phonemes and the second language phonemes based on the probability score of the phonemes corresponding to the second language phonemes; and a level determination unit, configured to determine the pronunciation accuracy level corresponding to each phoneme in the speech audio based on the confidence and a set confidence threshold.
[0014] In some embodiments of the present application, based on the foregoing solution, the confidence unit is configured to: determine the maximum probability score from the probability scores of the phonemes corresponding to the second language phonemes; calculate the ratio between the probability score of a specified phoneme corresponding to the second language phonemes and the maximum probability score; and determine the confidence between the specified phoneme and the second language phonemes based on the ratio.
[0015] In some embodiments of the present application, based on the foregoing solution, the pronunciation detection device further includes a display unit, configured to determine the pronunciation accuracy level corresponding to each phoneme based on the probability score, and display the text corresponding to the phoneme based on the display method corresponding to the pronunciation accuracy level.
[0016] In some embodiments of the present application, based on the foregoing solution, the display unit is configured to: obtain the text corresponding to the speech audio; segment the text based on the phonemes in the speech audio to generate the text corresponding to each phoneme; and display the text corresponding to the phoneme based on the display method corresponding to the pronunciation accuracy level of each phoneme.
[0017] In some embodiments of the present application, based on the foregoing solution, the pronunciation detection device is further configured to: query a target phoneme with the lowest pronunciation accuracy level from the pronunciation accuracy levels corresponding to each phoneme; obtain pronunciation teaching information corresponding to the target phoneme, where the pronunciation teaching information includes at least one of the following information: phonetic text, correct pronunciation, and demonstration video; and display the pronunciation teaching information.
[0018] In some embodiments of the present application, based on the foregoing solution, the pronunciation detection device is further configured to: obtain a target sentence containing the target phoneme from the sentence library corresponding to the second language; display the target sentence; obtain a practice audio sent by the user based on the target sentence; and detect the practice audio to obtain the pronunciation accuracy level corresponding to the target phoneme.
[0019] According to one aspect of the embodiments of the present application, there is provided a computer-readable medium having a computer program stored thereon, and when the computer program is executed by a processor, the pronunciation detection method described in the above embodiments is implemented.
[0020] According to one aspect of the embodiments of the present application, there is provided an electronic device, including: one or more processors; a storage device for storing one or more programs, and when the one or more programs are executed by the one or more processors, the one or more processors implement the pronunciation detection method described in the above embodiments.
[0021] According to one aspect of the embodiments of the present application, there is provided a computer program product or a computer program, the computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the pronunciation detection method provided in the above various alternative implementation manners.
[0022] In the technical solutions provided by some embodiments of the present application, after obtaining the speech audio to be detected, audio frame features are extracted from the speech audio. Based on the matching degree between the audio frame features and the first language phonemes in the preset first language, a first posterior probability is generated. At the same time, based on the matching degree between the audio frame features and the second language phonemes in the preset second language, a second posterior probability is generated. Then, neural network regression processing is performed on the first posterior probability and the second posterior probability to generate the probability score that the phonemes in the speech audio correspond to the second language phonemes. Based on the probability score, the pronunciation accuracy level corresponding to each phoneme is determined. By determining the corresponding display method based on the pronunciation accuracy level, the text corresponding to the phonemes is displayed on the terminal. Through the above method, the influence of the pronunciation habits of the first language on the pronunciation detection result of the second language can be avoided, the similar phonemes in the first language and the second language can be effectively distinguished, and then an accurate pronunciation detection result can be obtained, improving the accuracy of pronunciation detection and the practice efficiency of the speaker.
[0023] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] The accompanying drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts. In the drawings:
[0025] Figure 1 A schematic diagram showing an exemplary system architecture to which the technical solutions of the embodiments of the present application can be applied;
[0026] Figure 2 A schematic diagram showing a system architecture based on a cloud platform according to an embodiment of the present application;
[0027] Figure 3 A flowchart showing a pronunciation detection method according to an embodiment of the present application;
[0028] Figure 4 A schematic diagram showing the acquisition of speech audio according to an embodiment of the present application;
[0029] Figure 5 A schematic diagram showing the acquisition of speech audio according to an embodiment of the present application;
[0030] Figure 6 A schematic diagram showing the acquisition of speech audio according to an embodiment of the present application;
[0031] Figure 7 Schematically shows a schematic diagram of pronunciation detection according to an embodiment of the present application;
[0032] Figure 8 Schematically shows a schematic diagram of displaying the pronunciation accuracy level according to an embodiment of the present application;
[0033] Figure 9 Schematically shows extracting audio frame features from a speech audio to be detected according to an embodiment of the present application;
[0034] Figure 10 Schematically shows a flowchart of constructing an acoustic model according to an embodiment of the present application;
[0035] Figure 11 Schematically shows a schematic diagram of an acoustic model according to an embodiment of the present application;
[0036] Figure 12 Schematically shows a flowchart of generating a first posterior probability according to an embodiment of the present application;
[0037] Figure 13 Schematically shows a schematic diagram of displaying the text corresponding to a speech audio according to an embodiment of the present application;
[0038] Figure 14 Schematically shows a schematic diagram of displaying speech audio teaching according to an embodiment of the present application;
[0039] Figure 15 Schematically shows a schematic diagram of speech practice according to an embodiment of the present application;
[0040] Figure 16 Schematically shows a block diagram of a pronunciation detection method according to an embodiment of the present application;
[0041] Figure 17 Shows a schematic diagram of the structure of a computer system of an electronic device suitable for implementing the embodiments of the present application. Detailed implementation manners
[0042] Now, example embodiments will be described more fully with reference to the accompanying drawings. However, the example embodiments can be implemented in various forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this application will be more complete and comprehensive, and will fully convey the concept of the example embodiments to those skilled in the art.
[0043] In addition, the described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are provided to give a thorough understanding of the embodiments of the present application. However, those skilled in the art will realize that the technical solutions of the present application may be practiced without one or more of the specific details, or other methods, components, devices, steps, etc. may be adopted. In other cases, well-known methods, devices, implementations, or operations are not shown or described in detail to avoid obscuring aspects of the present application.
[0044] The block diagrams shown in the drawings are merely functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities may be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.
[0045] The flowcharts shown in the drawings are only illustrative and do not necessarily include all the content and operations / steps, nor are they necessarily executed in the order described. For example, some operations / steps may be decomposed, while some operations / steps may be combined or partially combined, so the actual execution order may change according to the actual situation.
[0046] Artificial Intelligence (AI) is a theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a manner similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines to enable the machines to have the functions of perception, reasoning, and decision-making. Artificial intelligence technology is an interdisciplinary subject involving a wide range of fields, including both hardware-level technologies and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0047] The key technologies of speech technology include automatic speech recognition (ASR), text-to-speech (TTS), and voiceprint recognition technology. Enabling computers to listen, see, speak, and sense is the future development direction of human-computer interaction, and among them, speech has become one of the most promising human-computer interaction methods. Natural language processing (NLP) is an important direction in the fields of computer science and artificial intelligence. It studies various theories and methods that can achieve effective communication between humans and computers using natural language. Natural language processing is a science that integrates linguistics, computer science, and mathematics. Therefore, the research in this field will involve natural language, that is, the language people use in daily life, so it has a close connection with the research of linguistics. Natural language processing technologies usually include text processing, semantic understanding, machine translation, robot question answering, knowledge graphs, and other technologies. Machine learning (ML) is an interdisciplinary subject that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and rote learning.
[0048] The solution provided by the embodiments of this application involves technologies such as speech technology, natural language processing, and machine learning in artificial intelligence. By using the first speech model corresponding to the first language and the second speech model corresponding to the second language obtained through pre-training, the matching degree between the phonemes of the first language and the second language in the speech audio emitted by the user is identified through natural language processing. Then, based on these two matching degrees, the matching degree of the speech audio based on the phonemes of the second language is determined through machine learning to determine the pronunciation accuracy of the speech audio, and then it is displayed on the user terminal to improve the accuracy of speech audio detection and teaching.
[0049] Figure 1 The figure shows a schematic diagram of an exemplary system architecture to which the technical solution of the embodiments of this application can be applied.
[0050] As Figure 1 shown, the system architecture may include terminal devices (such as Figure 1One or more of the smart phone 101, tablet computer 102, and portable computer 103 shown (which can of course also be a desktop computer, etc.), network 104, and server 105. The network 104 serves as a medium for providing a communication link between the terminal device and the server 105. The network 104 can include various connection types, such as wired communication links, wireless communication links, etc.
[0051] It should be understood that Figure 1 the number of terminal devices, networks, and servers in is merely illustrative. According to implementation needs, there can be any number of terminal devices, networks, and servers. For example, the server 105 can be a server cluster composed of multiple servers, etc.
[0052] The user can use the terminal device to interact with the server 105 via the network 104 to receive or send messages, etc. The server 105 can be a server that provides various services. For example, the user uses the terminal device 103 (which can also be the terminal device 101 or 102) to upload the audio frame features extracted from the voice audio to be detected to the server 105; based on the matching degree between the audio frame features and the first language phonemes in the preset first language, a first posterior probability is generated, and based on the matching degree between the audio frame features and the second language phonemes in the preset second language, a second posterior probability is generated; the first posterior probability and the second posterior probability are subjected to neural network regression processing to generate the probability score that the phonemes in the voice audio correspond to the second language phonemes; based on the probability score, the pronunciation accuracy level corresponding to each phoneme is determined, and the pronunciation accuracy level is sent to the terminal device, and the corresponding display method is determined based on the pronunciation accuracy level to display the text corresponding to the phonemes on the terminal.
[0053] In the above solution, after obtaining the voice audio to be detected, the audio frame features are extracted from the voice audio, a first posterior probability is generated based on the matching degree between the audio frame features and the first language phonemes in the preset first language, and at the same time, a second posterior probability is generated based on the matching degree between the audio frame features and the second language phonemes in the preset second language. Then, the first posterior probability and the second posterior probability are subjected to neural network regression processing to generate the probability score that the phonemes in the language audio correspond to the second language phonemes, so as to determine the pronunciation accuracy level corresponding to each phoneme based on the probability score, and send the pronunciation accuracy level to the terminal device, and determine the corresponding display method based on the pronunciation accuracy level to display the text corresponding to the phonemes on the terminal. By the above method, the influence of the pronunciation habit of the first language on the pronunciation detection result of the second language can be avoided, the similar phonemes in the first language and the second language can be effectively distinguished, and then an accurate pronunciation detection result can be obtained, improving the accuracy of pronunciation detection and the practice efficiency of the speaker.
[0054] It should be noted that the pronunciation detection method provided by the embodiments of the present application is generally executed by the server 105. Correspondingly, the pronunciation detection device is generally set in the server 105. However, in other embodiments of the present application, the terminal device may also have a similar function to the server, so as to execute the pronunciation detection solution provided by the embodiments of the present application.
[0055] Figure 2 It is a schematic diagram of a system architecture based on a cloud platform provided by the embodiments of the present application.
[0056] Cloud computing is a computing model that distributes computing tasks on a resource pool composed of a large number of computing devices, enabling various application systems to obtain computing power, storage space, and information services as needed. The network that provides resources is called the "cloud". The resources in the "cloud" seem to be infinitely expandable to users, and can be obtained at any time, used on demand, expanded at any time, and paid according to usage. As a basic capacity provider of cloud computing, a cloud computing resource pool will be established, abbreviated as a cloud platform, generally called an Infrastructure as a Service (IaaS) platform. Various types of virtual resources are deployed in the resource pool for external customers to choose and use. The cloud computing resource pool mainly includes: computing devices (virtualized machines, including operating systems), storage devices, and network devices. According to logical function division, a Platform as a Service (PaaS) layer can be deployed on the IaaS layer, and a Software as a Service (SaaS) layer can be deployed above the PaaS layer. The SaaS layer can also be directly deployed on the IaaS. PaaS is a platform for software operation, such as databases, web containers, etc. SaaS is various business software, such as web portals, SMS mass senders, etc. Generally speaking, SaaS and PaaS are upper layers relative to IaaS.
[0057] Cloud computing refers to the delivery and usage model of IT infrastructure, which means obtaining the required resources in a on-demand and easily scalable manner through the network; in a broad sense, cloud computing refers to the delivery and usage model of services, which means obtaining the required services in a on-demand and easily scalable manner through the network. Such services can be related to IT and software, the Internet, or other services. Cloud computing is the product of the development and integration of traditional computer and network technologies such as Grid Computing, Distributed Computing, Parallel Computing, Utility Computing, Network Storage Technologies, Virtualization, and Load Balance.
[0058] With the development of the Internet, real-time data streams, and the diversification of connected devices, as well as the promotion of demands such as search services, social networks, mobile commerce, and open collaboration, cloud computing has developed rapidly. Different from the previous parallel distributed computing, the emergence of cloud computing will, in concept, drive a revolutionary change in the entire Internet model and enterprise management model.
[0059] As Figure 2 shown, in the system architecture of this embodiment, a first language phoneme library and a second language phoneme library are stored in cloud 204, and the storage method can be to store the first language phoneme library through the first language model corresponding to the first language, and to store the second language phoneme library through the second language model corresponding to the second language.
[0060] The system architecture also includes terminal devices such as smart phone 201, tablet computer 202, and portable computer 203. In addition, it can also be other terminal devices. After the terminal device obtains the voice audio to be detected, it extracts the audio frame features from the voice audio, generates a first posterior probability based on the matching degree between the audio frame features and the first language phonemes in the preset first language, and at the same time generates a second posterior probability based on the matching degree between the audio frame features and the second language phonemes in the preset second language. Then, neural network regression processing is performed on the first posterior probability and the second posterior probability to generate the probability score of the phonemes in the language audio corresponding to the second language phonemes. Finally, based on the probability score, the pronunciation accuracy level corresponding to each phoneme is determined, and the corresponding display method is determined based on the pronunciation accuracy level, and the text corresponding to the phoneme is displayed based on the display method. Through the above method, the influence of the pronunciation habits of the first language on the pronunciation detection result of the second language can be avoided, the similar phonemes in the first language and the second language can be effectively distinguished, and then an accurate pronunciation detection result can be obtained, improving the accuracy of pronunciation detection and the practice efficiency of the pronunciation speaker.
[0061] In this embodiment, the server corresponding to the cloud can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing cloud computing services. The terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, etc., but is not limited thereto. The terminal and the server can be directly or indirectly connected through wired or wireless communication means, and this application does not make any restrictions here.
[0062] The implementation details of the technical solution of the embodiment of the present application are elaborated in detail below:
[0063] Figure 3 The flowchart of the pronunciation detection method according to an embodiment of the present application is shown. This pronunciation detection method can be executed by a server, which can be Figure 1 the server shown in Figure 3 or can be directly executed by a mid-end device. Referring to
[0064] shown, this pronunciation detection method at least includes steps S310 to S330, which are introduced in detail as follows:
[0065] Figures 4 to 6 It is a schematic diagram of obtaining voice audio provided by an embodiment of the present application.
[0066] As Figure 4 shown, in an embodiment of the present application, first obtain the voice audio to be detected. The obtaining method is to first display the text to be read aloud, for example, "hello", on the interface of the application program; then the user triggers the button of "click to start following".
[0067] As Figure 5 shown, when it is detected that the user triggers the button of "click to start following", start obtaining the voice audio and display "Please get closer to the microphone and read aloud" to prompt the user's reading method. At the same time, display the "click to end" button so that after the user finishes reading, click this button to obtain the indication to end obtaining the audio.
[0068] As Figure 6 shown, in order to limit the duration of the user's reading and improve the efficiency of reading detection, in this embodiment, a countdown process can also be performed during the reading process, and the countdown duration is displayed on the interface. For example, "countdown 2 seconds" is displayed during the process of collecting the voice audio. Through the above method, the user can be reminded to read in time, and the efficiency of obtaining the voice audio and the detection efficiency can be improved.
[0069] Figure 7A schematic diagram of pronunciation detection provided by an embodiment of the present application.
[0070] As Figure 7 shown, in an embodiment of the present application, it is divided into two parts: a client 710 and a server 720. In the client 710 part, it is introduced that the user performs pronunciation practice operations on the software. After the software records the user's practice audio, it transmits it to the server 720. After the server detects pronunciation errors, it transmits the errors back to the user and prompts the user for modification opinions. The server describes the whole process of performing phoneme-level pronunciation error detection on the user's pronunciation after receiving the audio of the user's pronunciation practice, and also explains that after detecting the pronunciation error information at the server, it transmits it back to the client for the user to perform the next practice.
[0071] Specifically, after obtaining the speech audio through the client 710, feature extraction 730 is performed on the speech audio at the server 720 to extract the frame-level features 740 of the audio. Specifically, the audio frame features in this embodiment are used to represent the audio features corresponding to each frame in the speech audio, that is, a set of feature sequences that can represent the frame level of the user's pronunciation. Exemplarily, the method of extracting audio frames in this embodiment can be obtained through filtering or through speech recognition.
[0072] It should be noted that in this embodiment, by extracting the frame-level audio features as the frame-level features of the audio, forced alignment is performed on the frame-level features to determine the playback time corresponding to each audio feature, and then the playback time period corresponding to all audio features included in a phoneme is determined. Based on the posterior probability corresponding to each audio feature, the overall posterior probability of the phoneme corresponding to the playback time period is determined. Through the above-mentioned extraction of frame-level features of the audio, the accuracy of audio detection and recognition can be accurately refined to the frame level, further enhancing the pronunciation detection accuracy.
[0073] In addition to the above method of extracting frame-level features from the audio, it is also possible to extract the audio features corresponding to the time periods of each phoneme from it to directly identify the posterior probability corresponding to the phoneme. By this means, the efficiency of audio data processing can be increased, and thus the efficiency of audio detection and recognition can be improved.
[0074] In step S320, a first posterior probability is generated based on the matching degree between the audio frame features and the first language phonemes in the preset first language, and a second posterior probability is generated based on the matching degree between the audio frame features and the second language phonemes in the preset second language.
[0075] As Figure 7As shown, in an embodiment of the present application, after obtaining the frame-level feature 740, that is, the audio frame feature, the playback time corresponding to each phoneme in the speech audio is determined by forced alignment 750. After the terminal device or the server obtains the speech audio, it is not determined whether the speech audio is in the first language or the second language. In this case, it is necessary to detect the matching degree or similarity between the current speech audio and the language phonemes corresponding to the two models through the recognition models of the two languages, that is, the posterior probability of the current speech audio corresponding to the set language phonemes. As Figure 7 shown, in this embodiment, the posterior probability at the segment level is identified based on the acoustic model, that is, the matching degree between the audio frame feature and the first language phonemes in the preset first language is determined through the First language (L1) acoustic model 760 to generate the first posterior probability 771, and the matching degree between the audio frame feature and the second language phonemes in the preset second language is determined through the Second Language (L2) acoustic model 762 to generate the second posterior probability 772.
[0076] It should be noted that the first language in this embodiment can be the user's mother tongue, and the second language can be the foreign language that the user is practicing. For example, in a scenario where a Chinese-speaking child is practicing English, the first language in this embodiment corresponds to Chinese, and the second language corresponds to English.
[0077] In addition, the first language in this embodiment can be the user's common language, and the second language can be the user's practice language, that is, the user usually communicates in the first language and uses the second language for practice during the practice process.
[0078] In step S330, neural network regression processing is performed on the first posterior probability and the second posterior probability to generate the probability score of the phoneme in the speech audio corresponding to the second language phoneme.
[0079] In an embodiment of the present application, after obtaining the first posterior probability and the second posterior probability, based on Figure 7In the manner of the Deep Neural Networks (DNN) 780, regression processing is performed through the first posterior probability and the second probability, and then based on the Goodness of Pronunciation (GOP) 790, the probability score that the phoneme in the speech audio corresponds to the second language phoneme is determined. The probability score in this embodiment is used to represent the similarity between the speech audio uttered by the user and the second language, and this similarity is based on the same or similar phonemes in the first language and the second language. By filtering out the pronunciation habits of the first language corresponding to the first posterior probability in the second posterior probability, the probability score specifically for the second language phoneme is obtained.
[0080] In the actual pronunciation process, a speaker whose native language is the first language may carry the pronunciation habits of the first language when speaking the second language, that is, speak the corresponding phonemes in the second language in the pronunciation manner of the first language. In this case, it will lead to bias in pronunciation detection. Therefore, in this embodiment, the posterior probability is calculated to capture the phonemes in the second language speech of the speaker that are the same as or similar to those in the first language, based on the second posterior probability corresponding to the second language, filter out the pronunciation content of the first language corresponding to the first posterior probability in the second language, and calculate the probability score that the phoneme in the speech audio corresponds to the second language phoneme based on the language phonemes obtained after filtering. Through the above method, the influence of the pronunciation habits of the first language on the pronunciation judgment of the second language is avoided, and the accuracy and objectivity of pronunciation detection are improved.
[0081] Specifically, in this embodiment, the specific way to perform neural network regression processing on the first posterior probability and the second posterior probability to generate the probability score that the phoneme in the speech audio corresponds to the second language phoneme can be to perform logistic regression on the first posterior probability and the second posterior probability through the Sigmoid function to obtain the probability score.
[0082] In an embodiment of the present application, after step S330, step S340 may further be included, that is, based on the probability score, determine the pronunciation accuracy level corresponding to each phoneme, and display the text corresponding to the phoneme based on the display manner corresponding to the pronunciation accuracy level.
[0083] As Figure 7 shown, in an embodiment of the present application, probability score thresholds corresponding to each pronunciation accuracy level are preset to determine whether there is a pronunciation error 711. Specifically, the pronunciation accuracy levels in this embodiment may include levels corresponding to states such as precise, good, qualified, wrong, and missed pronunciation, and each pronunciation accuracy level has its corresponding display manner, such as color, shade, or text size, etc. In this embodiment, after generating the probability score, based on the display manner corresponding to the pronunciation accuracy level, the text corresponding to the phoneme is displayed on the terminal interface.
[0084] Figure 8 This is a schematic diagram for displaying the pronunciation accuracy level provided by the embodiments of the present application.
[0085] As Figure 8 shown, after generating the pronunciation accuracy level, the total level corresponding to the entire speech audio can be displayed by the number of stars. Moreover, the method in this embodiment can identify the pronunciation accuracy levels corresponding to each phoneme in the speech audio, and display them in different display manners. For example, after recognizing the speech audio corresponding to "Good afternoon", if the pronunciation of "Good" is accurate, it is displayed in bold; if the pronunciation of "after" has a deviation, it is displayed in gray; if the pronunciation of "noon" is qualified, it is displayed in a normal display manner. Through the above display manners, the pronunciation status of the user for each phoneme can be clearly indicated, improving the user's practice efficiency.
[0086] In an embodiment of the present application, as Figure 9 shown, the process of extracting the audio frame features from the speech audio to be detected in step S310 includes the following steps:
[0087] Step S910, perform signal enhancement processing on the speech audio to generate enhanced speech.
[0088] In an embodiment of the present application, through pre-emphasis and other preprocessing on the learner's speech, the principle is mainly to enhance the high frequency of the speech signal to a certain extent and remove the influence of oral radiation. Specifically, the process of performing signal enhancement processing on the speech audio to generate enhanced speech in step S910 specifically includes:
[0089] Step S9101, obtain the first signal corresponding to the first moment in the speech audio and the second signal corresponding to the second moment before the first moment;
[0090] Step S9102, calculate the weighted signal corresponding to the second signal based on the set signal coefficient and the second signal;
[0091] Step S9103, generate the enhanced signal corresponding to the first moment based on the difference between the first signal and the weighted signal;
[0092] Step S9104, combine the enhanced signals corresponding to each moment in the speech audio to obtain enhanced speech.
[0093] Specifically, in a continuous signal, n is used to represent the playback moment of speech. In this embodiment, the first moment is n, and the second signal corresponding to the second moment before the first moment is n-1. The first signal corresponding to the first moment is x(n), and the second signal corresponding to the second moment before the first moment is x(n-1); based on the set signal coefficient α and the second signal x(n-1), the weighted signal αx(n-1) corresponding to the second signal is calculated; based on the difference between the first signal and the weighted signal, the enhanced signal y(n) = x(n) - αx(n-1) corresponding to the first moment is generated; finally, the enhanced signals corresponding to each moment in the voice audio are combined to obtain enhanced speech.
[0094] Step S920: Based on the set frame length, perform frame segmentation on the enhanced speech to generate a speech sequence.
[0095] In an embodiment of the present application, then operations such as frame segmentation are performed on the signal. Exemplarily, with a frame length of 25 ms and a frame shift of 10 ms, a pronunciation of several seconds is decomposed into a sequence of speech segments each 25 ms long.
[0096] Step S930: Based on the set window length, perform windowing on the speech sequence to generate a windowed speech sequence.
[0097] In an embodiment of the present application, windowing is performed on each small segment of speech in the speech segment sequence obtained in the above steps, and it can be performed by adding a Hamming window.
[0098] Step S940: Perform Fourier transform on the windowed speech sequence to generate a frequency-domain speech signal.
[0099] In an embodiment of the present application, Fourier transform is performed on each small segment of speech, so that the speech signal can be transformed from the time domain to the frequency domain.
[0100] Step S950: Perform filtering on the frequency-domain speech signal to generate audio frame features.
[0101] In an embodiment of the present application, each speech frame sequence in this group in the frequency domain is respectively extracted by Mel filtering frame by frame into features available for subsequent models. Essentially, it is a process of information compression and abstraction. The features that can be extracted at this stage are diverse, such as spectral features (Mel Frequency Cepstral Coefficients MFCC, Filter Bank FBANK, Packet Level Protocol PLP, etc.), frequency features (fundamental frequency, formant, etc.), time-domain features (duration features), energy features, and so on. The features used in the experiments of this case are 40-dimensional FBANK features. After passing through this module, a learner's pronunciation becomes a sequence of features that can represent their pronunciation, that is, the frame-level features mentioned in the figure.
[0102] In an embodiment of the present application, such asFigure 10 As shown in the figure, the pronunciation detection method in this embodiment further includes:
[0103] Step S1010: Obtain a first speech sample generated based on a first language and a first speech text corresponding to the first speech sample, and obtain a second speech sample generated based on a second language and a second speech text corresponding to the second speech sample.
[0104] Step S1020: Construct an acoustic model for identifying phonemes included in an audio based on a time-delay neural network;
[0105] Step S1030: Input the first speech sample into the acoustic model, and adjust the parameters of the acoustic model based on a first loss function obtained from the output first phoneme and the first speech text to obtain a first acoustic model;
[0106] Step S1040: Input the second speech sample into the acoustic model, and adjust the parameters of the acoustic model based on a second loss function obtained from the output second phoneme and the second speech text to obtain a second acoustic model.
[0107] Figure 11 It is a schematic diagram of an acoustic model provided by an embodiment of the present application.
[0108] In an embodiment of the present application, when a learner is learning the pronunciation of a second language (L2), for phonemes in L2 that are similar to the learner's mother tongue, i.e., the first language (L1), the learner will use the phonemes of L1 for substitution, which is one of the important reasons for pronunciation errors.
[0109] To avoid the problem of confusion detection caused by such similar phonemes, in this embodiment, based on the second language data 1110 read by a user whose native language is the first language and the first language data 1120 read by a user whose native language is the first language, that is, the local L1 speech corpus and the local L2 speech corpus, these two corpora are introduced as training data in the input layer to ensure the integrity and accuracy of training. Speech recognition tasks for Chinese and English are respectively set in the output layer, and through the shared hidden layer formed by the transfer learning mechanism in the time-delay neural network 1150, an acoustic model with Chinese and English pronunciation generalization capabilities is obtained, that is, the first acoustic model 1140 and the second acoustic model 1130.
[0110] In this embodiment, since different data and tasks may have an inherent correlation, by using the hidden layer parameters of a deep neural network to obtain such a correlation, the knowledge obtained from one task can be applied to the solution of another task. By using the multi-task and multi-language transfer learning method, data that has a strong correlation with the target task - the detection of English pronunciation errors of learners is included as much as possible to construct an acoustic model with language generalization ability.
[0111] In one embodiment of the present application, as Figure 12 shown, the process of generating the first posterior probability based on the matching degree between the audio frame features and the preset first language phonemes in step S320 includes the following steps:
[0112] Step S3210: Input the audio frame features into the first acoustic model trained based on the first language samples, and output the first posterior probability corresponding to the matching degree between the audio frame features and the first language phonemes;
[0113] Step S3220: Based on the waveforms corresponding to each phoneme in the speech audio, identify the start and end moments corresponding to the phonemes;
[0114] Step S3230: Based on the start and end moments corresponding to the phonemes and the time frame information corresponding to the audio frame features, determine the audio frame features included in the phonemes;
[0115] Step S3240: Calculate the mean value of the first posterior probabilities corresponding to the audio frame features included in the phonemes to generate the first posterior probability of the phoneme corresponding to the first language phonemes.
[0116] In one embodiment of the present application, inputting the audio frame features into the first acoustic model trained based on the first language samples and outputting the first posterior probability corresponding to the matching degree between the audio frame features and the first language phonemes, this probability represents the matching degree between each frame of the learner's pronunciation and the phoneme distribution of the first language samples in the acoustic model. Based on the speech recognition framework and the forced alignment technology, the given speech and text are aligned at the phoneme level, so that the start time and end time of each phoneme in the speech segment can be known. Based on the start and end moments corresponding to the phonemes and the time frame information corresponding to the audio frame features, determine the audio frame features included in the phonemes; finally, calculate the mean value of the first posterior probabilities corresponding to the audio frame features included in the phonemes, and obtain the average value of the probabilities as the first posterior probability of the phoneme corresponding to the first language phonemes. In this embodiment, the posterior probability feature of L1 is introduced to better distinguish the same or similar phonemes in L1 and L2. Combining the two features, finally, the probability score on the L2 phoneme set is obtained through DNN regression.
[0117] Specifically, when calculating the first posterior probability, after the speech features at the frame level are input into the acoustic model, the posterior probability of each frame can be obtained. This probability represents the matching degree between each frame of the learner's pronunciation and the phoneme distribution in the acoustic model. Since the acoustic model is usually trained with native speaker data, it can be regarded as viewing what the learner's pronunciation looks like from the perspective of a native speaker. The acoustic model adopted in this embodiment is a part of a speech recognition framework based on the Hidden Markov Model - Time Delay Neural Network (HMM-TDNN), and its principle is as follows:
[0118]
[0119] Among them, p(x|w) represents the acoustic model part, w represents the pronunciation text, which is the current pronunciation of the learner, that is, the speech audio corresponding to the second language. The probability p(x|w) characterizes the goodness or badness of the learner's pronunciation of the phonemes represented by the current text.
[0120] In an embodiment of the present application, the process of generating the second posterior probability based on the matching degree between the audio frame features and the preset second language phonemes in step S320 includes the following steps: inputting the audio frame features into the second acoustic model trained based on the second language samples, and outputting the second posterior probability corresponding to the matching degree between the audio frame features and the second language phonemes; identifying the start and end moments corresponding to the phonemes based on the waveform of the speech audio; determining the audio frame features included in the phonemes based on the start and end moments corresponding to the phonemes and the time frame information corresponding to the audio frame features; calculating the mean value of the second posterior probabilities corresponding to the audio frame features in the phonemes based on the start and end moments corresponding to the phonemes, and determining the second posterior probability of the phoneme corresponding to the second language phoneme.
[0121] In an embodiment of the present application, inputting the audio frame features into the second acoustic model trained based on the second language samples, and outputting the second posterior probability corresponding to the matching degree between the audio frame features and the second language phonemes. This probability represents the matching degree between each frame of the learner's pronunciation and the phoneme distribution of the second language samples in the acoustic model. Based on the speech recognition framework and the forced alignment technology, the given speech and text are aligned at the phoneme level, so that the start time and end time of each phoneme in the speech segment can be known. Determining the audio frame features included in the phonemes based on the start and end moments corresponding to the phonemes and the time frame information corresponding to the audio frame features; finally, calculating the mean value of the second posterior probabilities corresponding to the audio frame features included in the phonemes, and obtaining the average value of the probabilities as the second posterior probability of the phoneme corresponding to the second language phoneme.
[0122] In an embodiment of the present application, the process of performing neural network regression processing on the first posterior probability and the second posterior probability in step S330 to generate the probability score of the phoneme in the speech audio corresponding to the second language phoneme includes the following steps: concatenating the first posterior probability and the second posterior probability to obtain a probability feature; performing neural network regression processing on the probability feature to generate the probability score of the phoneme in the speech audio corresponding to the second language phoneme.
[0123] In an embodiment of the present application, according to the time period of each phoneme in the user's pronunciation obtained in the second module and the phoneme posterior probability on each frame obtained by the L1 and L2 acoustic models in the third module, the phoneme posterior probability of each phoneme segment in the two acoustic models is further obtained. The L1 posterior feature represents the score of each phoneme in the L1 phoneme set for the user's pronunciation of this phoneme, and the L2 posterior feature represents the score of each phoneme in the L2 phoneme set for the user's pronunciation of this phoneme. Thus, for similar phonemes in L1 and L2, we can combine the L1 posterior feature with the L2 posterior feature for auxiliary enhanced detection.
[0124] In an embodiment of the present application, the process of determining the pronunciation accuracy level corresponding to each phoneme based on the probability score in step S340 includes the following steps: determining the confidence level between the phoneme and the second language phoneme based on the probability score of the phoneme corresponding to the second language phoneme; determining the pronunciation accuracy level corresponding to each phoneme in the speech audio based on the confidence level and a set confidence threshold.
[0125] Specifically, in an embodiment of the present application, the confidence interval of a probability sample is the interval estimate of a certain population parameter of this sample. The confidence interval shows the degree to which the true value of this parameter has a certain probability of falling around the measurement result. The confidence interval gives the credibility of the measured value of the measured parameter. In this embodiment, the pronunciation accuracy level corresponding to each phoneme in the speech audio is determined based on the confidence level and a set confidence threshold.
[0126] In an embodiment of the present application, determining the confidence level between the phoneme and the second language phoneme based on the probability score of the phoneme corresponding to the second language phoneme includes: determining the maximum probability score from the probability scores of the phoneme corresponding to the second language phoneme; calculating the ratio between the probability score of the specified phoneme corresponding to the second language phoneme and the maximum probability score; determining the confidence level between the specified phoneme and the second language phoneme based on the ratio.
[0127] Specifically, this module is based on the posterior probability at the phoneme level output by the DNN neural network and the corresponding alignment information at the phoneme level. Through the Goodness of Pronunciation (GOP) algorithm, by comparing the probabilities of the phonemes that the user should have pronounced and the phonemes that the user actually pronounced, it is possible to judge whether each pronunciation has a deviation. From the probability score P(p) of the phoneme corresponding to the second language phoneme, the maximum probability score P(q) is determined; the ratio between the probability score of the specified phoneme corresponding to the second language phoneme and the maximum probability score is calculated as:
[0128]
[0129] where p represents the current pronunciation phoneme; P(p) represents the probability of the current phoneme output by the DNN; S represents the entire phoneme set; q represents the phoneme corresponding to the maximum probability output by the DNN; P(q) represents the maximum probability output by the DNN. After GOP scoring, the deviation of the phonemes in the user's pronunciation is judged through a threshold, and then it is returned to the user on the client which phoneme in the current pronunciation is a deviation and what it was mispronounced as.
[0130] Through the above process, we can know which phoneme has a deviation in the learner's pronunciation. Among them, how to obtain the final phoneme score is very important, and the level of the phoneme score directly affects the deviation judgment of the system.
[0131] In an embodiment of the present application, the process of displaying the text corresponding to the phoneme based on the display method corresponding to the pronunciation accuracy level in step S340 includes: obtaining the text corresponding to the speech audio; segmenting the text based on the phonemes in the speech audio to generate the text corresponding to each phoneme; and displaying the text corresponding to the phoneme based on the display method corresponding to the pronunciation accuracy level corresponding to each phoneme.
[0132] Figure 13 It is a schematic diagram showing the text corresponding to the speech audio in the embodiment of the present application.
[0133] As Figure 13 shown, in this embodiment, first obtain the text corresponding to the speech audio: Good afternoon; segment the text based on the phonemes in the speech audio to obtain the texts corresponding to each phoneme, namely "Good", "after", and "noon", determine their corresponding display methods, that is, bold, grayscale, and normal display, and display the texts corresponding to each phoneme based on these display methods.
[0134] In one embodiment of the present application, after determining the pronunciation accuracy level corresponding to each phoneme based on the probability score and displaying the text corresponding to the phoneme based on the display method corresponding to the pronunciation accuracy level, the method further includes: querying the target phoneme with the lowest pronunciation accuracy level from the pronunciation accuracy levels corresponding to each phoneme; obtaining pronunciation teaching information corresponding to the target phoneme, where the pronunciation teaching information includes at least one of the following information: phonetic text, correct pronunciation, and demonstration video; and displaying the pronunciation teaching information.
[0135] Figure 14 It is a schematic diagram showing the voice audio teaching in the embodiment of the present application.
[0136] As Figure 14 shown, after determining the pronunciation level of a user for a certain phrase or sentence and displaying it in the interface 1410, targeted teaching can be carried out based on the specific pronunciation situation, or teaching can be carried out for all the pronunciations in the phrase or sentence. For example Figure 14 in it, each phoneme in "Good afternoon" is taught in the interface 1420. The teaching information in this embodiment includes at least one of phonetic text, correct pronunciation, and demonstration video. By teaching the phrase or sentence, the learning efficiency of the user can be improved, and the learning and practice effects of the user can be enhanced.
[0137] In one embodiment of the present application, after querying the target phoneme with the lowest pronunciation accuracy level from the pronunciation accuracy levels corresponding to each phoneme, the method further includes: obtaining a target phrase or sentence containing the target phoneme from the phrase or sentence library corresponding to the second language; displaying the target phrase or sentence; obtaining a practice audio sent by the user based on the target phrase or sentence; and detecting the practice audio to obtain the pronunciation accuracy level corresponding to the target phoneme.
[0138] Figure 15 It is a schematic diagram of a voice practice provided by the embodiment of the present application.
[0139] As Figure 15 shown, after querying the target phoneme with the lowest pronunciation accuracy level from the pronunciation accuracy level interface 1510 corresponding to each phoneme, obtaining a target phrase or sentence containing the target phoneme from the phrase or sentence library corresponding to the second language; displaying the target phrase or sentence 1520; obtaining the practice audio "fool" sent by the user based on the target phrase or sentence and displaying it in the interface 1530; and detecting the practice audio to obtain the pronunciation accuracy level corresponding to the target phoneme. By the above-mentioned intensive practice method, the practice effect of the user and the pronunciation accuracy can be further improved.
[0140] Compared with the traditional system that only uses the local L2 speech corpus, the overall performance of this embodiment is relatively improved by 8.82%. The improvement in consonants is very obvious, but the performance of vowels is not prominent. This is because most of the similar or identical phonemes in L1 and L2 exist in the consonants of L1. The performance improvement of phonemes such as consonants Z, JH, and F exceeds 20% relatively, indicating that this solution can effectively distinguish the same or similar phonemes between L1 and L2 and improve the robustness of the pronunciation error detection model. After being combined with the product, since learners are often prone to mispronouncing similar sounds, English Jun can more accurately detect the pronunciations in learners' pronunciation that are similar to their mother tongue, making the scoring based on pronunciation quality more well-founded. Thus, children can focus their limited attention on the most important error correction. In this way, they can improve their oral English ability more efficiently and confidently.
[0141] The device embodiments of the present application are introduced below, which can be used to execute the pronunciation detection method in the above embodiments of the present application. It can be understood that the device can be a computer program (including program code) running in a computer device. For example, the device is an application software; the device can be used to execute the corresponding steps in the method provided by the embodiments of the present application. For the details not disclosed in the device embodiments of the present application, please refer to the embodiments of the above pronunciation detection method of the present application.
[0142] Figure 16 The block diagram of a pronunciation detection device according to an embodiment of the present application is shown.
[0143] Refer to Figure 16 As shown, a pronunciation detection device 1600 according to an embodiment of the present application includes: an extraction unit 1610, configured to extract audio frame features from the speech audio to be detected; a probability unit 1620, configured to generate a first posterior probability based on the matching degree between the audio frame features and the first language phonemes in a preset first language, and generate a second posterior probability based on the matching degree between the audio frame features and the second language phonemes in a preset second language; a scoring unit 1630, configured to perform neural network regression processing on the first posterior probability and the second posterior probability to generate a probability score of the phonemes in the speech audio corresponding to the second language phonemes.
[0144] In some embodiments of the present application, based on the foregoing solution, the extraction unit 1610 includes: an enhancement unit configured to perform signal enhancement processing on the speech audio to generate enhanced speech; a framing unit configured to perform framing processing on the enhanced speech based on a set frame length to generate a speech sequence; a windowing unit configured to perform windowing processing on the speech sequence based on a set window length to generate a windowed speech sequence; a transformation unit configured to perform Fourier transform on the windowed speech sequence to generate a frequency-domain speech signal; and a filtering unit configured to perform filtering processing on the frequency-domain speech signal to generate the audio frame features.
[0145] In some embodiments of the present application, based on the foregoing solution, the enhancement unit is configured to: obtain a first signal corresponding to a first moment in the speech audio and a second signal corresponding to a second moment before the first moment; calculate a weighted signal corresponding to the second signal based on a set signal coefficient and the second signal; generate an enhanced signal corresponding to the first moment based on a difference between the first signal and the weighted signal; and combine the enhanced signals corresponding to each moment in the speech audio to obtain the enhanced speech.
[0146] In some embodiments of the present application, based on the foregoing solution, the probability unit 1620 includes: a first model unit configured to input the audio frame features into a first acoustic model trained based on a first language sample and output a first posterior probability corresponding to a matching degree between the audio frame features and the first language phonemes; a first moment unit configured to identify start and end moments corresponding to a phoneme based on waveforms corresponding to each phoneme in the speech audio; a first feature unit configured to determine audio frame features included in the phoneme based on the start and end moments corresponding to the phoneme and time frame information corresponding to the audio frame features; and a first probability unit configured to calculate an average value of the first posterior probabilities corresponding to the audio frame features included in the phoneme to generate a first posterior probability of the phoneme corresponding to the first language phonemes.
[0147] In some embodiments of the present application, based on the foregoing solution, the pronunciation detection device 1600 is further configured to: obtain a first speech sample generated based on a first language and a first speech text corresponding to the first speech sample, and obtain a second speech sample generated based on a second language and a second speech text corresponding to the second speech sample; construct an acoustic model for identifying phonemes included in an audio based on a time delay neural network; input the first speech sample into the acoustic model, and adjust parameters of the acoustic model based on a first loss function obtained from the first phonemes output and the first speech text to obtain the first acoustic model; and input the second speech sample into the acoustic model, and adjust parameters of the acoustic model based on a second loss function obtained from the second phonemes output and the second speech text to obtain a second acoustic model.
[0148] In some embodiments of the present application, based on the foregoing solution, the probability unit 1620 includes: a second model unit, configured to input the audio frame features into a second acoustic model trained based on a second language sample, and output a second posterior probability corresponding to the matching degree between the audio frame features and the second language phonemes; a second time unit, configured to identify the start and end times corresponding to the phonemes based on the waveform of the speech audio; a second feature unit, configured to determine the audio frame features included in the phonemes based on the start and end times corresponding to the phonemes and the time frame information corresponding to the audio frame features; and a second probability unit, configured to calculate the mean of the second posterior probabilities corresponding to the audio frame features in the phonemes based on the start and end times corresponding to the phonemes, and determine the second posterior probability of the phonemes corresponding to the second language phonemes.
[0149] In some embodiments of the present application, based on the foregoing solution, the scoring unit 1630 is configured to: splice the first posterior probability and the second posterior probability to obtain a probability feature; perform neural network regression processing on the probability feature to generate a probability score of the phonemes in the speech audio corresponding to the second language phonemes.
[0150] In some embodiments of the present application, based on the foregoing solution, the display unit includes: a confidence unit, configured to determine the confidence between the phonemes and the second language phonemes based on the probability score of the phonemes corresponding to the second language phonemes; and a level determination unit, configured to determine the pronunciation accuracy level corresponding to each phoneme in the speech audio based on the confidence and a set confidence threshold.
[0151] In some embodiments of the present application, based on the foregoing solution, the confidence unit is configured to: determine the maximum probability score from the probability scores of the phonemes corresponding to the second language phonemes; calculate the ratio between the probability score of a specified phoneme corresponding to the second language phonemes and the maximum probability score; and determine the confidence between the specified phoneme and the second language phonemes based on the ratio.
[0152] In some embodiments of the present application, based on the foregoing solution, the pronunciation detection device further includes a display unit, configured to determine the pronunciation accuracy level corresponding to each phoneme based on the probability score, and display the text corresponding to the phoneme based on the display method corresponding to the pronunciation accuracy level.
[0153] In some embodiments of the present application, based on the foregoing solution, the display unit is configured to: obtain the text corresponding to the speech audio; segment the text based on the phonemes in the speech audio to generate the text corresponding to each phoneme; and display the text corresponding to the phoneme based on the display method corresponding to the pronunciation accuracy level of each phoneme.
[0154] In some embodiments of the present application, based on the foregoing solution, the pronunciation detection device 1600 is further configured to: query a target phoneme with the lowest pronunciation accuracy level from the pronunciation accuracy levels corresponding to each phoneme; obtain pronunciation teaching information corresponding to the target phoneme, where the pronunciation teaching information includes at least one of the following information: phonetic text, correct pronunciation, and demonstration video; and display the pronunciation teaching information.
[0155] In some embodiments of the present application, based on the foregoing solution, the pronunciation detection device 1600 is further configured to: obtain a target sentence including the target phoneme from the sentence library corresponding to the second language; display the target sentence; obtain a practice audio sent by the user based on the target sentence; and detect the practice audio to obtain the pronunciation accuracy level corresponding to the target phoneme.
[0156] Figure 17 The structural schematic diagram of a computer system of an electronic device suitable for implementing the embodiments of the present application is shown.
[0157] It should be noted that Figure 17 The computer system 1700 of the electronic device shown is only an example and should not impose any limitations on the functions and usage scope of the embodiments of the present application.
[0158] As Figure 17 shown, the computer system 1700 includes a central processing unit (CPU) 1701, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 1702 or the program loaded from the storage section 1708 into the random access memory (RAM) 1703, such as executing the method described in the foregoing embodiments. In the RAM 1703, various programs and data required for system operation are also stored. The CPU 1701, ROM 1702, and RAM 1703 are connected to each other through a bus 1704. The input / output (I / O) interface 1705 is also connected to the bus 1704.
[0159] The following components are connected to the I / O interface 1705: an input section 1706 including a keyboard, a mouse, etc.; an output section 1707 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage section 1708 including a hard disk, etc.; and a communication section 1709 including a network interface card such as a LAN (Local Area Network) card, a modem, etc. The communication section 1709 performs communication processing via a network such as the Internet. A drive 1710 is also connected to the I / O interface 1705 as required. A removable medium 1711, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 1710 as required so that a computer program read therefrom is installed into the storage section 1708 as required.
[0160] Specifically, according to an embodiment of the present application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, an embodiment of the present application includes a computer program product including a computer program carried on a computer-readable medium, the computer program including a computer program for performing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 1709, and / or installed from the removable medium 1711. When the computer program is executed by a central processing unit (CPU) 1701, various functions defined in the system of the present application are executed.
[0161] It should be noted that the computer-readable medium shown in the embodiments of the present application can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of a computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, a computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device. In the present application, a computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries a computer-readable computer program. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, and this computer-readable medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The computer program contained on a computer-readable medium can be transmitted using any appropriate medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.
[0162] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present application. Among them, each block in a flowchart or block diagram can represent a module, a program segment, or a part of code, and the above module, program segment, or part of code contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, as well as the combination of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0163] The units involved in the embodiments described in this application can be implemented in software or in hardware, and the described units can also be provided in a processor. In some cases, the names of these units do not constitute a limitation on the unit itself.
[0164] According to one aspect of the present application, there is provided a computer program product or a computer program, the computer program product or the computer program including computer instructions, the computer instructions being stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the methods provided in the above various alternative implementations.
[0165] As another aspect, the present application further provides a computer-readable medium, which may be included in the electronic device described in the above embodiments; or may exist alone without being assembled into the electronic device. The above computer-readable medium carries one or more programs, and when the above one or more programs are executed by an electronic device, the electronic device implements the methods described in the above embodiments.
[0166] It should be noted that although several modules or units of a device for action execution are mentioned in the above detailed description, this division is not mandatory. In fact, according to the embodiments of the present application, the features and functions of the two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.
[0167] From the description of the above embodiments, those skilled in the art can easily understand that the example embodiments described herein can be implemented by software or by a combination of software and necessary hardware. Therefore, the technical solutions according to the embodiments of the present application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which may be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several instructions to enable a computing device (which may be a personal computer, a server, a touch terminal, or a network device, etc.) to execute the methods according to the embodiments of the present application.
[0168] After considering the specification and practicing the disclosed embodiments herein, those skilled in the art will readily conceive of other embodiments of the present application. The present application is intended to cover any variations, uses, or adaptations of the present application, which follow the general principles of the present application and include common general knowledge or conventional technical means in the technical field not disclosed in the present application.
[0169] It should be understood that the present application is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present application is only limited by the appended claims.
Claims
1. A pronunciation detection method, characterized in that, Including: Extracting audio frame features from the speech audio to be detected; Generating a first posterior probability based on the matching degree between the audio frame features and the first language phonemes in a preset first language, and generating a second posterior probability based on the matching degree between the audio frame features and the second language phonemes in a preset second language; Based on the second posterior probability, filtering out the pronunciation content of the preset first language corresponding to the first posterior probability in the preset second language, and calculating a probability score that the phonemes in the speech audio correspond to the second language phonemes based on the language phonemes obtained after filtering; The probability score is used to represent the similarity between the speech audio and the preset second language.
2. The method according to claim 1, wherein Extracting audio frame features from the speech audio to be detected includes: Performing signal enhancement processing on the speech audio to generate enhanced speech; Performing frame segmentation on the enhanced speech based on a set frame length to generate a speech sequence; Performing windowing processing on the speech sequence based on a set window length to generate a windowed speech sequence; Performing Fourier transform on the windowed speech sequence to generate a frequency-domain speech signal; Performing filtering processing on the frequency-domain speech signal to generate the audio frame features.
3. The method according to claim 2, wherein Performing signal enhancement processing on the speech audio to generate enhanced speech includes: Obtaining a first signal corresponding to a first moment in the speech audio and a second signal corresponding to a second moment before the first moment; Calculating a weighted signal corresponding to the second signal based on a set signal coefficient and the second signal; Generating an enhanced signal corresponding to the first moment based on the difference between the first signal and the weighted signal; Combining the enhanced signals corresponding to each moment in the speech audio to obtain the enhanced speech.
4. The method according to claim 1, characterized in that Generating a first posterior probability based on the matching degree between the audio frame features and the preset first language phonemes includes: Inputting the audio frame features into a first acoustic model trained based on first language samples, and outputting a first posterior probability corresponding to the matching degree between the audio frame features and the first language phonemes; Identifying the start and end moments corresponding to the phonemes based on the waveforms corresponding to the phonemes in the speech audio; Determining the audio frame features included in the phonemes based on the start and end moments corresponding to the phonemes and the time frame information corresponding to the audio frame features; Calculating the mean value of the first posterior probabilities corresponding to the audio frame features included in the phonemes to generate a first posterior probability that the phonemes correspond to the first language phonemes.
5. The method according to claim 4, wherein Before inputting the audio frame features into a first acoustic model trained based on first language samples and outputting a first posterior probability corresponding to the matching degree between the audio frame features and the first language phonemes, it further includes: Obtaining a first speech sample generated based on a first language and a first speech text corresponding to the first speech sample, and obtaining a second speech sample generated based on a second language and a second speech text corresponding to the second speech sample; Constructing an acoustic model for identifying phonemes included in audio based on a time-delay neural network; Input the first speech sample into the acoustic model, and adjust the parameters of the acoustic model based on the first loss function obtained from the output first phoneme and the first speech text to obtain the first acoustic model; Input the second speech sample into the acoustic model, and adjust the parameters of the acoustic model based on the second loss function obtained from the output second phoneme and the second speech text to obtain the second acoustic model.
6. The method according to claim 1, wherein Generate a second posterior probability based on the matching degree between the audio frame features and the preset second language phonemes, including: Input the audio frame features into the second acoustic model trained based on the second language samples, and output the second posterior probability corresponding to the matching degree between the audio frame features and the second language phonemes; Based on the recognition of the speech audio, determine the start and end times corresponding to the phoneme; Based on the waveform of the speech audio, recognize the start and end times corresponding to the phoneme; Based on the start and end times corresponding to the phoneme, calculate the mean value of the second posterior probabilities corresponding to the audio frame features in the phoneme, and determine the second posterior probability corresponding to the phoneme for the second language phoneme.
7. The method according to claim 1, wherein Determine the pronunciation accuracy level corresponding to each phoneme based on the probability score, including: Based on the probability score corresponding to the phoneme for the second language phoneme, determine the confidence level between the phoneme and the second language phoneme; Based on the confidence level and the set confidence threshold, determine the pronunciation accuracy level corresponding to each phoneme in the speech audio.
8. The method according to claim 7, characterized in that, Based on the probability score corresponding to the phoneme for the second language phoneme, determine the confidence level between the phoneme and the second language phoneme, including: Determine the maximum probability score from the probability scores corresponding to the phoneme for the second language phoneme; Calculate the ratio between the probability score corresponding to the specified phoneme for the second language phoneme and the maximum probability score; Based on the ratio, determine the confidence level between the specified phoneme and the second language phoneme.
9. The method according to claim 1, wherein After calculating the probability score corresponding to the phoneme in the speech audio for the second language phoneme based on the filtered language phonemes, it further includes: Determine the pronunciation accuracy level corresponding to each phoneme based on the probability score, and display the text corresponding to the phoneme based on the display method corresponding to the pronunciation accuracy level.
10. The method according to claim 9, characterized in that, Display the text corresponding to the phoneme based on the display method corresponding to the pronunciation accuracy level, including: Obtain the text corresponding to the speech audio; Based on the phonemes in the speech audio, segment the text to generate the text corresponding to each phoneme; Based on the pronunciation accuracy level corresponding to each phoneme, display the text corresponding to the phoneme through the display method corresponding to the pronunciation accuracy level.
11. The method according to claim 9, wherein After determining the pronunciation accuracy level corresponding to each phoneme based on the probability score and displaying the text corresponding to the phoneme based on the display method corresponding to the pronunciation accuracy level, the method further includes: Query the target phoneme with the lowest pronunciation accuracy level from the pronunciation accuracy levels corresponding to each phoneme; Obtain the pronunciation teaching information corresponding to the target phoneme, where the pronunciation teaching information includes at least one of the following information: phonetic text, correct pronunciation, and demonstration video; Display the pronunciation teaching information.
12. The method according to claim 11, wherein After querying the target phoneme with the lowest pronunciation accuracy level from the pronunciation accuracy levels corresponding to each phoneme, the method further includes: Obtain a target sentence containing the target phoneme from the sentence library corresponding to the second language; Display the target sentence; Obtain a practice audio sent by the user based on the target sentence; Detect the practice audio to obtain the pronunciation accuracy level corresponding to the target phoneme.
13. A pronunciation detection device, characterized in that, Includes: An extraction unit for extracting audio frame features from the speech audio to be detected; A probability unit for generating a first posterior probability based on the matching degree between the audio frame features and the first language phonemes in the preset first language, and generating a second posterior probability based on the matching degree between the audio frame features and the second language phonemes in the preset second language; A scoring unit for filtering out the pronunciation content of the preset first language corresponding to the first posterior probability in the preset second language based on the second posterior probability, and calculating the probability score of the phonemes in the speech audio corresponding to the second language phonemes based on the language phonemes obtained after filtering; The probability score is used to represent the similarity between the speech audio and the preset second language; A display unit for determining the pronunciation accuracy level corresponding to each phoneme based on the probability score, and displaying the text corresponding to the phoneme based on the display method corresponding to the pronunciation accuracy level.
14. A computer-readable medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the pronunciation detection method according to any one of claims 1 to 12.
15. An electronic device, characterized in that, Includes: One or more processors; A storage device for storing one or more programs, which, when executed by the one or more processors, cause the one or more processors to implement the pronunciation detection method according to any one of claims 1 to 12.
16. A computer program product, characterized in that, The computer program product includes a computer program, and the computer program is suitable for being loaded and executed by a processor to implement the pronunciation detection method according to any one of claims 1 to 12.
Citation Information
Patent Citations
Pronunciation bias error detection method and device, storage medium and equipment
CN107610720A
Voice analysis method and device and storage medium
CN109686383A