Pronunciation error detection method and device, speech scoring method and device
By generating multi-scale phoneme aggregate data and combining it with a bidirectional long short-term memory model and an attention model, the accuracy and robustness issues of existing pronunciation error detection schemes are solved, achieving more efficient pronunciation error detection and speech scoring.
Patent Information
- Application Number
- CN202111678431.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-31
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2041-12-31
AI Technical Summary
The existing pronunciation error detection schemes have poor error detection accuracy and robustness, and are particularly affected by interference from similar phonemes within the set and phonemes outside the set.
By determining the state sequence and phoneme time boundary information of the speech to be checked for errors, multi-scale phoneme aggregation data is generated, and aggregation operations are performed using a bidirectional long short-term memory model and an attention model. The error detection loss function is designed in combination with domain expert knowledge to improve the accuracy and stability of error detection.
It improves the accuracy and stability of pronunciation error detection, enhances the error detection capability in noisy environments, reduces the amount of computation and improves the decision-making ability of the model.
Smart Images

Figure CN114495986B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical field of audio processing, and in particular to a pronunciation error detection method and device, and a speech scoring method and device. Background Art
[0002] As we all know, phonemes are the data foundation for implementing specific technical solutions in the field of audio processing technology, especially in the field of pronunciation error detection. Currently, existing pronunciation error detection solutions are mainly based on the GOP (Goodness of Pronunciation) algorithm. However, due to the influence of similar phonemes within the set and phonemes outside the set, the error detection accuracy and robustness of existing pronunciation error detection solutions are poor. Summary of the Invention
[0003] In view of this, the present disclosure provides a pronunciation error detection method and device, and a speech scoring method and device to solve the problem that the existing pronunciation error detection schemes have poor error detection accuracy and error detection robustness.
[0004] In a first aspect, an embodiment of the present disclosure provides a pronunciation error detection method, which includes: determining a state sequence of a speech to be detected for error; determining N phoneme time boundary information corresponding to each phoneme contained in a reading text corresponding to the speech to be detected for error, wherein N is a positive integer; generating phoneme aggregation data based on the state sequence and the N phoneme time boundary information corresponding to each phoneme contained in the reading text; and determining error detection information corresponding to each phoneme contained in the reading text based on the phoneme aggregation data.
[0005] In combination with the first aspect, in certain implementations of the first aspect, phoneme aggregation data is generated based on the state sequence and the N phoneme time boundary information corresponding to each phoneme contained in the read-aloud text, including: performing at least one short-time aggregation operation based on the N phoneme time boundary information corresponding to each phoneme contained in the read-aloud text to obtain short-time aggregation data; determining the full phoneme time boundary information corresponding to the phonemes contained in the read-aloud text based on the N phoneme time boundary information corresponding to each phoneme contained in the read-aloud text; performing at least one long-time aggregation operation based on the full phoneme time boundary information to obtain long-time aggregation data; and generating phoneme aggregation data based on the short-time aggregation data and the long-time aggregation data.
[0006] In combination with the first aspect, in certain implementations of the first aspect, based on the N phoneme time boundary information corresponding to each phoneme contained in the read aloud text, at least one short-time aggregation operation is performed to obtain short-time aggregation data, including: dividing the state sequence based on the N phoneme time boundary information corresponding to each phoneme contained in the read aloud text to obtain the state sequence segments corresponding to each phoneme contained in the read aloud text; based on the state sequence segments corresponding to each phoneme contained in the read aloud text, at least one short-time aggregation operation is performed to obtain the aggregation data corresponding to at least one short-time aggregation operation; and fusing the aggregation data corresponding to at least one short-time aggregation operation to obtain the short-time aggregation data.
[0007] In combination with the first aspect, in certain implementations of the first aspect, the short-time aggregation operation includes: calculating the mean of the state sequence segments corresponding to the phonemes contained in the read text based on the state sequence segments corresponding to the phonemes contained in the read text; and concatenating the mean of the state sequence segments corresponding to the phonemes contained in the read text.
[0008] In combination with the first aspect, in certain implementations of the first aspect, the short-time aggregation operation includes: converting the state sequence segments corresponding to each phoneme contained in the read text to obtain the phoneme sequences corresponding to each phoneme contained in the read text; determining the state sequence segments corresponding to the longest phoneme sequence for each phoneme contained in the read text; calculating the mean of the state sequence segments corresponding to the state sequence segments corresponding to the longest phoneme sequence; and updating the state sequence based on the mean of the state sequence segments corresponding to the state sequence segments corresponding to the longest phoneme sequence.
[0009] In combination with the first aspect, in some implementations of the first aspect, the short-term aggregation operation includes: inputting state sequence segments corresponding to each phoneme contained in the read text into a bidirectional long short-term memory model.
[0010] In conjunction with the first aspect, in certain implementations of the first aspect, before inputting the state sequence segments corresponding to the phonemes contained in the read text into the bidirectional long short-term memory model, the short-term aggregation operation further includes: inputting the state sequence segments corresponding to the phonemes contained in the read text into the attention model. Inputting the state sequence segments corresponding to the phonemes contained in the read text into the bidirectional long short-term memory model includes: inputting model output data of the attention model into the bidirectional long short-term memory model.
[0011] In combination with the first aspect, in some implementations of the first aspect, the long-term aggregation operation includes: inputting the state sequence into a bidirectional long short-term memory model.
[0012] In conjunction with the first aspect, in certain implementations of the first aspect, before inputting the state sequence into the bidirectional long short-term memory model, the long-term aggregation operation further includes: inputting the state sequence into an attention model. Inputting the state sequence into the bidirectional long short-term memory model includes: inputting model output data of the attention model into the bidirectional long short-term memory model.
[0013] In conjunction with the first aspect, in certain implementations of the first aspect, determining error detection information corresponding to each phoneme contained in the text being read aloud based on the phoneme aggregate data includes: utilizing a decision model to determine the error detection information corresponding to each phoneme contained in the text being read aloud based on the phoneme aggregate data. The decision model includes an error detection loss function designed based on domain expert knowledge data, the input of the decision model is the phoneme aggregate data, and the output of the decision model is the error detection information corresponding to each phoneme contained in the text being read aloud.
[0014] In conjunction with the first aspect, in certain implementations of the first aspect, determining a state sequence of the speech to be error-checked for reading aloud includes: inputting the speech to be error-checked for reading aloud into an acoustic model to obtain a state sequence. Furthermore, generating phoneme aggregation data based on the state sequence and N phoneme time boundary information corresponding to each phoneme contained in the read text; and determining error detection information corresponding to each phoneme contained in the read text based on the phoneme aggregation data, including: utilizing the error detection model to obtain the phoneme aggregation data based on the state sequence and N phoneme time boundary information corresponding to each phoneme contained in the read text; and utilizing the error detection model to determine the error detection information corresponding to each phoneme contained in the read text based on the phoneme aggregation data.
[0015] In combination with the first aspect, in certain implementations of the first aspect, during the process of training the error detection model, the training output data of the error detection model is used to update the training acoustic model.
[0016] In a second aspect, an embodiment of the present disclosure provides a speech scoring method, which includes: determining, for a reading speech to be scored and a reading text corresponding to the reading speech to be scored, error detection information corresponding to each phoneme contained in the reading text, wherein the error detection information corresponding to each phoneme contained in the reading text is calculated based on the method mentioned in the first aspect; and determining the speech scoring information of the reading speech to be scored based on the error detection information corresponding to each phoneme contained in the reading text.
[0017] In a third aspect, an embodiment of the present disclosure provides a pronunciation error detection device, which includes: a first determination module for determining a state sequence of a speech to be detected for error; a second determination module for determining N phoneme time boundary information corresponding to each phoneme contained in a reading text corresponding to the speech to be detected for error, wherein N is a positive integer; a generation module for generating phoneme aggregation data based on the state sequence and the N phoneme time boundary information corresponding to each phoneme contained in the reading text; and a third determination module for determining, based on the phoneme aggregation data, the error detection information corresponding to each phoneme contained in the reading text.
[0018] In a fourth aspect, an embodiment of the present disclosure provides a speech scoring device, which includes: a first determination module, used to determine, for a speech to be scored and a reading text corresponding to the speech to be scored, error detection information corresponding to each phoneme contained in the reading text, wherein the error detection information corresponding to each phoneme contained in the reading text is calculated based on the method mentioned in the first aspect; and a second determination module, used to determine speech scoring information of the reading speech to be scored based on the error detection information corresponding to each phoneme contained in the reading text.
[0019] In a fifth aspect, an embodiment of the present disclosure provides an electronic device, comprising: a processor; and a memory for storing instructions executable by the processor, wherein the processor is used to execute the method mentioned in the first aspect and / or the second aspect above.
[0020] In a sixth aspect, an embodiment of the present disclosure provides a computer-readable storage medium. An embodiment of the present disclosure provides a computer-readable storage medium, which stores a computer program for executing the method mentioned in the first aspect and / or the second aspect above.
[0021] Because the phoneme aggregation data is generated by performing a multi-scale aggregation operation on the state sequence based on the N phoneme time boundary information corresponding to each phoneme contained in the read text, the phoneme aggregation data can contain local phoneme information and global phoneme information at different scales. Thus, the disclosed embodiments utilize the phoneme aggregation data to improve the accuracy and stability of the error detection information corresponding to each phoneme contained in the read text, thereby improving the error detection accuracy and error detection stability (also known as error detection robustness). BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 1 Shown is a schematic diagram of an application scenario provided by an embodiment of the present disclosure.
[0023] Figure 2 FIG2 is a flow chart of a pronunciation error detection method provided by an embodiment of the present disclosure.
[0024] Figure 3The figure shows a flowchart of pronunciation error detection provided by an embodiment of the present disclosure.
[0025] Figure 4 The figure shows a flow chart of generating phoneme aggregation data provided by an embodiment of the present disclosure.
[0026] Figure 5 The figure shows a flow chart of obtaining short-term aggregated data provided by an embodiment of the present disclosure.
[0027] Figure 6 The figure shows a flow chart of obtaining short-term aggregated data provided by another embodiment of the present disclosure.
[0028] Figure 7 The figure shows a schematic diagram of a method for generating short-term aggregated data and long-term aggregated data provided by another embodiment of the present disclosure.
[0029] Figure 8 FIG2 is a flow chart of a speech scoring method provided in accordance with an embodiment of the present disclosure.
[0030] Figure 9 FIG2 is a schematic diagram of the structure of a pronunciation error detection device provided by an embodiment of the present disclosure.
[0031] Figure 10 FIG2 is a schematic diagram of the structure of a speech scoring device provided by an embodiment of the present disclosure.
[0032] Figure 11 Shown is a structural schematic diagram of an electronic device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION
[0033] The technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the drawings in the embodiments of the present disclosure. Obviously, the described embodiments are only part of the embodiments of the present disclosure, rather than all the embodiments.
[0034] Phonemes are the smallest units of speech divided according to the natural properties of speech. From the perspective of acoustic properties, phonemes are the smallest units of speech divided from the perspective of sound quality. From the perspective of physiological properties, one pronunciation action forms a phoneme. Based on the above properties of phonemes, it can be seen that the importance of phonemes in technical fields such as pronunciation error detection and speech recognition is self-evident. Especially in the field of pronunciation error detection, phonemes are the basis for the implementation of pronunciation error detection schemes. At present, existing pronunciation error detection schemes are mainly implemented based on the GOP algorithm. The basic idea of the GOP algorithm is to force the alignment of the speech to be detected and the text to be read, and then determine the likelihood ratio based on the likelihood score value obtained by the forced alignment, and then use the likelihood ratio (LR) to evaluate the quality of pronunciation.
[0035] However, since similar phonemes within the set are difficult to distinguish, and the GOP algorithm is easily interfered with by phonemes outside the set caused by noise, the accuracy and stability of the GOP algorithm are relatively poor. On this basis, the error detection accuracy and stability of the pronunciation error detection scheme that relies on the GOP algorithm are also poor.
[0036] In order to solve the above problems, the embodiments of the present disclosure provide a pronunciation error detection method and device, and a speech scoring method and device to achieve the purpose of improving the error detection accuracy and error detection stability.
[0037] The following combination Figure 1 The application scenarios of the embodiments of the present disclosure are briefly introduced.
[0038] Figure 1 The figure shows an application scenario diagram provided by an embodiment of the present disclosure. Figure 1 As shown, this application scenario is a pronunciation error detection scenario. Specifically, the pronunciation error detection scenario includes a server 110 and a user terminal 120 in communication with the server 110. The server 110 stores a text to be read aloud, and the server 110 is used to execute the pronunciation error detection method mentioned in the embodiment of the present disclosure. The user terminal 120 can be a user's terminal device. It is understood that the user can be a user practicing oral English and / or a user taking an oral English test, and the present embodiment of the present disclosure will not further describe them one by one.
[0039] For example, in actual application, the user uses the user terminal 120 to record the speech to be read aloud for error detection, and sends the speech to be read aloud for error detection to the server 110 via the user terminal 120. After receiving the speech to be read aloud for error detection, the server 110 first determines the state sequence of the speech to be read aloud for error detection, then determines the N phoneme time boundary information (such as triphone time boundary information) corresponding to each phoneme contained in the reading text corresponding to the speech to be read aloud for error detection, and then generates phoneme aggregation data based on the state sequence and the N phoneme time boundary information corresponding to each phoneme contained in the reading text, and finally determines the error detection information corresponding to each phoneme contained in the reading text based on the phoneme aggregation data. In addition, the server 110 sends the error detection information corresponding to each phoneme contained in the reading text to the user terminal 120. After receiving the error detection information corresponding to each phoneme contained in the reading text, the user terminal 120 presents the error detection information corresponding to each phoneme contained in the reading text to the user so that the user can know whether the pronunciation is accurate and other information.
[0040] Exemplarily, the user terminal 120 mentioned above is a mobile terminal such as a tablet computer or a mobile phone.
[0041] The following combination Figures 2 to 7 The pronunciation error detection method disclosed in the present invention is briefly introduced.
[0042] Figure 2FIG. 1 is a flow chart of a pronunciation error detection method provided by an embodiment of the present disclosure. Figure 2 As shown, the pronunciation error detection method provided by the embodiment of the present disclosure includes the following steps.
[0043] Step S210: determining a state sequence of the speech to be checked for errors.
[0044] For example, the speech to be read aloud with errors is the speech read aloud by the oral practice user with reference to the reading text. It can be understood that the speech to be read aloud with errors can be collected by a pickup (such as a microphone) of the user terminal.
[0045] The state sequence is also called a speech sequence. Exemplarily, a state sequence model is used to determine the state sequence of the speech to be read aloud with errors. The state sequence model is a neural network model, the input of the state sequence model is the speech to be read aloud with errors, and the output of the state sequence model is the state sequence of the speech to be read aloud with errors.
[0046] In some embodiments, determining the state sequence of the speech to be detected as an error in reading aloud may be performed by obtaining the state sequence of the speech to be detected as an error in reading aloud.
[0047] Step S220 , determining N phoneme time boundary information corresponding to each phoneme contained in the reading text corresponding to the speech to be detected for error.
[0048] Exemplarily, N is a positive integer. For example, N is equal to three, that is, the N-phoneme time boundary information is triphone time boundary information. A value of three for N can significantly improve error detection. Furthermore, N can also be equal to five, which is not uniformly limited in the present embodiment.
[0049] The reading text corresponding to the error-checked reading voice refers to the reference reading text of the error-checked reading voice. In actual application, the user reads the reading text presented by the user terminal, and the user terminal collects the voice emitted by the user to obtain the error-checked reading voice.
[0050] In some embodiments, determining the monophone time boundary information corresponding to each phoneme contained in the reading text corresponding to the to-be-detected error reading voice can be performed by obtaining the monophone time boundary information corresponding to each phoneme contained in the reading text corresponding to the to-be-detected error reading voice.
[0051] Step S230 : Generate phoneme aggregation data based on the state sequence and N phoneme time boundary information corresponding to each of the phonemes included in the read text.
[0052] Phoneme aggregation data, also known as phoneme aggregation data, is the data obtained by performing a series of multi-scale phoneme aggregation operations on the state sequence based on the N phoneme time boundary information corresponding to each phoneme contained in the read text (i.e., the data obtained by executing a multi-scale fusion strategy on the state sequence). For example, the multi-scale phoneme aggregation operations mentioned above include aggregation operations at different time scales, such as short-term aggregation operations and long-term aggregation operations (the specific implementation methods of short-term aggregation operations and long-term aggregation operations can be found in the following embodiments).
[0053] Step S240 : determining the error detection information corresponding to each phoneme contained in the read text based on the phoneme aggregation data.
[0054] For example, if the read text contains 20 phonemes, each of the 20 phonemes corresponds to an error detection information. The error detection information can be a numerical value between 0 and 1, where 1 indicates complete correctness and 0 indicates complete error.
[0055] For example, in actual application, the state sequence of the speech to be detected for error is first determined, and then the N phoneme time boundary information corresponding to each phoneme contained in the reading text corresponding to the speech to be detected for error is determined. Then, based on the state sequence and the N phoneme time boundary information corresponding to each phoneme contained in the reading text, phoneme aggregation data is generated, and based on the phoneme aggregation data, the error detection information corresponding to each phoneme contained in the reading text is determined.
[0056] Because the phoneme aggregation data is generated by performing a multi-scale aggregation operation on the state sequence based on the N phoneme time boundary information corresponding to each phoneme contained in the read text, the phoneme aggregation data can contain local phoneme information and global phoneme information at different scales. Thus, the disclosed embodiments utilize the phoneme aggregation data to improve the accuracy and stability of the error detection information corresponding to each phoneme contained in the read text, thereby improving the error detection accuracy and error detection stability (also known as error detection robustness).
[0057] It should be noted that some steps in the pronunciation error detection method mentioned in the above embodiment can be implemented with the help of a neural network model with corresponding functions. Figure 3 Give an example.
[0058] Figure 3 The figure shows a flow chart of pronunciation error detection provided by an embodiment of the present disclosure. Figure 3 As shown, in the embodiment of the present disclosure, Figure 1 In the illustrated embodiment, step S210 is implemented with the aid of an acoustic model, and, Figure 1 In the embodiment shown, steps S230 and S240 are implemented with the aid of an error detection model comprising an aggregation layer and a decision layer. Figure 1In the illustrated embodiment, step S220 is implemented by a segmentation alignment algorithm, which is used to time-align the phoneme sequences of the error-checked speech and the text to be read aloud, thereby ultimately obtaining N phoneme time boundary information corresponding to each phoneme contained in the text to be read aloud.
[0059] Specifically, the step of determining the state sequence of the speech to be read aloud with errors is performed as follows: inputting the speech to be read aloud with errors into an acoustic model to obtain a state sequence. The acoustic model is a neural network model, the input of the acoustic model is the speech to be read aloud with errors, and the output of the acoustic model is the state sequence of the speech to be read aloud with errors. In addition, based on the state sequence and the N phoneme time boundary information corresponding to each phoneme contained in the read aloud text, phoneme aggregation data is generated; based on the phoneme aggregation data, the step of determining the error detection information corresponding to each phoneme contained in the read aloud text is performed as follows: using the error detection model (i.e., using the aggregation layer in the error detection model), based on the state sequence and the N phoneme time boundary information corresponding to each phoneme contained in the read aloud text, the phoneme aggregation data is obtained; and using the error detection model (i.e., using the decision layer in the error detection model), based on the phoneme aggregation data, the error detection information corresponding to each phoneme contained in the read aloud text is determined. Among them, the error detection model is a neural network model, the input of the error detection model is the state sequence of the speech to be detected and the N phoneme time boundary information corresponding to each phoneme contained in the reading text, and the output of the error detection model is the error detection information corresponding to each phoneme contained in the reading text.
[0060] It should be noted that the aggregation layer can also be designed as an independent aggregation model, and the decision layer can also be designed as an independent decision model to improve the flexibility of the solution implementation. The embodiments of the present disclosure do not make unified limitations on this.
[0061] The disclosed embodiments implement a multi-model hybrid pronunciation error detection scheme, which not only fully utilizes the advantages of the models themselves, but also utilizes the data association relationship between models. In particular, during the training phase, some models can learn to utilize the data of other models, thereby further improving the accuracy and robustness of the model output data, thereby improving the error detection accuracy and error detection robustness. For example, in some embodiments, during the process of training the error detection model, the training output data of the error detection model is used to update the training acoustic model. Such a setting can enable the acoustic model to fully learn and utilize the data of the error detection model during the training process, thereby further optimizing the effect of the acoustic model.
[0062] In some embodiments, the decision model includes an error detection loss function designed based on domain expert knowledge data. The domain expert knowledge data refers to prior knowledge data that is conducive to improving the accuracy of error detection. For example, the language domain of the speech to be detected for error reading is English, and English includes 48 phonemes. Then, for a pronunciation sequence segment in the speech to be detected for error reading determined based on the time boundary information of the three phonemes, it should be considered that the pronunciation sequence segment corresponds to one of the 48 phonemes. However, if combined with the domain expert knowledge data, it can be determined that the pronunciation sequence segment can only belong to one of the 10 phonemes among the 48 phonemes, that is, the combination of domain expert knowledge data greatly narrows the scope of discrimination. It can be seen that the combination of domain expert knowledge data can not only effectively improve the accuracy of error detection, improve the ability of the decision model to remove interference items, thereby improving the error detection effect in a noisy environment, but also greatly reduce the amount of calculation.
[0063] In some embodiments, the decision model is a sigmoid model, and the model uses cross entropy loss as a loss function. More specifically, the data dimension of the sigmoid model input is P*S, where P is the number of phonemes in the read text, and S is the number of all phonemes in the language to be detected (if the language to be detected is English, then S=48). In actual application, the sigmoid model can reorder the S phonemes based on the domain expert knowledge data. After reordering, the first dimension is the score of the phoneme to be detected, and the remaining S-1 dimension is sorted according to the possible error type of the phoneme to be detected. The sorted data dimension is still P*S. Such a design can improve the decision-making ability of the model and thus improve the accuracy of error detection.
[0064] The following combination Figures 4 to 7 An example is given to illustrate how to generate phoneme aggregation data.
[0065] Figure 4 The figure shows a flow chart of generating phoneme aggregation data provided by an embodiment of the present disclosure. Figure 2 Based on the embodiment shown Figure 4 The embodiment shown is described below in detail. Figure 4 The embodiment shown is Figure 2 The differences and similarities between the illustrated embodiments are not described in detail.
[0066] like Figure 4 As shown, in the embodiment of the present disclosure, based on the state sequence and the N phoneme time boundary information corresponding to each phoneme contained in the read text, the step of generating phoneme aggregation data includes the following steps.
[0067] Step S410 : performing at least one short-term aggregation operation based on N phoneme time boundary information corresponding to each phoneme contained in the read text to obtain short-term aggregation data.
[0068] Short-term aggregation refers to local phoneme aggregation, that is, the aggregation of local phonemes (i.e., partial phonemes) among all the phonemes contained in the read text. For example, the short-term aggregation operation is performed using the state sequence segment corresponding to each three-phoneme time boundary as the aggregation unit.
[0069] The short-term aggregation operation can be performed multiple times, and the time scales of the multiple short-term aggregation operations are different from each other, so that multiple groups of short-term aggregation data with different time scales can be obtained. For example, the short-term aggregation operation is performed multiple times. In actual application, the triphone sequence of the read text is segmented. Specifically, N phonemes are used as a window length, and the triphone sequence is segmented in a sliding window manner. Then, based on multiple short-term aggregation operations, multiple short-term phoneme embedding values corresponding to each segmented triphone sequence are obtained. For each triphone sequence, the short-term phoneme embedding values of the corresponding positions of the triphone sequence are summed to obtain the sum of the short-term phoneme embeddings corresponding to the triphone sequence. Finally, the sum of the short-time phoneme embeddings corresponding to all triphone sequences is spliced together to obtain the short-term aggregation data.
[0070] Step S420 : determining the full phoneme time boundary information corresponding to the phonemes contained in the read text based on the N phoneme time boundary information corresponding to each phoneme contained in the read text.
[0071] The full phoneme time boundary information refers to the overall time boundary information corresponding to all phonemes contained in the read text.
[0072] Step S430 : performing at least one long-term aggregation operation based on the full phoneme time boundary information to obtain long-term aggregation data.
[0073] For example, the long-term aggregation operation refers to a global phoneme aggregation operation, that is, an aggregation operation performed on all phonemes contained in the read text. Furthermore, the long-term aggregation operation can be performed multiple times, each in a different manner. The long-term phoneme embedding values corresponding to each of the multiple long-term aggregation operations are summed to obtain the long-term aggregated data.
[0074] Step S440: Generate phoneme aggregation data based on the short-term aggregation data and the long-term aggregation data.
[0075] Exemplarily, the short-term aggregated data and the long-term aggregated data are fused in a splicing manner to obtain the phoneme aggregated data.
[0076] The pronunciation error detection method provided by the embodiment of the present disclosure takes into account the local information and global information contained in the state sequence and time boundary information with the help of short-term aggregated data and long-term aggregated data, generates phoneme aggregated data containing multi-dimensional information, and thus provides a data basis for improving the error detection accuracy and error detection robustness.
[0077] Figure 5 The figure shows a flow chart of obtaining short-term aggregated data provided by an embodiment of the present disclosure. Figure 4 Based on the embodiment shown Figure 5 The embodiment shown is described below in detail. Figure 5 The embodiment shown is Figure 4 The differences and similarities between the illustrated embodiments are not described in detail.
[0078] like Figure 5 As shown, in an embodiment of the present disclosure, based on the N phoneme time boundary information corresponding to each phoneme contained in the read text, at least one short-time aggregation operation is performed to obtain short-time aggregation data, which includes the following steps.
[0079] Step S510 : dividing the state sequence based on N phoneme time boundary information corresponding to each phoneme included in the read-aloud text to obtain state sequence segments corresponding to each phoneme included in the read-aloud text.
[0080] It can be understood that by dividing the state sequence based on the N phoneme time boundary information corresponding to each phoneme, a state sequence segment corresponding to the phoneme can be obtained.
[0081] Step S520: Based on the state sequence segments corresponding to the phonemes contained in the read text, at least one short-term aggregation operation is performed to obtain aggregated data corresponding to each of the at least one short-term aggregation operations. In other words, the data obtained after each short-term aggregation operation is the aggregated data.
[0082] Step S530: Merge the aggregated data corresponding to at least one short-term aggregation operation to obtain short-term aggregated data.
[0083] Exemplarily, the fusion mentioned in step S530 refers to summing the short-time phoneme embedding values of all short-time fusion operations corresponding to each phoneme to obtain the sum of the short-time phoneme embeddings, and then concatenating the sum of the short-time phoneme embeddings corresponding to all phonemes to obtain short-time aggregated data.
[0084] The embodiment of the disclosure generates multi-dimensional short-time aggregation data by dividing the state sequence based on the N-phoneme time boundary information corresponding to each phoneme contained in the read text, obtaining the state sequence segment corresponding to each phoneme contained in the read text, performing at least one short-time aggregation operation based on the state sequence segment corresponding to each phoneme contained in the read text, obtaining the aggregation data corresponding to each short-time aggregation operation, and then fusing the aggregation data corresponding to each short-time aggregation operation to obtain short-time aggregation data.
[0085] To further clarify the meaning of short-time aggregation operation, the following will be combined with Figure 6 The four short-time aggregation operations each have a specific implementation method. It can be understood that in actual application, at least one of the four short-time aggregation operations shown in the figure can be selected according to actual conditions, and the embodiment of the disclosure does not make unified limitation. Figure 6 The four short-time aggregation operations each have a specific implementation method. It can be understood that in actual application, at least one of the four short-time aggregation operations shown in the figure can be selected according to actual conditions, and the embodiment of the disclosure does not make unified limitation.
[0086] Figure 6 The embodiment of the disclosure another embodiment provides a flowchart for obtaining short-time aggregation data. In Figure 5 The embodiment shown in the figure is extended from Figure 6 The embodiment shown in the figure is extended from Figure 6 The embodiment shown in the figure is extended from Figure 5 The embodiment shown in the figure is extended from
[0087] As Figure 6 The embodiment of the disclosure mentions four short-time aggregation operations, namely short-time aggregation operations one to four, and the specific implementation methods of the four short-time aggregation operations will be introduced below.
[0088] Short-time aggregation operation one (corresponding to steps S610 and S620)
[0089] Step S610, based on the state sequence segment corresponding to each phoneme contained in the read text, calculate the mean value of the state sequence segment corresponding to each phoneme contained in the read text.
[0090] Step S620, splice the mean value of the state sequence segment corresponding to each phoneme contained in the read text to obtain the corresponding aggregation data.
[0091] That is, in the state sequence, the mean value of the audio embedding in the same N-phoneme time boundary is taken, and then the mean value of the state sequence segment corresponding to each phoneme contained in the read text is spliced to obtain the aggregation data corresponding to the short-time aggregation operation.
[0092] Short-time aggregation operation two (corresponding to steps S630 to S660)
[0093] Step S630: convert the state sequence segments corresponding to the phonemes contained in the read text to obtain the phoneme sequences corresponding to the phonemes contained in the read text. It is understood that the state sequence segments can be converted into phoneme sequences.
[0094] Step S640: for each of the phoneme sequences corresponding to the phonemes contained in the read text, determine the state sequence segment corresponding to the longest phoneme sequence. The longest phoneme sequence refers to the phoneme sequence with the longest continuous phoneme length.
[0095] Step S650: Calculate the mean of the state sequence segments corresponding to the longest phoneme sequence. If there are multiple longest phoneme sequences in parallel, the mean of the state sequence segments corresponding to each longest phoneme sequence can be calculated separately.
[0096] Step S660 : Based on the state sequence segment mean corresponding to the state sequence segment corresponding to the longest phoneme sequence, the state sequence is updated to obtain corresponding aggregated data.
[0097] That is, in the state sequence, the audio embeddings within the same N-phoneme time boundary are converted into phoneme sequences, and the audio embeddings corresponding to the phoneme sequence with the longest continuous phoneme length are averaged to obtain the aggregated embedding representation (i.e., the corresponding aggregated data).
[0098] Short-term aggregation operation three (corresponding to step S670)
[0099] In step S670 , the state sequence segments corresponding to the phonemes contained in the read text are input into a bidirectional long short-term memory (BLSTM) model to obtain corresponding aggregated data.
[0100] For example, in actual application, in the state sequence, after the audio embedding within the same N-phoneme time boundary (i.e., the state sequence segment corresponding to a phoneme) passes through the BLSTM model, the last layer of the output model is used as the aggregated data corresponding to the short-time aggregation operation (also known as the fusion strategy).
[0101] Short-term aggregation operation four (corresponding to steps S680 and S690)
[0102] In step S680 , the state sequence segments corresponding to the phonemes contained in the read text are input into the attention model (AM).
[0103] Step S690: Input the model output data of the attention model into the bidirectional long short-term memory model to obtain corresponding aggregated data.
[0104] That is to say, in the short-term aggregation operation four, the AM model and the BLSTM model connected to the AM model are used to further enrich the amount of information contained in the obtained aggregated data.
[0105] By utilizing the four short-term aggregation operations mentioned in the embodiments of the present disclosure, aggregated data from different aggregation paths can be obtained, thereby further enriching the amount of information in the short-term aggregation data, and providing a prerequisite for subsequently improving the error detection accuracy and error detection robustness.
[0106] It should be noted that the long-term aggregation operation mentioned in the above embodiment can be performed with reference to the short-term aggregation operation mentioned above, and the embodiments of the present disclosure will not be repeated. For example, in some embodiments, the long-term aggregation operation includes: inputting the state sequence into the BLSTM model. Furthermore, in some embodiments, before inputting the state sequence into the BLSTM model, it also includes: inputting the state sequence into the AM model. In this case, inputting the state sequence into the BLSTM model includes: inputting the model output data of the AM model into the BLSTM model.
[0107] The following combination Figure 7 Give examples to illustrate how to generate short-term aggregate data and long-term aggregate data. Figure 7 FIG. 1 is a schematic diagram of a method for generating short-term aggregated data and long-term aggregated data according to another embodiment of the present disclosure. Figure 7 As shown, based on the state sequence of the read speech containing time boundary information, short-term aggregation data on the left and long-term aggregation data on the right are generated respectively.
[0108] Figure 8 FIG. 1 is a flow chart of a speech scoring method according to an embodiment of the present disclosure. Figure 8 As shown, the speech scoring method provided by the embodiment of the present disclosure includes the following steps.
[0109] Step S810 : determining error detection information corresponding to each phoneme contained in the read text for the speech to be scored and the read text corresponding to the speech to be scored.
[0110] For example, the error detection information corresponding to each phoneme contained in the read text is calculated based on the pronunciation error detection method mentioned in any of the above embodiments. It can be understood that when the pronunciation error detection method mentioned in any of the above embodiments is applied, the read speech to be scored mentioned in step S810 can be regarded as the read speech to be error-checked mentioned in the above embodiments.
[0111] Step S820: Determine the speech scoring information of the speech to be scored based on the error detection information corresponding to each phoneme contained in the read text. For example, the speech scoring information is percentage score information.
[0112] The speech scoring method provided by the embodiments of the present disclosure can be applied to oral practice and / or oral examination scenarios.
[0113] The disclosed embodiments achieve the purpose of speech scoring by determining error detection information corresponding to each phoneme contained in the speech to be scored and the corresponding text of the speech to be scored, and then determining speech scoring information for the speech to be scored based on the error detection information corresponding to each phoneme contained in the text. Because the error detection information corresponding to each phoneme contained in the text is both highly accurate and robust, the speech scoring information determined using the disclosed embodiments is also relatively accurate.
[0114] Combined with the above Figures 2 to 8 , describes the method embodiment of the present disclosure in detail, and the following is combined with Figures 9 to 11 , describes the device embodiment of the present disclosure in detail. In addition, it should be understood that the description of the method embodiment corresponds to the description of the device embodiment, so that parts not described in detail can refer to the previous method embodiment.
[0115] Figure 9 The figure shows a schematic diagram of the structure of a pronunciation error detection device provided by an embodiment of the present disclosure. Figure 9 As shown, the pronunciation error detection device provided by the embodiment of the present disclosure includes a first determination module 910, a second determination module 920, a generation module 930 and a third determination module 940. Specifically, the first determination module 910 is used to determine the state sequence of the speech to be read aloud for error detection. The second determination module 920 is used to determine the N phoneme time boundary information corresponding to each phoneme contained in the reading text corresponding to the speech to be read aloud for error detection, where N is a positive integer. The generation module 930 is used to generate phoneme aggregation data based on the state sequence and the N phoneme time boundary information corresponding to each phoneme contained in the reading text. The third determination module 940 is used to determine the error detection information corresponding to each phoneme contained in the reading text based on the phoneme aggregation data.
[0116] In some embodiments, the generation module 930 is also used to perform at least one short-time aggregation operation based on the N phoneme time boundary information corresponding to each phoneme contained in the read-aloud text to obtain short-time aggregation data; determine the full phoneme time boundary information corresponding to the phonemes contained in the read-aloud text based on the N phoneme time boundary information corresponding to each phoneme contained in the read-aloud text; perform at least one long-time aggregation operation based on the full phoneme time boundary information to obtain long-time aggregation data; and generate phoneme aggregation data based on the short-time aggregation data and the long-time aggregation data.
[0117] In some embodiments, the generation module 930 is also used to divide the state sequence based on the N phoneme time boundary information corresponding to each phoneme contained in the read text, and obtain the state sequence segments corresponding to each phoneme contained in the read text; perform at least one short-time aggregation operation based on the state sequence segments corresponding to each phoneme contained in the read text, and obtain the aggregation data corresponding to at least one short-time aggregation operation; and fuse the aggregation data corresponding to at least one short-time aggregation operation to obtain short-time aggregation data.
[0118] In some embodiments, the short-term aggregation operation includes: calculating the mean of the state sequence segments corresponding to the phonemes contained in the read text based on the state sequence segments corresponding to the phonemes contained in the read text; and concatenating the mean of the state sequence segments corresponding to the phonemes contained in the read text.
[0119] In some embodiments, the short-term aggregation operation includes: converting the state sequence segments corresponding to each phoneme contained in the read text to obtain the phoneme sequences corresponding to each phoneme contained in the read text; determining the state sequence segments corresponding to the longest phoneme sequence for each phoneme contained in the read text; calculating the mean of the state sequence segments corresponding to the state sequence segments corresponding to the longest phoneme sequence; and updating the state sequence based on the mean of the state sequence segments corresponding to the state sequence segments corresponding to the longest phoneme sequence.
[0120] In some embodiments, the short-term aggregation operation includes: inputting the state sequence segments corresponding to the phonemes contained in the read text into the bidirectional long short-term memory model.
[0121] In some embodiments, before inputting the state sequence segments corresponding to the phonemes contained in the read text into the bidirectional long short-term memory model, the short-term aggregation operation further includes: inputting the state sequence segments corresponding to the phonemes contained in the read text into the attention model. Inputting the state sequence segments corresponding to the phonemes contained in the read text into the bidirectional long short-term memory model includes: inputting model output data of the attention model into the bidirectional long short-term memory model.
[0122] In some embodiments, the long-term aggregation operation includes: inputting the state sequence into a bidirectional long short-term memory model.
[0123] In some embodiments, before inputting the state sequence into the bidirectional LSTM model, the long-term aggregation operation further includes: inputting the state sequence into an attention model. Inputting the state sequence into the bidirectional LSTM model includes: inputting model output data of the attention model into the bidirectional LSTM model.
[0124] In some embodiments, the third determination module 940 is further configured to determine, using a decision model and based on the phoneme aggregation data, error detection information corresponding to each phoneme contained in the text being read aloud. The decision model includes an error detection loss function designed based on domain expert knowledge data, the input of the decision model is the phoneme aggregation data, and the output of the decision model is error detection information corresponding to each phoneme contained in the text being read aloud.
[0125] In some embodiments, the first determination module 910 is further configured to input the speech to be error-checked into the acoustic model to obtain a state sequence. Furthermore, the third determination module 940 is further configured to utilize the error detection model to obtain phoneme aggregation data based on the state sequence and N phoneme time boundary information corresponding to each phoneme contained in the read text; and utilize the error detection model to determine error detection information corresponding to each phoneme contained in the read text based on the phoneme aggregation data.
[0126] In some embodiments, during the process of training the error detection model, the training output data of the error detection model is used to update the training acoustic model.
[0127] Figure 10 The figure shows a schematic diagram of the structure of a speech scoring device provided by an embodiment of the present disclosure. Figure 10 As shown, the speech scoring device provided by the embodiment of the present disclosure includes a first determination module 1010 and a second determination module 1020. Specifically, the first determination module 1010 is used to determine, for the speech to be scored and the reading text corresponding to the speech to be scored, the error detection information corresponding to each phoneme contained in the reading text, wherein the error detection information corresponding to each phoneme contained in the reading text is calculated based on the pronunciation error detection method mentioned in the above embodiment. The second determination module 1020 is also used to determine the speech scoring information of the speech to be scored based on the error detection information corresponding to each phoneme contained in the reading text.
[0128] Figure 11 Shown is a structural schematic diagram of an electronic device provided by an embodiment of the present disclosure. Figure 11 The electronic device 1100 shown (which may be a computer device) includes a memory 1101, a processor 1102, a communication interface 1103, and a bus 1104. The memory 1101, the processor 1102, and the communication interface 1103 are communicatively connected to each other via the bus 1104.
[0129] The memory 1101 may be a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 1101 may store programs. When the program stored in the memory 1101 is executed by the processor 1102, the processor 1102 and the communication interface 1103 are used to perform the various steps of the pronunciation error detection method and / or the speech scoring method of the embodiments of the present disclosure.
[0130] The processor 1102 can adopt a general-purpose central processing unit (CPU), a microprocessor, an application-specific integrated circuit (ASIC), a graphics processing unit (GPU) or one or more integrated circuits to execute relevant programs to implement the functions required to be performed by the units in the pronunciation error detection device and / or the speech scoring device of the embodiment of the present disclosure.
[0131] The processor 1102 may also be an integrated circuit chip with signal processing capabilities. During implementation, the various steps of the pronunciation error detection method and / or speech scoring method disclosed herein may be performed by hardware integrated logic circuits or software instructions in the processor 1102. The aforementioned processor 1102 may also be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The various methods, steps, and logic block diagrams disclosed in the embodiments of the present disclosure may be implemented or executed. The general-purpose processor may be a microprocessor or any conventional processor. The steps of the methods disclosed in conjunction with the embodiments of the present disclosure may be directly implemented and executed by a hardware decoding processor, or by a combination of hardware and software modules in the decoding processor. The software modules may be located in a storage medium well-known in the art, such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, an electrically erasable programmable memory, or a register. The storage medium is located in the memory 1101, and the processor 1102 reads the information in the memory 1101 and combines its hardware to complete the functions required to be performed by the units included in the pronunciation error detection device and / or speech scoring device of the embodiment of the present disclosure, or executes the pronunciation error detection method and / or speech scoring method of the method embodiment of the present disclosure.
[0132] The communication interface 1103 uses a transceiver such as, but not limited to, a transceiver to implement communication between the electronic device 1100 and other devices or a communication network. For example, the communication interface 1103 can be used to obtain a text to be read aloud.
[0133] The bus 1104 may include a path for transmitting information between various components of the electronic device 1100 (eg, the memory 1101 , the processor 1102 , and the communication interface 1103 ).
[0134] It should be noted that although Figure 11 The electronic device 1100 shown only shows a memory, a processor, and a communication interface. However, in the specific implementation process, those skilled in the art should understand that the electronic device 1100 also includes other devices necessary for normal operation. At the same time, according to specific needs, those skilled in the art should understand that the electronic device 1100 may also include hardware devices that implement other additional functions. In addition, those skilled in the art should understand that the electronic device 1100 may also only include the devices necessary to implement the embodiments of the present disclosure, and does not necessarily include Figure 11 All devices shown in .
[0135] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this disclosure.
[0136] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0137] In the several embodiments provided in the present disclosure, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0138] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0139] In addition, each functional unit in each embodiment of the present disclosure may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
[0140] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present disclosure, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present disclosure. The aforementioned storage medium includes various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory, a random access memory, a magnetic disk, or an optical disk.
[0141] The above description is merely a specific embodiment of the present disclosure, but the scope of protection of the present disclosure is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this disclosure should be included in the scope of protection of the present disclosure. Therefore, the scope of protection of the present disclosure should be based on the scope of protection of the claims.
Claims
1. A pronunciation error detection method, characterized in that: include: Determine a state sequence of the speech to be detected for error reading; Determine N phoneme time boundary information corresponding to each phoneme contained in the reading text corresponding to the error-detected reading voice, where N is a positive integer; performing aggregation operations of different time scales on the state sequence based on N phoneme time boundary information corresponding to each phoneme contained in the read text to generate phoneme aggregation data; Determining error detection information corresponding to each of the phonemes included in the read text based on the phoneme aggregation data; The aggregation operations at different time scales include short-time aggregation operations, and the phoneme aggregation data include short-time aggregation data; performing aggregation operations at different time scales on the state sequence based on N phoneme time boundary information corresponding to each phoneme contained in the read text to generate phoneme aggregation data includes: Dividing the state sequence based on N phoneme time boundary information corresponding to each phoneme included in the read-aloud text to obtain state sequence segments corresponding to each phoneme included in the read-aloud text; Based on the state sequence segments corresponding to the phonemes contained in the read text, performing at least one short-time aggregation operation to obtain aggregated data corresponding to the at least one short-time aggregation operation; The aggregated data corresponding to at least one of the short-term aggregation operations are fused to obtain the short-term aggregated data.
2. The pronunciation error detection method according to claim 1, wherein The aggregation operations at different time scales further include a long-term aggregation operation, and the phoneme aggregation data further include long-term aggregation data; performing aggregation operations at different time scales on the state sequence based on N phoneme time boundary information corresponding to each phoneme contained in the read text to generate phoneme aggregation data further includes: Determining full phoneme time boundary information corresponding to the phonemes contained in the read text based on N phoneme time boundary information corresponding to each phoneme contained in the read text; Based on the full phoneme time boundary information, performing at least one long-term aggregation operation to obtain the long-term aggregation data; The phoneme aggregate data is generated based on the short-term aggregate data and the long-term aggregate data.
3. The pronunciation error detection method according to claim 1, wherein: The short-term aggregation operation includes: Calculating the mean of the state sequence segments corresponding to the phonemes contained in the read aloud text based on the state sequence segments corresponding to the phonemes contained in the read aloud text; The mean values of the state sequence segments corresponding to the phonemes contained in the read text are concatenated.
4. The pronunciation error detection method according to claim 1, wherein: The short-term aggregation operation includes: Converting the state sequence segments corresponding to the phonemes contained in the read-aloud text to obtain the phoneme sequences corresponding to the phonemes contained in the read-aloud text; For each phoneme sequence corresponding to each phoneme in the read text, determining a state sequence segment corresponding to the longest phoneme sequence; Calculating the mean of the state sequence segments corresponding to the state sequence segments corresponding to the longest phoneme sequence; The state sequence is updated based on the mean of the state sequence segments corresponding to the state sequence segments corresponding to the longest phoneme sequence.
5. The pronunciation error detection method according to claim 1, wherein: The short-term aggregation operation includes: Inputting the state sequence segments corresponding to the phonemes contained in the read text into the attention model; The model output data of the attention model is input into the bidirectional long short-term memory model.
6. The pronunciation error detection method according to claim 2, wherein: The long-term aggregation operation includes: inputting the state sequence into a bidirectional long short-term memory model.
7. The pronunciation error detection method according to claim 6, wherein: Before inputting the state sequence into the bidirectional long short-term memory model, the method further includes: Inputting the state sequence into an attention model; The step of inputting the state sequence into a bidirectional long short-term memory model includes: The model output data of the attention model is input into the bidirectional long short-term memory model.
8. The pronunciation error detection method according to any one of claims 1 to 7, characterized in that: The determining, based on the phoneme aggregation data, error detection information corresponding to each phoneme contained in the read text includes: A decision model is used to determine the error detection information corresponding to each phoneme contained in the read-aloud text based on the phoneme aggregation data, wherein the decision model includes an error detection loss function designed based on domain expert knowledge data, the input of the decision model is the phoneme aggregation data, and the output of the decision model is the error detection information corresponding to each phoneme contained in the read-aloud text.
9. The pronunciation error detection method according to any one of claims 1 to 7, characterized in that: The state sequence of determining the speech to be read aloud with errors includes: Inputting the error-checked speech into the acoustic model to obtain the state sequence; Furthermore, generating phoneme aggregation data based on the state sequence and N phoneme time boundary information corresponding to each phoneme included in the read text, and determining error detection information corresponding to each phoneme included in the read text based on the phoneme aggregation data, including: Obtaining the phoneme aggregation data based on the state sequence and N phoneme time boundary information corresponding to the phonemes contained in the read text using an error detection model; The error detection model is used to determine error detection information corresponding to each phoneme contained in the read text based on the phoneme aggregation data.
10. The pronunciation error detection method according to claim 9, characterized in that: During the process of training the error detection model, the training output data of the error detection model is used to update the training of the acoustic model.
11. A speech scoring method, characterized in that: include: Determining, for the speech to be scored and the text to be read aloud corresponding to the speech to be scored, error detection information corresponding to each phoneme contained in the text to be read aloud, wherein the error detection information corresponding to each phoneme contained in the text to be read aloud is calculated based on the method according to any one of claims 1 to 10 above; The speech scoring information of the speech to be scored is determined based on the error detection information corresponding to each phoneme contained in the read text.
12. A pronunciation error detection device, characterized in that: include: A first determining module is used to determine a state sequence of the speech to be detected for error reading; The second determining module is used to determine N phoneme time boundary information corresponding to each phoneme contained in the reading text corresponding to the error-detected reading voice, wherein N is a positive integer; A generating module, configured to perform aggregation operations of different time scales on the state sequence based on N phoneme time boundary information corresponding to each phoneme contained in the read text, to generate phoneme aggregation data; A third determining module is configured to determine error detection information corresponding to each of the phonemes included in the read text based on the phoneme aggregation data; The aggregation operations at different time scales include short-time aggregation operations, and the phoneme aggregation data include short-time aggregation data; performing aggregation operations at different time scales on the state sequence based on N phoneme time boundary information corresponding to each phoneme contained in the read text to generate phoneme aggregation data includes: Dividing the state sequence based on N phoneme time boundary information corresponding to each phoneme included in the read-aloud text to obtain state sequence segments corresponding to each phoneme included in the read-aloud text; Based on the state sequence segments corresponding to the phonemes contained in the read text, performing at least one short-time aggregation operation to obtain aggregated data corresponding to the at least one short-time aggregation operation; The aggregated data corresponding to at least one of the short-term aggregation operations are fused to obtain the short-term aggregated data.
13. A speech scoring device, characterized in that: include: a first determining module configured to determine, for a speech to be scored and a reading text corresponding to the speech to be scored, error detection information corresponding to each phoneme contained in the reading text, wherein the error detection information corresponding to each phoneme contained in the reading text is calculated based on the method described in any one of claims 1 to 10 above, and the speech to be scored is obtained by a user reading the reading text; The second determining module is configured to determine the speech scoring information of the speech to be scored based on the error detection information corresponding to each phoneme contained in the read text.
14. An electronic device, characterized in that: include: processor; a memory for storing instructions executable by the processor, The processor is configured to execute the method according to any one of claims 1 to 11.
15. A computer-readable storage medium, characterized in that The storage medium stores a computer program, and the computer program is used to execute the method according to any one of claims 1 to 11.
Citation Information
Patent Citations
Pronunciation error detection method and device, electronic equipment and storage medium
CN111862960A