Pronunciation assessment methods, devices, equipment and storage media

By constructing a decoding network to analyze the optimal path of audio data and extracting segments of children's pronunciation for evaluation, the problem of indistinguishable pronunciation between parents and children is solved, thus achieving accuracy in pronunciation evaluation results and correction of children's pronunciation problems.

CN116631441BActive Publication Date: 2025-12-02GUANGZHOU XIBEISI INTELLIGENT TECHNOLOGY CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310519182.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-05-09
Publication Date
2025-12-02
Estimated Expiration
2043-05-09

AI Technical Summary

Technical Problem

In existing technologies, the pronunciation assessment results for parents and children cannot be accurately distinguished, resulting in inaccurate assessment results and affecting the correction of children's pronunciation problems.

Method used

By constructing a decoding network, analyzing the optimal path of audio data, and determining the input and output sequences, the system can extract segments of the child's pronunciation for evaluation, filtering out the influence of the parents' pronunciation.

Benefits of technology

It improves the accuracy of pronunciation assessment results, accurately reflects children's pronunciation problems, and helps children correct their own pronunciation problems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116631441B_ABST
    Figure CN116631441B_ABST
Patent Text Reader

Abstract

This application discloses a pronunciation evaluation method, apparatus, device, and storage medium. The method includes: acquiring audio data and corresponding pronunciation text; acquiring the phoneme sequence of the pronunciation text and a pre-constructed decoding network, wherein the last node in the decoding network is connected to the first node through an unconditional directed path; searching for the optimal path of the audio data based on the decoding network, and determining the input sequence and output sequence of the optimal path; determining the number of times the pronunciation text in the audio data is read aloud based on the output sequence; if the number of readings is greater than one, extracting audio segments of each reading of the pronunciation text from the audio data based on the input sequence and phoneme sequence; determining the target audio segment of the target object from multiple audio segments; evaluating the pronunciation of the target object based on the target audio segment; and outputting the evaluation result. Through the above technical means, the evaluation result can accurately reflect the child's pronunciation problems, which is beneficial for the child to correct their pronunciation problems.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular to a pronunciation evaluation method, apparatus, device, and storage medium. Background Technology

[0002] Pronunciation assessment technology is a subfield of computer-assisted language learning. It analyzes a user's audio recordings and outputs scores on indicators such as pronunciation accuracy, fluency, and completeness, allowing the user to correct their pronunciation problems based on these scores.

[0003] In existing technology, when parents guide their children to read aloud the text displayed on a learning device, the device collects the child's audio and evaluates the child's pronunciation based on the audio. When parents guide their children to read the text, they may first demonstrate the pronunciation themselves and then have the child repeat after them. At this time, the learning device collects audio containing both the parent's and child's pronunciations. The pronunciation evaluation based on this audio only provides a result for both the parent's and child's pronunciations, which is not accurate enough. Summary of the Invention

[0004] This application provides a pronunciation evaluation method, apparatus, device, and storage medium to solve the problem in the prior art where the comprehensive evaluation of children's and parents' pronunciation affects the accuracy of the evaluation results, thereby improving the accuracy of the evaluation results.

[0005] Firstly, this application provides a pronunciation assessment method, including:

[0006] Acquire audio data and the corresponding pronunciation text, acquire the phoneme sequence of the pronunciation text and a pre-constructed decoding network, wherein the last node in the decoding network is connected to the first node through an unconditional directed path;

[0007] The optimal path for the audio data is searched based on the decoding network, and the input and output sequences of the optimal path are determined.

[0008] The number of times the pronunciation text in the audio data is read is determined based on the output sequence. If the number of readings is greater than one, an audio segment of each reading of the pronunciation text is extracted from the audio data based on the input sequence and the phoneme sequence.

[0009] The target audio segment of the target object is determined from the multiple audio segments, the pronunciation of the target object is evaluated based on the target audio segment, and the evaluation result is output.

[0010] Secondly, this application provides a pronunciation evaluation device, including:

[0011] The data acquisition module is configured to acquire audio data and corresponding pronunciation text, acquire the phoneme sequence of the pronunciation text and a pre-constructed decoding network, wherein the last node in the decoding network is connected to the first node through an unconditional directed path;

[0012] The path search module is configured to search for the optimal path of the audio data frame by frame based on the decoding network, and to determine the input sequence and output sequence of the optimal path;

[0013] The segment extraction module is configured to determine the number of times the phonic text in the audio data is read aloud based on the output sequence, and if the number of readings is greater than one, extract an audio segment from the audio data for each reading of the phonic text based on the input sequence and the phoneme sequence;

[0014] The first evaluation module is configured to determine the target audio segment of the target object from the multiple audio segments, evaluate the pronunciation of the target object based on the target audio segment, and output the evaluation result.

[0015] Thirdly, this application provides a pronunciation evaluation device, including:

[0016] One or more processors; a memory storing one or more programs that, when executed by the one or more processors, cause the one or more processors to implement the pronunciation evaluation method as described in the first aspect.

[0017] Fourthly, this application provides a storage medium containing computer-executable instructions, which, when executed by a computer processor, are used to perform the pronunciation evaluation method as described in the first aspect.

[0018] In this application, audio data of a user reading aloud text is collected, and the phoneme sequence of the text and a decoding network pre-constructed based on the phoneme sequence are obtained. The optimal path of the audio data is searched based on the decoding network to obtain the input and output sequences of the optimal path. The input sequence consists of the target phonemes corresponding to each audio frame of the audio data in temporal order. The number of character sequences in the output sequence is the same as the number of times the user reads the text aloud. Based on the number of character sequences in the output sequence, the number of times the text was read aloud in the audio data can be determined. If the number of readings is greater than one, the audio data may contain audio of a non-target object reading the text. Therefore, based on the input sequence and phoneme sequence, audio segments of each reading of the text are extracted from the audio data. The target audio segment of the target object is determined from multiple audio segments, and the pronunciation of the target object is evaluated based on the target audio segment to obtain the evaluation result of the target object reading the text. The audio segments used for evaluation in this application only include audio segments of the target object reading the text aloud, filtering out the influence of non-target object audio on the evaluation results, improving the accuracy of the evaluation results, and solving the problem in the prior art that the comprehensive evaluation of children's pronunciation and parents' pronunciation affects the accuracy of the evaluation results. The evaluation results can accurately reflect the child's pronunciation problems and help the child correct their own pronunciation problems. Attached Figure Description

[0019] Figure 1 This is a flowchart of a pronunciation evaluation method provided in an embodiment of this application;

[0020] Figure 2 This is a schematic diagram of the decoding network provided in an embodiment of this application;

[0021] Figure 3 This is a flowchart of the input and output sequences for determining the optimal path provided in an embodiment of this application;

[0022] Figure 4 This is a schematic diagram of the first path of the second audio frame provided in an embodiment of this application;

[0023] Figure 5 This is a schematic diagram of the first path of the third audio frame provided in an embodiment of this application;

[0024] Figure 6 This is a flowchart of extracting audio segments of read-aloud text provided in an embodiment of this application;

[0025] Figure 7 This is a flowchart of determining a target audio segment based on voiceprint features provided in an embodiment of this application;

[0026] Figure 8 This is a flowchart of determining the target audio segment based on the binary search method provided in an embodiment of this application;

[0027] Figure 9 This is a schematic diagram of the structure of a pronunciation evaluation device provided in an embodiment of this application;

[0028] Figure 10 This is a schematic diagram of the structure of a pronunciation evaluation device provided in an embodiment of this application. Detailed Implementation

[0029] To make the objectives, technical solutions, and advantages of this application clearer, specific embodiments of this application will be described in further detail below with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are merely for explaining this application and not for limiting it. It should also be noted that, for ease of description, only the parts relevant to this application are shown in the drawings, not all of them. Before discussing exemplary embodiments in more detail, it should be mentioned that some exemplary embodiments are described as processes or methods depicted as flowcharts. Although the flowcharts describe operations (or steps) as sequential processes, many of these operations can be performed in parallel, concurrently, or simultaneously. Furthermore, the order of the operations can be rearranged. A process can be terminated when its operation is completed, but it may also have additional steps not included in the drawings. A process can correspond to a method, function, procedure, subroutine, subroutine, etc.

[0030] The terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such use of data can be interchanged where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class and the number of objects is not limited; for example, a first object can be one or more. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.

[0031] In a common existing implementation, when parents guide their children to read text displayed on a learning device, they first read the text aloud to demonstrate the correct pronunciation, and the child repeats after them. The learning device then captures audio containing both the parent's and child's pronunciations. For example, when the device displays the word "apple," the parent reads "apple" once, and the child repeats it. The audio captured by the device contains two pronunciations of "apple," but only the second pronunciation is the child's. Current learning devices perform a comprehensive evaluation of the entire audio clip, i.e., the two "apple" pronunciations. The output evaluation results address the pronunciation problems of both the parent and child, but this evaluation cannot accurately reflect the child's pronunciation problems and is not conducive to the child correcting their own pronunciation issues.

[0032] To address the aforementioned issues, this embodiment provides a pronunciation assessment method that extracts individual segments of a child's pronunciation from the collected audio data and assesses these segments to improve the accuracy of the assessment results.

[0033] The pronunciation evaluation method provided in this embodiment can be executed by a pronunciation evaluation device, which can be implemented by software and / or hardware. The pronunciation evaluation device can consist of two or more physical entities, or it can consist of a single physical entity. For example, the pronunciation evaluation device can be the learning device itself, or it can be the processor of the learning device.

[0034] The pronunciation assessment device is equipped with at least one type of operating system, including but not limited to Android, Linux, and Windows. The device can install at least one application based on the operating system; this application can be a built-in application of the operating system or an application downloaded from a third-party device or server. In this embodiment, the pronunciation assessment device has at least one application capable of executing pronunciation assessment methods.

[0035] For ease of understanding, this embodiment uses a learning device as the main body for performing the pronunciation evaluation method as an example.

[0036] Figure 1 A flowchart of a pronunciation evaluation method provided in an embodiment of this application is given. (Reference) Figure 1 The pronunciation assessment method specifically includes:

[0037] S110. Obtain audio data and corresponding pronunciation text, obtain the phoneme sequence of the pronunciation text and a pre-built decoding network, and connect the last node in the decoding network to the first node through an unconditional directed path.

[0038] This embodiment uses the word "apple" as an example for pronunciation. When the learning device displays "apple" on the screen, it begins recording. Recording ends when the parent or child clicks the pronunciation assessment control on the learning device. During recording, the parent and / or child read "apple" aloud, and the learning device collects the audio data of their reading.

[0039] The phoneme sequence of apple is ae-p-ah-l. A decoding network is constructed based on ae-p-ah-l, which is connected end to end. This decoding network is used to parse the number of times apple is pronounced in the audio data and the audio segment of the text apple in each pronunciation. Figure 2 This is a schematic diagram of the decoding network provided in an embodiment of this application. Figure 2 As shown, the decoding network consists of six nodes. Nodes 0 through 6 are the first, second, third, fourth, fifth, and last nodes of the decoding network, respectively. The last node is connected to the first node via an unconditional directed path, meaning the last and first nodes can be considered the same node. Except for the first and last nodes, each of the remaining nodes has two candidate input phonemes, two candidate directed paths, and two candidate output characters. Each candidate input phoneme corresponds to one candidate directed path and one candidate output character. The candidate input phoneme for the first node is sil, and its candidate directed path is a path pointing to the next node. The candidate directed paths corresponding to a node include paths pointing to the next node and directed paths pointing to itself. The candidate input phonemes corresponding to the paths pointing to the next node are determined according to the order of the phonemes in the phoneme sequence. For example, the candidate input phonemes corresponding to the paths pointing to the next node for nodes 1 through 4 are ae, p, ah, and l, respectively. The candidate input phonemes corresponding to the directed path pointing to itself are the candidate input phonemes corresponding to the path pointing to the next node from the previous node. For example, the candidate input phonemes corresponding to the directed paths pointing to themselves from nodes 1 to 5 are sil, ae, p, ah, and l, respectively. The candidate input phoneme corresponding to the path pointing to the next node from node 5 is the sil phoneme. The candidate output character of the first node is SIL, the candidate output characters of the third-to-last node include the character sequence of the pronunciation text and eps, the candidate output characters of the second-to-last node include SIL and eps, and the candidate output characters of other nodes are all eps. The candidate output character corresponding to the path pointing to the next node from the third-to-last node is the character sequence of the pronunciation text, and the candidate output character corresponding to the directed path pointing to itself from the third-to-last node is eps. The candidate output character corresponding to the path pointing to the next node from the second-to-last node is SIL, and the candidate output character corresponding to the directed path pointing to itself from the second-to-last node is eps. Figure 2In this example, sil represents a silence phoneme, SIL represents a silence character, and eps represents a meaningless character. In "ae:sil", the character on the left represents the node's input, and the character on the right represents the node's output. This embodiment analyzes the target phonemes corresponding to each audio frame of the audio data frame by frame using a decoding network and determines the number of times the spoken text in the audio data is read aloud.

[0040] S120. Search for the optimal path of audio data based on the decoding network, and determine the input sequence and output sequence of the optimal path.

[0041] The optimal path refers to the directed path formed by the target paths of each audio frame corresponding to the target node in the audio data. The target directed path is the directed path corresponding to the target input phoneme of the target node, and the target input phoneme is the target phoneme of the audio frame corresponding to the target node. After finding the optimal path, the input sequence can be obtained based on the target input phonemes of each target node on the optimal path, and the output sequence can be obtained based on the target output characters of each target node. The target output character is the output character corresponding to the target input phoneme of the node.

[0042] Figure 3 This is a flowchart illustrating the input and output sequences for determining the optimal path, as provided in an embodiment of this application. Figure 3 As shown, the steps for determining the input and output sequences of the optimal path specifically include S1201-S1204:

[0043] S1201. Determine the acoustic score of each audio frame in the audio data based on a pre-trained acoustic model. The acoustic score includes the predicted score of each preset phoneme.

[0044] The acoustic model is used to predict the probability of each preset phoneme corresponding to an audio frame. The neural network model is trained beforehand using open-source or commercial data to obtain the acoustic model. In this embodiment, the audio data is divided into multiple fixed-length audio frames, and the audio features of each audio frame are extracted. These audio features are input into the acoustic model to obtain the probability of each preset phoneme corresponding to the audio frame. The probability of each preset phoneme is the prediction score of the corresponding preset phoneme. A phoneme refers to the smallest unit of speech derived from the natural attributes of speech. Preset phonemes include multiple vowel phonemes, multiple consonant phonemes, and one silence phoneme. The preset phonemes encompass the phonemes in various phoneme sequences corresponding to different spoken texts.

[0045] S1202. Based on the prediction scores of each preset phoneme corresponding to each audio frame and the candidate input phonemes, candidate directed paths and candidate output characters of each node in the decoding network, search for the optimal path of the audio data and determine the target phoneme corresponding to each audio frame and the target output character of the target node corresponding to each audio frame.

[0046] The optimal path search process is as follows: The first node in the decoding network is identified as the target node corresponding to the first audio frame in the audio data. Based on the prediction scores of each preset phoneme corresponding to the first audio frame, the prediction scores of the candidate input phonemes of the target node corresponding to the first audio frame are determined. Each candidate directed path of the target node is taken as the first path of the first audio frame, and the prediction scores of the candidate input phonemes are taken as the path score of the first path to which the corresponding candidate directed path belongs. The node pointed to by the candidate directed path of the target node corresponding to the first audio frame is taken as the candidate node corresponding to the second audio frame. Based on the prediction scores of each preset phoneme corresponding to the second audio frame, the prediction scores of the candidate input phonemes of the corresponding candidate node are determined. The first path of the first audio frame to which the candidate directed path points to the candidate node belongs is combined with each candidate directed path of the candidate node, and the resulting combined path is taken as the first path of the second audio frame. The sum of the prediction scores of the candidate input phonemes corresponding to all candidate directed paths included in the first path is taken as the path score of the first path. Next, based on the nodes pointed to by the candidate directed paths of the candidate nodes in the previous audio frame, candidate nodes for the next audio frame are determined. Based on the prediction scores of each preset phoneme in the next audio frame, the prediction scores of each candidate input phoneme for the candidate node are determined. The first path of the previous audio frame to which the candidate directed path pointing to the candidate node of the next audio frame belongs is combined with each candidate directed path of the candidate node, and the resulting combined path is used as the first path of the next audio frame. The sum of the prediction scores of the candidate input phonemes corresponding to all candidate directed paths contained in the first path is used as the path score of the first path. When the candidate directed path of a candidate node in an audio frame points to the last node of the decoding network, the first node in the decoding network is determined as the candidate node for the next audio frame based on the unconditional directed path. After determining the first path of the last audio frame in the audio data, the first path with the highest path score is determined as the optimal path of the audio data.

[0047] refer to Figure 2Node 0 is designated as the target node corresponding to the first audio frame of the audio data. The candidate input phoneme for node 0 is 'sil'. The prediction score of the 'sil' phoneme is obtained from the prediction scores of each preset phoneme corresponding to the first audio frame. The directed path from node 0 to node 1 is taken as the first path of the first audio frame, and the prediction score of the 'sil' phoneme is taken as the path score of the first path of the first audio frame. The candidate directed path of the target node corresponding to the first audio frame points to node 1, so node 1 can be identified as a candidate node for the second audio frame. The candidate input phonemes of node 1 include the 'sil' and 'ae' phonemes. Based on the prediction scores of each preset phoneme corresponding to the second audio frame, the prediction scores of the 'sil' and 'ae' phonemes are determined. The directed path from node 1 to node 1 and the directed path from node 1 to node 2 are combined with the directed path from node 0 to node 1 to obtain the first path A and the first path B of the second audio frame. Figure 4 As shown, Figure 4 The first path A and the first path B of the second audio frame are shown. The path score of the first path A is obtained by adding the predicted score of the sil phoneme corresponding to the directed path from node 0 to node 1 in the first audio frame to the predicted score of the ae phoneme corresponding to the directed path from node 1 to node 2 in the second audio frame. Similarly, the path score of the first path B is obtained by adding the predicted score of the sil phoneme corresponding to the directed path from node 0 to node 1 in the first audio frame to the predicted score of the sil phoneme corresponding to the directed path from node 1 to node 1 in the second audio frame.

[0048] Two candidate directed paths from node 1 point to node 1 and node 2 respectively, thus identifying nodes 1 and 2 as candidate nodes for the third audio frame. The candidate input phonemes for node 2 include the p and ae phonemes, while those for node 1 include the sil and ae phonemes. Based on the prediction scores of each preset phoneme in the third audio frame, the prediction scores for the p, sil, and ae phonemes are determined. Since the candidate directed path pointing to node 2 (the directed path from node 1 to node 2) belongs to the first path A, the first path A is combined with the candidate directed paths of node 2 (including the directed path from node 2 to node 3 and the directed path from node 2 to node 2) to obtain the first paths for the second third audio frame. Since the candidate directed path pointing to node 1 (the directed path from node 1 to node 1) belongs to the first path B, the first path B is combined with the candidate directed paths of node 1 (including the directed path from node 1 to node 2 and the directed path from node 1 to node 1) to obtain the first paths for the third audio frame. Figure 5 As shown, Figure 5The first path C of the third audio frame is shown. The path score of the first path C is obtained by adding the path score of the first path A to the predicted score of the p-phoneme corresponding to the directed path from node 2 to node 3. The calculation principle for the path scores of the other first paths in the third audio frame is the same.

[0049] If a candidate node for a given audio frame includes node 5, and the candidate directed paths for node 5 include a directed path from node 5 to node 6, then based on the unconditional directed path from node 6 to node 0, node 0 is selected as the candidate node for the next audio frame. Based on the above steps, the first path and path score for each audio frame can be derived. After determining the first path and path score for the last audio frame, the first path with the highest path score is selected as the optimal path for the audio data from among the first paths corresponding to the last audio frame.

[0050] Following the temporal sequence of the audio frames, the optimal path sequentially passes through a candidate node of each audio frame. The nodes on the optimal path become the target nodes of the corresponding audio frames; for example, the third node on the optimal path is the target node of the third audio frame. The candidate directed paths included in the optimal path are then used as the target directed paths of the corresponding nodes; for example, the third candidate directed path on the optimal path is the target directed path of the third node. The candidate input phonemes corresponding to the target directed paths are then used as the target phonemes of the corresponding audio frames, and the candidate output characters corresponding to the target directed paths are then used as the target output characters of the corresponding nodes. For example, the candidate input phoneme corresponding to the third candidate directed path is used as the target phoneme of the third audio frame, and the candidate output character corresponding to the third candidate directed path is used as the target output character of the third node.

[0051] It should be noted that the later an audio frame appears in the audio data, the more first paths it corresponds to, and the greater the computational load. To reduce computational load, when determining the first path corresponding to an audio frame, the first path with the highest path score can be selected as the preferred first path for the audio frame. Then, based on a preset threshold, the remaining first paths are filtered out, and the first paths whose score difference from the preferred first path is less than the preset threshold are selected as backup first paths for the audio frame. Removing the first paths other than the backup and preferred paths from the first paths corresponding to the audio frame reduces the number of first paths for the audio frame and improves the efficiency of finding the optimal path.

[0052] S1203. Sort the target phonemes corresponding to each audio frame according to the order of each audio frame in the audio data to obtain the input sequence of the optimal path.

[0053] For example, the input sequence of the optimal path can be understood as a sequence of target input phonemes of each node on the optimal path, that is, a sequence of target phonemes of each audio frame according to the temporal order of the audio frames. The target phonemes corresponding to each audio frame are sorted according to the temporal order of the audio frames to obtain the input sequence of the optimal path.

[0054] S1204. Based on the order of each audio frame in the audio data, sort the target output characters of the target node corresponding to each audio frame to obtain the output sequence of the optimal path.

[0055] For example, the output sequence of the optimal path can be understood as a sequence of target output characters from each node on the optimal path. After determining the optimal path, the target output characters of the target nodes corresponding to each audio frame are sorted according to the temporal order of each audio frame to obtain the output sequence of the optimal path. Alternatively, the target output characters of each target node are sorted according to the order of the target nodes on the optimal path to obtain the output sequence of the optimal path.

[0056] S130. Determine the number of times the text is read aloud in the audio data based on the output sequence. If the number of readings is greater than one, extract audio segments of the text read aloud each time from the audio data based on the input sequence and the phoneme sequence.

[0057] In this embodiment, the number of times the spoken text in the audio data is read is determined based on the number of character sequences in the spoken text in the output sequence. Assume that the directed path from the first audio frame to the target node corresponding to the i-th audio frame can form a unidirectional path from node 0 to node 5, where the unidirectional path can include self-looping paths of nodes. Accordingly, concatenating the target output characters of the corresponding target nodes according to the order of the audio frames yields the character sequence "SIL apple". If the directed path from the (i+1)-th audio frame to the target node corresponding to the last audio frame can also form a unidirectional path from node 0 to node 5 or 6, then the output sequence of the optimal path is "SIL apple SIL SIL apple" or "SIL apple SIL SIL apple SIL". Since the output sequence "SIL apple SIL SIL apple" or "SIL apple SIL apple SIL apple SIL" contains two "apple", the number of times the spoken text is read is determined to be twice. Correspondingly, when the output sequence is "SIL apple" or "SIL apple SIL", the number of times the spoken text is read is determined to be once.

[0058] It should be noted that this embodiment describes the process of determining the output sequence and the number of readings in detail using one and two readings as examples. However, the number of readings can also be 0 or more than two. Regardless of the number of readings, the principle of determining the output sequence and the number of readings is the same as that of the above embodiment.

[0059] If the number of times the pronounced text is read aloud in the audio data is greater than one, it can be determined that the audio data includes at least two audio segments of the pronounced text. Based on the input sequence, multiple audio segments of the pronounced text are extracted from the audio data to determine which audio segment belongs to the segment from which the child read the pronounced text. In this embodiment, Figure 6 This is a flowchart illustrating the process of extracting audio segments from read-aloud text, as provided in an embodiment of this application. For example... Figure 6 As shown, the steps for extracting audio segments from the read-aloud text specifically include S1301-S1302:

[0060] S1301. Determine the target phoneme sequence in the input sequence based on the phoneme sequence. The order of different phonemes in the target phoneme sequence is the same as the order of different phonemes in the phoneme sequence. The adjacent audio frames of the target phoneme sequence in the input sequence are silent phonemes.

[0061] S1302. Based on the timestamp of the audio frame corresponding to the target phoneme sequence, extract the audio segment corresponding to the target phoneme sequence from the audio data.

[0062] The target phoneme sequence refers to the sequence of target phonemes composed of each audio frame in the audio segment of the read-aloud text, which is contained in the input sequence. When the read-aloud text is read, the user reads each phoneme in the phoneme sequence according to the phoneme order of the read-aloud text. Each phoneme lasts for a period of time, so the types of phonemes included in the target phoneme sequence are the same as those included in the phoneme sequence. The number of the same phoneme in the target phoneme sequence and the phoneme sequence can be different, and the order between different phonemes in the target phoneme sequence is the same as the order between different phonemes in the phoneme sequence. For example, if there is a sequence in the input sequence as "sil-sil-ae-ae-ae-ae-ae-pppp-ah-ah-ah-lll-sil-sil-sil", then "ae-ae-ae-ae-ae-ae-pppp-ah-ah-ah-lll" can be identified as the target phoneme sequence corresponding to the phoneme sequence. Assuming that “ae-ae-ae-ae-ae-pppp-ah-ah-ah-lll” corresponds to the third to the seventeenth audio frames respectively, the audio frame sequence from the third to the seventeenth audio frames is extracted from the audio data as the audio segment of the read-aloud text.

[0063] S140. Identify the target audio segment of the target object from multiple audio segments, evaluate the pronunciation of the target object based on the target audio segment, and output the evaluation result.

[0064] In this embodiment, the target object refers to a child learning pronunciation using a learning device, and the target audio segment refers to an audio segment recorded by the learning device when the child reads the pronunciation text aloud. When the audio data includes multiple audio segments of read-aloud texts, some audio segments may be recorded by the learning device when a parent reads the pronunciation text aloud. Evaluating these audio segments would affect the accuracy of the evaluation results. Therefore, this embodiment identifies the target audio segment from the multiple audio segments of read-aloud texts included in the audio data and evaluates the target audio segment to ensure the accuracy of the evaluation results.

[0065] This embodiment provides two methods for determining target audio segments: one is to filter target audio segments from multiple audio segments based on the voiceprint features of the target object, and the other is to filter target audio segments from multiple audio segments based on a binary classification method of children's pronunciation or adult pronunciation.

[0066] Figure 7 This is a flowchart illustrating the process of determining a target audio segment based on voiceprint features, as provided in an embodiment of this application. Figure 7 As shown, the steps for determining the target audio segment based on voiceprint features specifically include S1401-S1402:

[0067] S1401. Extract the voiceprint features of each audio segment and determine the similarity between each voiceprint feature and the reference voiceprint features of the target object.

[0068] S1402. If the similarity is greater than or equal to a preset similarity threshold, determine the corresponding audio segment as the target audio segment of the target object.

[0069] In this embodiment, the preset similarity threshold is the minimum similarity of the voiceprint features of two audio recordings belonging to the same object. For example, voiceprint recognition technology extracts corresponding voiceprint feature vectors from each audio segment, and the distance between each voiceprint feature vector and the reference voiceprint feature vector of the target object is determined as the similarity between the voiceprint features of the audio segment and the reference voiceprint features. When the similarity between the voiceprint features of the audio segment and the reference voiceprint features is greater than or equal to the preset similarity threshold, it can be determined that the voiceprint features and the reference voiceprint features belong to the same object, and thus the corresponding audio segment is identified as the target audio segment. When the similarity between the voiceprint features of the audio segment and the reference voiceprint features is less than the preset similarity threshold, it can be determined that the voiceprint features and the reference voiceprint features do not belong to the same object, and thus the corresponding audio segment is not identified as the target audio segment.

[0070] In this embodiment, the reference voiceprint features of the target object can be obtained by pre-collecting audio of the target object speaking and extracting voiceprint features from the audio as reference voiceprint features, or by extracting voiceprint features from historical audio data containing only one audio segment as reference voiceprint features. It can be understood that when the historical audio data contains only one audio segment of read-aloud text, this audio segment is the target audio segment recorded by the learning device when the target object reads the text aloud, and correspondingly, the voiceprint features extracted from this audio segment are the reference voiceprint features of the target object.

[0071] In another embodiment, Figure 8 This is a flowchart illustrating the determination of a target audio segment based on a binary search method, provided in an embodiment of this application. For example... Figure 8 As shown, the steps for determining the target audio segment based on the binary search method specifically include S1403-S1404:

[0072] S1403. Identify audio segments as either adult or child pronunciations using a pre-trained sound classification model.

[0073] S1404. When the audio segment is spoken by a child, the audio segment is identified as the target audio segment of the target object.

[0074] For example, a neural network model is pre-trained using various sample audio segments and corresponding pronunciation type labels to obtain a sound classification model for distinguishing whether an audio segment is spoken by an adult or a child. The pronunciation type labels include adult and child pronunciations. After inputting an audio segment into the sound classification model, the probability that the audio segment is spoken by an adult or a child is obtained. If the probability that the audio segment is spoken by an adult is high, the audio segment is determined to be spoken by an adult and is not the target audio segment; if the probability that the audio segment is spoken by a child is high, the audio segment is determined to be spoken by a child and is the target audio segment.

[0075] If multiple target audio segments can be identified from all the audio segments in the audio data, the pronunciation accuracy of each target audio segment is evaluated using GOP (goodness of pronunciation) technology. The number of target audio segments, the evaluation results of each target audio segment, and the segment number in the audio data are displayed on the student's device screen. If only one target audio segment can be identified from all the audio segments in the audio data, the pronunciation accuracy of the target audio segment is evaluated using GOP technology. The evaluation result is then displayed on the student's device screen along with the number of target audio segments and the evaluation result. If no target audio segment is identified from all the audio segments in the audio data, a pre-set prompt is displayed on the student's device screen to remind the user to read the pronunciation text again.

[0076] In one embodiment, when the number of readings is equal to one, an audio segment of the text to be read aloud is extracted from the audio data based on the input sequence and phoneme sequence as the target audio segment. The target audio segment is then evaluated, and the evaluation result is output. For example, when the number of readings of the text in the audio data is equal to one, the audio segment of the text to be read aloud in the audio data can be determined as the target audio segment recorded by the learning device when the target object reads the text aloud. After determining the target phoneme sequence corresponding to the audio segment in the input sequence based on the phoneme sequence, the target audio segment is extracted from the audio data based on the audio frame corresponding to the target phoneme sequence. The target audio segment is evaluated using GOP technology, and the evaluation result is displayed on the screen of the learning device so that the user can correct their pronunciation problems based on the evaluation result.

[0077] In summary, the pronunciation evaluation method provided in this application collects audio data of a user reading aloud text, obtains the phoneme sequence of the text and a decoding network pre-constructed based on the phoneme sequence, searches for the optimal path of the audio data based on the decoding network, and obtains the input sequence and output sequence of the optimal path. The input sequence is composed of the target phonemes corresponding to each audio frame of the audio data in temporal order, and the number of character sequences in the output sequence is the same as the number of times the user reads the text aloud. Based on the number of character sequences in the output sequence, the number of times the text was read aloud in the audio data can be determined. If the number of readings is greater than one, the audio data may contain audio of a non-target object reading the text. Therefore, based on the input sequence and phoneme sequence, audio segments of each reading of the text are extracted from the audio data. The target audio segment of the target object is determined from multiple audio segments, and the pronunciation of the target object is evaluated based on the target audio segment to obtain the evaluation result of the target object reading the text. The audio segments used for evaluation in this application only include audio segments of the target object reading the text aloud, filtering out the influence of non-target object audio on the evaluation results, improving the accuracy of the evaluation results, and solving the problem in the prior art that the comprehensive evaluation of children's pronunciation and parents' pronunciation affects the accuracy of the evaluation results. The evaluation results can accurately reflect the child's pronunciation problems and help the child correct their own pronunciation problems.

[0078] Based on the above embodiments, Figure 9 This is a schematic diagram of a pronunciation evaluation device provided in an embodiment of this application. (Reference) Figure 9 The pronunciation evaluation device provided in this embodiment specifically includes: a data acquisition module 21, a path search module 22, a segment extraction module 23, and a first evaluation module 24.

[0079] Among them, the data acquisition module 21 is configured to acquire audio data and corresponding pronunciation text, acquire the phoneme sequence of the pronunciation text and a pre-constructed decoding network, and the last node in the decoding network is connected to the first node through an unconditional directed path;

[0080] The path search module 22 is configured to search for the optimal path of audio data frame by frame based on the decoding network, and determine the input sequence and output sequence of the optimal path;

[0081] The segment extraction module 23 is configured to determine the number of times the phonic text in the audio data is read based on the output sequence, and if the number of readings is greater than one, extract audio segments of the phonic text for each reading from the audio data based on the input sequence and the phoneme sequence.

[0082] The first evaluation module 24 is configured to identify the target audio segment of the target object from multiple audio segments, evaluate the pronunciation of the target object based on the target audio segment, and output the evaluation result.

[0083] Based on the above embodiments, the path search module 22 includes: an acoustic score determination unit configured to determine the acoustic score of each audio frame in the audio data based on a pre-trained acoustic model, wherein the acoustic score includes the predicted score of each preset phoneme; a path search unit configured to search for the optimal path of the audio data based on the predicted scores of each preset phoneme corresponding to each audio frame and the candidate input phonemes, candidate directed paths, and candidate output characters of each node in the decoding network, thereby determining the target phoneme corresponding to each audio frame and the target output character of the node corresponding to each audio frame; an input sequence determination unit configured to sort the target phonemes corresponding to each audio frame according to the order of each audio frame in the audio data to obtain the input sequence of the optimal path; and an output sequence determination unit configured to sort the output characters of the node corresponding to each audio frame according to the order of each audio frame in the audio data to obtain the output sequence of the optimal path.

[0084] Based on the above embodiments, the path search unit includes: a first node determination subunit, configured to determine the first node in the decoding network as the node corresponding to the first audio frame in the audio data, and determine the prediction score of the candidate input phonemes of the node corresponding to the first audio frame according to the prediction scores of each preset phoneme corresponding to the first audio frame; and a phoneme determination subunit, configured to determine the candidate input phoneme with the highest preset score of the corresponding node as the target phoneme of the first audio frame, determine the candidate directed path corresponding to the candidate input phoneme with the highest preset score of the corresponding node as the optimal directed path of the first audio frame, and determine the candidate output character corresponding to the candidate input phoneme with the highest preset score of the corresponding node as the target output character of the node corresponding to the first audio frame. The second node determination subunit is configured to determine the node corresponding to the next audio frame based on the node pointed to by the optimal directed path of the first audio frame, and to determine the target phoneme, optimal directed path, and target output character of the next audio frame based on the prediction scores of each preset phoneme corresponding to the next audio frame, as well as the candidate input phonemes and candidate directed paths of the corresponding node. The third node determination subunit is configured to determine the node corresponding to the next audio frame as the first node in the decoding network based on the unconditional directed path when the optimal directed path of the audio frame points to the last node in the decoding network. The stop search subunit is configured to stop the search and obtain the optimal path of the audio data after determining the target phoneme of the last audio frame in the audio data.

[0085] Based on the above embodiments, the segment extraction module 23 includes: a reading count determination unit, configured to determine the reading count of the audio text in the audio data based on the number of character sequences of the audio text in the output sequence.

[0086] Based on the above embodiments, the segment extraction module 23 includes: a target phoneme sequence determination unit, configured to determine the target phoneme sequence corresponding to the input sequence based on the phoneme sequence, wherein the order between different phonemes in the target phoneme sequence is the same as the order between different phonemes in the phoneme sequence, and the adjacent audio frames of the target phoneme sequence in the input sequence are silent phonemes; and an audio segment extraction unit, configured to extract the audio segment corresponding to the target phoneme sequence from the audio data based on the timestamp of the audio frame corresponding to the target phoneme sequence.

[0087] Based on the above embodiments, the first evaluation module 24 includes: a similarity determination unit, configured to extract the voiceprint features of each audio segment and determine the similarity between each voiceprint feature and the reference voiceprint features of the target object; and a first segment determination unit, configured to determine the corresponding audio segment as the target audio segment of the target object when the similarity is greater than or equal to a preset similarity threshold.

[0088] Based on the above embodiments, the first evaluation module 24 includes: a pronunciation type determination unit, configured to identify whether an audio segment is an adult pronunciation or a child pronunciation through a pre-trained sound classification model; and a second segment determination unit, configured to determine the audio segment as the target audio segment of the target object when the audio segment is a child pronunciation.

[0089] Based on the above embodiments, the pronunciation evaluation device further includes: a second evaluation module, configured to, when the number of readings is equal to one, extract an audio segment of the read-out text from the audio data as a target audio segment based on the input sequence and the phoneme sequence, evaluate the target audio segment, and output the evaluation result.

[0090] The pronunciation evaluation device provided in this application, as described above, acquires audio data of a user reading aloud text, obtains the phoneme sequence of the text and a decoding network pre-constructed based on the phoneme sequence, searches for the optimal path of the audio data based on the decoding network, and obtains the input sequence and output sequence of the optimal path. The input sequence is composed of the target phonemes corresponding to each audio frame of the audio data in temporal order, and the number of character sequences in the output sequence is the same as the number of times the user reads the text aloud. Based on the number of character sequences in the output sequence, the number of times the text was read aloud in the audio data can be determined. If the number of readings is greater than one, the audio data may contain audio of a non-target object reading the text. Therefore, based on the input sequence and phoneme sequence, audio segments of each reading of the text are extracted from the audio data. The target audio segment of the target object is determined from multiple audio segments, and the pronunciation of the target object is evaluated based on the target audio segment to obtain the evaluation result of the target object reading the text. The audio segments used for evaluation in this application only include audio segments of the target object reading the text aloud, filtering out the influence of non-target object audio on the evaluation results, improving the accuracy of the evaluation results, and solving the problem in the prior art that the comprehensive evaluation of children's pronunciation and parents' pronunciation affects the accuracy of the evaluation results. The evaluation results can accurately reflect the child's pronunciation problems and help the child correct their own pronunciation problems.

[0091] The pronunciation evaluation device provided in this application embodiment can be used to execute the pronunciation evaluation method provided in the above embodiment, and has corresponding functions and beneficial effects.

[0092] Figure 10 This is a schematic diagram of the structure of a pronunciation evaluation device provided in an embodiment of this application, with reference to... Figure 10 The pronunciation evaluation device includes a processor 31, a memory 32, a communication device 33, an input device 34, and an output device 35. The number of processors 31 and the number of memories 32 in the pronunciation evaluation device can be one or more. The processor 31, memory 32, communication device 33, input device 34, and output device 35 of the pronunciation evaluation device can be connected via a bus or other means.

[0093] The memory 32, as a computer-readable storage medium, can be used to store software programs, computer-executable programs, and modules, such as program instructions / modules corresponding to the pronunciation evaluation method in any embodiment of this application (e.g., the data acquisition module 21, path search module 22, segment extraction module 23, and first evaluation module 24 in the pronunciation evaluation device). The memory 32 may primarily include a program storage area and a data storage area. The program storage area may store the operating system and at least one application program required for a function; the data storage area may store data created based on the use of the device, etc. Furthermore, the memory 32 may include high-speed random access memory and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other non-volatile solid-state storage device. In some instances, the memory may further include memory remotely located relative to the processor, and these remote memories can be connected to the device via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.

[0094] The communication device 33 is used for data transmission.

[0095] The processor 31 executes various functional applications and data processing of the device by running software programs, instructions and modules stored in the memory 32, thereby realizing the above-mentioned pronunciation evaluation method.

[0096] Input device 34 can be used to receive input digital or character information, and to generate key signal inputs related to user settings and function control of the device. Output device 35 may include display devices such as a display screen.

[0097] The pronunciation evaluation device provided above can be used to perform the pronunciation evaluation method provided in the above embodiments, and has corresponding functions and beneficial effects.

[0098] This application embodiment also provides a storage medium containing computer-executable instructions. When executed by a computer processor, the computer-executable instructions are used to perform a pronunciation evaluation method. The pronunciation evaluation method includes: acquiring audio data and corresponding pronunciation text; acquiring the phoneme sequence of the pronunciation text and a pre-constructed decoding network, wherein the last node in the decoding network is connected to the first node through an unconditional directed path; searching for the optimal path of the audio data based on the decoding network, and determining the input sequence and output sequence of the optimal path; determining the number of times the pronunciation text in the audio data is read aloud based on the output sequence; if the number of readings is greater than one, extracting audio segments of each reading of the pronunciation text from the audio data based on the input sequence and the phoneme sequence; determining the target audio segment of the target object from multiple audio segments; evaluating the pronunciation of the target object based on the target audio segment; and outputting the evaluation result.

[0099] Storage medium – any type of memory device or storage device. The term “storage medium” is intended to include: mounting media, such as CD-ROM, floppy disk, or magnetic tape devices; computer system memory or random access memory, such as DRAM, DDR RAM, SRAM, EDO RAM, Rambus RAM, etc.; non-volatile memory, such as flash memory, magnetic media (e.g., hard disk or optical storage); registers or other similar types of memory elements, etc. Storage medium may also include other types of memory or combinations thereof. Furthermore, storage medium may reside in a first computer system in which the program is executed, or it may reside in a different second computer system connected to the first computer system via a network (such as the Internet). The second computer system can provide program instructions to the first computer for execution. The term “storage medium” can include two or more storage media residing in different locations (e.g., in different computer systems connected via a network). Storage medium may store program instructions (e.g., specifically implemented as a computer program) executable by one or more processors.

[0100] Of course, the computer-executable instructions provided in the embodiments of this application are not limited to the pronunciation evaluation method described above, but can also perform related operations in the pronunciation evaluation method provided in any embodiment of this application.

[0101] The pronunciation evaluation device, storage medium, and pronunciation evaluation equipment provided in the above embodiments can execute the pronunciation evaluation method provided in any embodiment of this application. For technical details not described in detail in the above embodiments, please refer to the pronunciation evaluation method provided in any embodiment of this application.

[0102] The above description is merely a preferred embodiment and the technical principles employed in this application. This application is not limited to the specific embodiments described herein, and various obvious changes, readjustments, and substitutions that can be made by those skilled in the art will not depart from the scope of protection of this application. Therefore, although this application has been described in detail through the above embodiments, this application is not limited to the above embodiments, and may include many other equivalent embodiments without departing from the concept of this application. The scope of this application is determined by the scope of the claims.

Claims

1. A pronunciation assessment method, characterized in that, include: Acquire audio data and the corresponding pronunciation text, acquire the phoneme sequence of the pronunciation text and a pre-constructed decoding network, wherein the last node in the decoding network is connected to the first node through an unconditional directed path; The optimal path for the audio data is searched based on the decoding network, and the input and output sequences of the optimal path are determined. The number of times the pronunciation text in the audio data is read is determined based on the output sequence. If the number of readings is greater than one, an audio segment of each reading of the pronunciation text is extracted from the audio data based on the input sequence and the phoneme sequence. The target audio segment of the target object is determined from the multiple audio segments, the pronunciation of the target object is evaluated based on the target audio segment, and the evaluation result is output.

2. The pronunciation evaluation method according to claim 1, characterized in that, The step of searching for the optimal path of the audio data based on the decoding network and determining the input and output sequences of the optimal path includes: The acoustic score of each audio frame in the audio data is determined based on a pre-trained acoustic model, and the acoustic score includes the predicted score of each preset phoneme. Based on the prediction scores of each preset phoneme corresponding to each audio frame and the candidate input phonemes, candidate directed paths and candidate output characters of each node in the decoding network, the optimal path of the audio data is searched to determine the target phoneme corresponding to each audio frame and the target output character of the target node corresponding to each audio frame. According to the order of each audio frame in the audio data, the target phonemes corresponding to each audio frame are sorted to obtain the input sequence of the optimal path; Based on the order of each audio frame in the audio data, the target output characters of the target node corresponding to each audio frame are sorted to obtain the output sequence of the optimal path.

3. The pronunciation evaluation method according to claim 2, characterized in that, The step of determining the number of times the spoken text in the audio data is read aloud based on the output sequence includes: The number of times the spoken text in the audio data is read aloud is determined based on the number of character sequences in the spoken text in the output sequence.

4. The pronunciation evaluation method according to claim 2, characterized in that, The step of extracting audio segments from the audio data for each reading of the pronounced text based on the input sequence and the phoneme sequence includes: Based on the phoneme sequence, a target phoneme sequence is determined in the input sequence. The order between different phonemes in the target phoneme sequence is the same as the order between different phonemes in the phoneme sequence. The adjacent audio frames before and after the target phoneme sequence in the input sequence are silence phonemes. Based on the timestamp of the audio frame corresponding to the target phoneme sequence, an audio segment corresponding to the target phoneme sequence is extracted from the audio data.

5. The pronunciation evaluation method according to claim 1, characterized in that, The step of determining the target audio segment of the target object from the plurality of audio segments includes: Extract the voiceprint features of each audio segment and determine the similarity between each voiceprint feature and the reference voiceprint features of the target object; If the similarity is greater than or equal to a preset similarity threshold, the corresponding audio segment is determined to be the target audio segment of the target object.

6. The pronunciation evaluation method according to claim 1, characterized in that, The step of determining the target audio segment of the target object from the plurality of audio segments includes: The audio segment is identified as either an adult's or a child's voice using a pre-trained sound classification model. If the audio segment is spoken by a child, the audio segment is identified as the target audio segment of the target object.

7. The pronunciation evaluation method according to claim 1, characterized in that, After determining the number of times the pronounced text in the audio data is read aloud based on the output sequence, the method further includes: When the number of readings is equal to one, based on the input sequence and the phoneme sequence, an audio segment of reading the pronounced text is extracted from the audio data as a target audio segment, the target audio segment is evaluated, and the evaluation result is output.

8. A pronunciation evaluation device, characterized in that, include: The data acquisition module is configured to acquire audio data and corresponding pronunciation text, acquire the phoneme sequence of the pronunciation text and a pre-constructed decoding network, wherein the last node in the decoding network is connected to the first node through an unconditional directed path; The path search module is configured to search for the optimal path of the audio data frame by frame based on the decoding network, and to determine the input sequence and output sequence of the optimal path; The segment extraction module is configured to determine the number of times the phonic text in the audio data is read aloud based on the output sequence, and if the number of readings is greater than one, extract an audio segment from the audio data for each reading of the phonic text based on the input sequence and the phoneme sequence; The first evaluation module is configured to determine the target audio segment of the target object from the multiple audio segments, evaluate the pronunciation of the target object based on the target audio segment, and output the evaluation result.

9. A pronunciation evaluation device, characterized in that, include: One or more processors; A memory that stores one or more programs, which, when executed by one or more processors, cause the one or more processors to implement the pronunciation evaluation method as described in any one of claims 1-7.

10. A storage medium containing computer-executable instructions, characterized in that, The computer-executable instructions, when executed by a computer processor, are used to perform the pronunciation evaluation method as described in any one of claims 1-7.

Citation Information

Patent Citations

  • Spoken language pronunciation evaluation method and device, equipment and storage medium

    CN113571094A

  • Method and device for evaluating chapter recitation quality, electronic equipment and storage medium

    CN115331662A