Monosyllabic reading determination device, monosyllabic reading determination system, monosyllabic reading determination method, and program
The monosyllabic reading determination device provides an objective and flexible evaluation of kana character pronunciation by considering output probabilities, pronunciation and shape similarities, and speaker characteristics, addressing inaccuracies in conventional methods and supporting dyslexia learning.
Patent Information
- Application Number
- JP2025111506
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2025-07-01
- Publication Date
- 2025-08-29
- Estimated Expiration
- 2045-07-01
AI Technical Summary
Conventional reading evaluation methods lack accuracy in assessing the pronunciation of monosyllabic sounds, particularly for individuals with dyslexia, due to unclear correspondence between speech recognition results and reference characters, and do not account for voice fluctuations or character skipping during reading.
A monosyllabic reading determination device that includes a display control unit to show reference kana characters, a speech recognition unit to identify monosyllabic sounds with output probabilities, and a determination unit to evaluate correct reading based on probability, pronunciation similarity, shape similarity, and speaker characteristics.
Enables objective and flexible evaluation of reading aloud, accurately determining correct or incorrect pronunciation of kana characters, including cases of skipping or misreading, and supporting learning for dyslexia.
Smart Images

Figure 0007731182000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to a monosyllabic reading determination device, a monosyllabic reading determination method, and a program for determining whether a speaker has correctly read a monosyllabic. [Background technology]
[0002] Conventionally, cognitive processing related to reading behavior has been explained by the "dual-route model of reading." This model assumes that reading involves two processing routes: a non-lexical route (phonological route) in which characters are converted into sounds one by one, and a lexical route (visual-lexical route) in which familiar words are visually recognized and their phonological meanings are acquired all at once through semantic processing.
[0003] In the early stages of reading, non-lexical pathways that link individual letters with sounds are primarily used. At this stage, speakers must read letters sequentially, which makes reading slow and prone to errors. However, as readers become more proficient, they can read familiar words and letters using the lexical pathway, allowing for faster and more accurate reading.
[0004] It has been pointed out that people with dyslexia (developmental reading disorder) have delayed development of the non-lexical pathway, making it difficult to accurately decode each letter. In response to this, learning support technologies using reading aloud training and audio feedback have been proposed to promote the automation of decoding processes in the non-lexical pathway.
[0005] For example, a technology is known in which a speaker reads aloud displayed kana characters, recognizes the speech, and evaluates the accuracy of the reading based on the degree of agreement between the speech recognition result and a reference sound. Furthermore, technology is being considered to prevent misidentification by analyzing the accuracy of speech recognition for each syllable and misreading trends, taking into account the pronunciation tendencies unique to people with underdeveloped palates and the confusion with visually similar kana characters.
[0006] Furthermore, as related prior art, Patent Document 1 proposes a speech recognition device that can accurately recognize monosyllabic speech. This technology separates input speech into consonant and vowel parts, performs spectral analysis on each speech section, and compares the generated patterns with a dictionary to recognize the consonants and vowels separately, ultimately integrating them into monosyllabic sounds. While attempts are made to improve recognition accuracy through methods such as majority voting and vowel correction, these are intended to output identification results and are not intended for use in learning support, such as determining whether a reading is correct or not, or detecting skipped parts.
[0007] However, in conventional technology, there is no fully established method for accurately assessing the accuracy of pronunciation at the monosyllable level from multiple perspectives, such as the output probability included in the speech recognition results, pronunciation similarity, visual shape similarity, and speech characteristics at different developmental stages. [Prior art documents] [Patent documents]
[0008] [Patent Document 1] Japanese Patent Application Publication No. 62-269198 Summary of the Invention [Problem to be solved by the invention]
[0009] In recent years, technology to assess reading aloud ability has become increasingly important in fields such as learning support and diagnosing language development. In particular, the ability to appropriately assess the accuracy of reading aloud for individual kana characters has attracted attention as it can be useful for detecting early literacy difficulties and measuring the effectiveness of reading aloud training.
[0010] However, in conventional reading evaluation methods, the correspondence between the results of speech recognition and the reference characters is often unclear, and voice fluctuations due to the tendency for misreading or individual differences in speech are often not properly taken into account, making it difficult to achieve accurate and reliable evaluation.
[0011] Furthermore, when reading multiple kana characters aloud in succession, some characters may be skipped or the order may be changed, but there is no established technology that can reliably determine whether the reading is correct even in such cases.
[0012] The present invention has been made in consideration of the above points, and an object of the present invention is to provide a reading aloud determination technique that enables a more objective and flexible evaluation of characters that are read aloud. [Means for solving the problem]
[0013] According to the present invention, a display control means for displaying on a display device reference kana characters indicating reference monosyllabic sounds; a speech recognition means for recognizing the speech of a speaker who reads the reference kana characters aloud and outputting identification information for identifying each of a plurality of monosyllabic sounds obtained by the recognition together with the output probability of each sound; a determination means for determining whether the speaker is reading the reference kana characters correctly based on identification information for identifying the reference monosyllabic sounds, and the identification information and output probabilities for the plurality of types of monosyllabic sounds; Equipped with A monosyllabic reading determination device is provided. [Effects of the Invention]
[0014] According to the present invention, a reading aloud determination technique is provided that enables more objective and flexible evaluation of characters that are read aloud. [Brief explanation of the drawings]
[0015] [Figure 1] 1 is a functional block diagram showing the configuration of a monosyllabic reading determination device according to a first embodiment of the present invention. [Figure 2] FIG. 1 is a diagram for explaining a monosyllabic reading determination method according to a first embodiment of the present invention. [Figure 3] FIG. 4 is another diagram for explaining the monosyllabic reading determination method according to the first embodiment of the present invention. [Figure 4] FIG. 2 is a diagram showing a table used in the monosyllabic reading determination device according to the first embodiment of the present invention. [Figure 5] 4 is a flowchart illustrating a first monosyllabic reading determination method executed by the monosyllabic reading determination device according to the first embodiment of the present invention. [Figure 6] 10 is a first part of a flowchart illustrating a second monosyllabic reading determination method executed by the monosyllabic reading determination device according to the first embodiment of the present invention. [Figure 7] 10 is a second part of a flowchart illustrating a second monosyllabic reading determination method executed by the monosyllabic reading determination device according to the first embodiment of the present invention. [Figure 8] 10 is a first part of a flowchart illustrating a third monosyllabic reading determination method executed by the monosyllabic reading determination device according to the first embodiment of the present invention. [Figure 9] 10 is a second part of a flowchart illustrating a third monosyllabic reading determination method executed by the monosyllabic reading determination device according to the first embodiment of the present invention. [Figure 10] 10 is a first part of a flowchart illustrating a fourth monosyllabic reading determination method executed by the monosyllabic reading determination device according to the first embodiment of the present invention. [Figure 11] 10 is a second part of a flowchart illustrating a fourth monosyllabic reading determination method executed by the monosyllabic reading determination device according to the first embodiment of the present invention. [Figure 12] FIG. 10 is a diagram showing an example of a screen displayed by the monosyllabic reading determination device according to the second embodiment of the present invention. [Figure 13] 10 is a flowchart illustrating a monosyllabic reading determination method executed by a monosyllabic reading determination system according to a second embodiment of the present invention. [Figure 14] FIG. 1 is a functional block diagram showing a speech recognition unit and a model generation device for generating a trained model for recognition used in the speech recognition unit, which are included in the monosyllabic reading determination device according to the first or second embodiment of the present invention. [Figure 15]FIG. 10 is a functional block diagram showing a speech recognition unit and a model generation device for generating a trained model for recognition used in the speech recognition unit, both of which are included in a monosyllabic reading determination device according to a third embodiment of the present invention. [Figure 16] FIG. 10 is a functional block diagram showing a speech recognition unit and a model generation device for generating a trained model for recognition used in the speech recognition unit, both of which are included in a monosyllabic reading determination device according to a fourth embodiment of the present invention. [Figure 17] FIG. 11 is a functional block diagram showing a speech recognition unit and a model generation device for generating a trained model for recognition used in the speech recognition unit, both of which are included in a monosyllabic reading determination device according to a fifth embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0016] DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS Hereinafter, embodiments of the present invention will be described in detail with reference to the drawings.
[0017] [First embodiment] As shown in FIG. 1, the monosyllabic reading determination device according to the first embodiment of the present invention includes, in addition to a processor 101, a memory 102, a storage device 103, a display control unit 104, a display device 105, a voice input unit 106, a voice recognition unit 107, a determination unit 108, an input / output unit 109, an input device 110, an output device 111, and a bus 112 connecting these devices.
[0018] The processor 101 controls the entire monosyllabic reading determination device in accordance with a program stored in the storage device 103. The display device 105 is provided with a UI for presenting reference kana characters (monosyllabic characters) and prompts the speaker to read aloud.
[0019] The speech input unit 106 digitizes the speech of the speaker 121 acquired by a microphone or the like into speech data and sends it to the speech recognition unit 107. The speech recognition unit 107 classifies the input speech into multiple types of monosyllables based on the trained model for recognition 1414 (see FIG. 14 ) and calculates the output probability of each type.
[0020] The determination unit 108 determines whether the speaker is reading the displayed reference kana characters correctly using the methods shown in Figures 5 to 11. The determination unit 108 also determines whether the reading is correct from multiple angles, taking into account skipping (see Figure 13), misreading based on shape similarity (see Figures 4 and 5), and misreading due to pronunciation similarity or that specific to people with underdeveloped palates (see Figures 6 and 7). The determination unit 108 also reflects the possibility of ambiguous pronunciation (see Figure 9) in the record.
[0021] 1, the speech recognition unit 107 and the determination unit 108 are depicted as functional blocks separate from the processor 101. However, this is not limited to this, and the processor 101 may load and execute a program that gives the processor 101 the functions of the speech recognition unit 107 and the determination unit 108, as indicated by the dashed lines in FIG.
[0022] Figures 2 and 3 show the kana characters representing multiple monosyllabic sounds obtained by recognition of the kana character "a" and their output probabilities, as well as the kana characters representing monosyllabic sounds obtained by judgment. Also shown are examples of correct / incorrect judgments based on shape similarity and probability distribution. Figure 4 shows classifications of shape and sound similarity for each reference kana character, which are used for judgment.
[0023] As shown in Figure 13, even if the speaker skips some of the kana characters displayed side by side on one screen, the system repeats evaluation for all combinations and outputs the most accurate skipped combination and judgment result. This allows for flexible response even in situations where skipping occurs.
[0024] As shown in FIG. 14, a speech recognition calculation unit 1411 included in the speech recognition unit 107 uses an acoustic model 1413 and a trained model for recognition 1414 to probabilistically identify multiple types of monosyllabic sounds from speech.
[0025] With these configurations, this embodiment can accurately determine whether a speaker has correctly read the displayed kana characters by comprehensively taking into consideration output probability, similarity in pronunciation, similarity in shape, speaker characteristics, etc., and can provide an effective evaluation means for reading aloud training and support for dyslexia.
[0026] FIG. 2 is a diagram showing an example of a speech recognition result for a monosyllable that is read aloud, and a reading aloud determination process based on the result.
[0027] 2, the reference kana character "a" is displayed on the display device 105, and the voice of the speaker reading it aloud is input to the voice recognition unit 107. The voice recognition unit 107 inputs a plurality of monosyllable candidates such as "a," "o," and "wa" together with their output probabilities to the determination unit 108.
[0028] In the example shown by reference numeral 202, "a" has the highest output probability, followed by "wa" and "wa" in that order. Here, "wa" and "wa" are similar in pronunciation to "a." Since the monosyllabic sound with the highest output probability is the same as the reference monosyllabic sound (also referred to as the "reference sound"), the determination unit 108 determines that the speaker is reading the reference kana character correctly.
[0029] In the example shown by reference numeral 203, "wa" has the highest output probability, followed by "a" and "wa" in that order. The determination unit 108 determines that the speaker is reading the standard kana character correctly because the monosyllabic sound with the highest output probability is a sound that is pronunciation-similar to the standard sound.
[0030] In the example shown by reference numeral 204, "o" has the highest output probability, followed by "ko" and "so" in that order. Here, "o" is similar in shape to "a." Furthermore, "ko" and "so" are similar in pronunciation to "o." The determination unit 108 determines that the speaker is misreading the reference kana character because the monosyllabic sound with the highest output probability is identical to the monosyllabic sound represented by "o," a kana character whose shape is similar to the reference kana character. In particular, the determination unit 108 determines that the speaker is mistaking the reference kana character for a kana character whose shape is similar to the reference kana character.
[0031] In the example shown by reference numeral 205, "ko" has the highest output probability, followed by "o" and "so" in that order. The determination unit 108 determines that the speaker is misreading the reference kana character because the monosyllabic sound with the highest output probability is a sound similar in pronunciation to the monosyllabic sound represented by "o," a kana character whose shape is similar to the reference kana character. In particular, the determination unit 108 determines that the speaker is mistaking the reference kana character for a kana character whose shape is similar to the reference kana character.
[0032] As such, Figures 2 and 3 specifically show the probabilistic output results in speech recognition of monosyllabic sounds and the factors considered in determining whether a sound is correct or not based on pronunciation and visual similarity, and are an example that will help understand the determination processing logic of the present invention.
[0033] FIG. 4 is a diagram showing an example of classification of other kana characters that are phonetically and visually similar to a reference kana character.
[0034] In this diagram, kana characters are classified into four categories based on the base kana characters. Each category is as follows: [Group A]: Kana characters that sound similar and have similar shapes (not applicable) [Group B]: Kana characters that sound similar but have different shapes [Group C]: Kana characters that are similar in shape but not in sound [Group D]: Kana characters that are similar in both sound and shape Such a table is used to implement the determination method described with reference to Figures 2 and 3. That is, once a reference kana character is selected, kana characters belonging to Group B and kana characters belonging to Group C are used for determination of that reference kana character.
[0035] Note that, as identification information for identifying a monosyllabic sound that can be represented by one kana character, such as "a," the character code corresponding to that one kana character can be used. Also, as identification information for identifying a monosyllabic sound that can be represented by two kana characters, such as "kya," the combined character codes corresponding to those two kana characters can be used.
[0036] When comparing a first monosyllabic sound with a second monosyllabic sound, the identification information for identifying the first monosyllabic sound is compared with the identification information for identifying the second monosyllabic sound.
[0037] FIG. 5 shows a flowchart for explaining the first determination method performed by the determination unit .
[0038] In S501, the determination unit 108 starts a timer (not shown).
[0039] In S502, the determination unit 108 displays a reference kana character indicating a reference monosyllabic sound. Specifically, in response to an instruction from the determination unit 108, the display control unit 104 causes the display device 105 to display a new reference kana character.
[0040] If the speech input unit 106 inputs speech from the speaker (YES in S503) before the timer times out (NO in S504), the judgment unit 108 acquires, in S505, multiple types of monosyllabic sounds obtained by speech recognition by the speech recognition unit 107, along with their output probabilities.
[0041] If the timer times out (YES in S504) before the speech input unit 106 inputs the speech of the speaker (NO in S503), the determination unit 108 determines in S516 that Displayed kana characters (standard kana characters) -Determined reading: NULL · Verification result: Skip ·Similar characters: NULL are recorded in a set in the storage device 103.
[0042] In S506, the determination unit 108 determines whether the monosyllabic sound with the highest output probability is the same as a reference monosyllabic sound (reference sound).
[0043] If the determination result in S506 is affirmative, the determination unit 108 advances the process to S510, and if negative, the determination unit 108 advances the process to S507.
[0044] In S507, the determination unit 108 determines whether the monosyllabic sound with the highest output probability is a sound similar in pronunciation to the reference monosyllabic sound.
[0045] If the determination result of S507 is affirmative, the determination unit 108 advances the process to S510, and if negative, the determination unit 108 advances the process to S508.
[0046] In S508, the determination unit 108 determines whether the monosyllabic sound with the highest output probability is the same as the monosyllabic sound represented by any kana character that is similar in shape to the reference kana character.
[0047] If the determination result of S508 is affirmative, the determination unit 108 advances the process to S512, and if negative, the determination unit 108 advances the process to S509.
[0048] In S509, the determination unit 108 determines whether the monosyllabic sound with the highest output probability is a sound whose pronunciation is similar to that of a monosyllabic sound represented by any kana character whose shape is similar to that of the reference kana character.
[0049] If the determination result of S509 is affirmative, the determination unit 108 advances the process to S512, and if negative, the determination unit 108 advances the process to S514.
[0050] In S510, the determination unit 108 determines that the speaker is reading the reference kana characters correctly.
[0051] In S511, the determination unit 108 Displayed kana characters (standard kana characters) - Determined "reading": Displayed kana characters ·Judgment result: Positive ·Similar characters: NULL are recorded in a set in the storage device 103.
[0052] In S512, the determination unit 108 determines that the speaker is reading the reference kana characters incorrectly. In particular, the determination unit 108 determines that the speaker is reading the kana characters incorrectly due to similarities in their shapes.
[0053] In S513, the determination unit 108 Displayed kana characters (standard kana characters) - Determined "reading": The kana character whose shape is similar to the displayed kana character ·Judgment result: False ·Similar characters:YES are recorded in a set in the storage device 103.
[0054] In S514, the determination unit 108 determines that the speaker has read the reference kana character incorrectly. In particular, the determination unit 108 determines that the speaker has read the reference kana character incorrectly due to reasons other than the similarity of the kana character's shape. Note that the determination unit 108 does not need to determine the sound of the kana character that the speaker has uttered.
[0055] In S515, the determination unit 108 Displayed kana characters (standard kana characters) -Determined reading: NULL ·Judgment result: False ·Similar characters: NO are recorded in a set in the storage device 103.
[0056] 5, the determination unit 108 classifies and evaluates the accuracy of reading for each syllable in multiple stages, taking into account the output probability of the speech recognition result and the similarity of the shape and pronunciation of the kana characters. This enables highly practical reading determination for grasping misreading tendencies and for supporting learning.
[0057] 6 and 7 show flowcharts for explaining the second determination method by the determination unit 108. The second method is the first method to which a part corresponding to pronunciation by a speaker with an underdeveloped palate is added.
[0058] Among the steps included in the second method, the steps that are the same as the steps included in the first method are given the same reference numerals and redundant explanations will be omitted.
[0059] In S609, the determination unit 108 determines whether the monosyllabic sound with the highest output probability is a sound whose pronunciation is similar to that of a monosyllabic sound represented by any kana character whose shape is similar to that of the reference kana character.
[0060] If the determination result of S609 is affirmative, the determination unit 108 advances the process to S512, and if negative, the determination unit 108 advances the process to S701 (FIG. 7).
[0061] In S701, the determination unit 108 determines whether the sound of the monosyllable with the highest output probability is the same as the sound when the reference kana character is read by a person with an underdeveloped palate.
[0062] If the determination result in S701 is affirmative, the determination unit 108 advances the process to S703, and if negative, the determination unit 108 advances the process to S702.
[0063] In S702, the determination unit 108 determines whether the monosyllabic sound with the highest output probability is the same as a sound whose pronunciation is similar to the sound when the reference kana character is read by a person with an underdeveloped palate.
[0064] If the determination result in S702 is affirmative, the determination unit 108 advances the process to S703, and if negative, the determination unit 108 advances the process to S706.
[0065] In S703, the determination unit 108 determines whether or not the speaker has an underdeveloped palate based on known information.
[0066] If the determination result in S703 is positive, the determination unit 108 advances the process to S704, and if negative, the determination unit 108 advances the process to S706. Note that S703 may be omitted.
[0067] In S704, the determination unit 108 determines that the speaker is reading the reference kana characters correctly.
[0068] In S705, the determination unit 108 Displayed kana characters (standard kana characters) - Determined "reading": Displayed kana characters ·Judgment result: Positive ·Similar characters: NULL ·Undeveloped palate:YES are recorded in a set in the storage device 103.
[0069] In S706, the determination unit 108 determines that the speaker has read the reference kana character incorrectly. In particular, the determination unit 108 determines that the speaker has read the reference kana character incorrectly due to reasons other than the similarity of the kana character's shape. Note that the determination unit 108 does not need to determine the sound of the kana character that the speaker has uttered.
[0070] In S707, the determination unit 108 Displayed kana characters (standard kana characters) -Determined reading: NULL ·Judgment result: False ·Similar characters: NO ·Undeveloped palate: NO are recorded in a set in the storage device 103.
[0071] 6 are modifications of S511, S513, and S516, respectively. S611 differs from S511 in that "pronunciation by a person with an underdeveloped palate: NO" is added to the set, S613 differs from S513 in that "pronunciation by a person with an underdeveloped palate: NO" is added to the set, and S616 differs from S516 in that "pronunciation by a person with an underdeveloped palate: NULL" is added to the set.
[0072] According to the second method shown in Figures 6 and 7, pronunciations by a speaker with an underdeveloped palate that would normally be judged to be incorrectly pronounced sounds can be judged to be correctly pronounced sounds by a speaker with an underdeveloped palate.
[0073] 8 and 9 show flowcharts for explaining a third determination method performed by the determination unit 108. The third method is the second method to which a section for dealing with ambiguous pronunciation has been added. Here, ambiguous pronunciation refers to a pronunciation in which a speaker does not know how to read a reference kana character and pronounces a sound that is intermediate between the sound that would be produced if the reference kana character were read correctly and the sound that would be produced if the reference kana character were read incorrectly. For example, this refers to a pronunciation in which a speaker pronounces a sound that is intermediate between the sound that would be produced if the reference kana character were read correctly and the sound that would be produced if a kana character that is similar in shape to the reference kana character were read correctly.
[0074] Among the steps included in the third method, the steps that are the same as the steps included in the first or second method are given the same reference numerals and redundant explanations will be omitted.
[0075] If the determination in S702 is negative, the determination unit 108 proceeds to S901. If the determination in S703 is negative, the determination unit 108 proceeds to S901, and if the determination in S703 is positive, the determination unit 108 proceeds to S903.
[0076] In S901, the determination unit 108 determines whether the sound of the monosyllable with the highest output probability is the same as the sound when the reference kana character is read ambiguously.
[0077] If the determination result of S901 is affirmative, the determination unit 108 advances the process to S905, and if negative, the determination unit 108 advances the process to S902.
[0078] In S902, the determination unit 108 determines whether the monosyllabic sound with the highest output probability is a sound whose pronunciation is similar to the sound when the reference kana character is read ambiguously.
[0079] If the determination result in S902 is affirmative, the determination unit 108 advances the process to S905, and if the determination result is negative, the determination unit 108 advances the process to S907.
[0080] In S903, the determination unit 108 determines that the speaker is reading the reference kana characters correctly.
[0081] In S904, the determination unit 108 Displayed kana characters (standard kana characters) - Determined "reading": Displayed kana characters ·Judgment result: Positive ·Similar characters: NULL ·Undeveloped palate:YES ·Ambiguous: NO are recorded in a set in the storage device 103.
[0082] In S905, the determination unit 108 determines that the speaker is reading the reference kana characters ambiguously and incorrectly.
[0083] In S906, the determination unit 108 Displayed kana characters (standard kana characters) -Determined reading: NULL ·Judgment result: False ·Similar characters: NO ·Undeveloped palate: NO ·Ambiguous:YES are recorded as a pair in the storage device 103. The determined "reading" may be two kana characters indicating two sounds that make up an ambiguous middle sound.
[0084] In S907, the determination unit 108 determines that the speaker has read the reference kana character incorrectly. In particular, the determination unit 108 determines that the speaker has read the reference kana character incorrectly due to reasons other than the similarity of the kana character's shape. Note that the determination unit 108 does not need to determine the sound of the kana character that the speaker has uttered.
[0085] In S908, the determination unit 108 Displayed kana characters (standard kana characters) -Determined reading: NULL ·Judgment result: False ·Similar characters: NO ·Undeveloped palate: NO ·Ambiguous: NO are recorded in a set in the storage device 103.
[0086] 8 is a modified version of S616. S816 differs from S616 in that "ambiguous:NULL" is added to the set.
[0087] In the third method, steps S701, S702, and S703 corresponding to pronunciation by a speaker with an underdeveloped palate may be omitted.
[0088] 10 and 11 show flowcharts for explaining the fourth determination method performed by the determination unit 108. FIG.
[0089] In the first determination method, if the monosyllabic sound with the highest output probability in S507 is a sound that is similar in pronunciation to the reference monosyllabic sound, it is always determined in S510 that the speaker is reading the reference kana character correctly. In contrast, in the fourth determination method, even if the monosyllabic sound with the highest output probability in S507 is a sound that is similar in pronunciation to the reference monosyllabic sound, there are cases where it is determined that the speaker is not reading the reference kana character correctly.
[0090] Furthermore, in the first determination method, whenever the monosyllabic sound with the highest output probability in S508 is identical to the monosyllabic sound represented by any kana character whose shape is similar to the reference kana character, it is determined in S512 that the speaker is reading it incorrectly due to the similarity in the shape of the kana character. In contrast, in the fourth determination method, even if the monosyllabic sound with the highest output probability in S508 is identical to the monosyllabic sound represented by any kana character whose shape is similar to the reference kana character, there are cases where it is determined that the speaker is reading it incorrectly due to a reason other than the similarity in the shape of the kana character.
[0091] Among the steps included in the fourth method, the steps that are the same as the steps included in the first method are given the same reference numerals and redundant explanations will be omitted.
[0092] If the monosyllabic sound with the highest output probability in S507 is a sound whose pronunciation is similar to that of the reference monosyllabic sound, the determination unit 108 advances the process to S1101.
[0093] In S1101, the determination unit 108 determines whether the monosyllabic sound having the second highest output probability is the same as the reference monosyllabic sound.
[0094] If the determination result of S1101 is positive, the determination unit 108 proceeds to S510; if the determination result is negative, the determination unit 108 proceeds to S508. That is, the determination unit 108 controls the process flow so that if the monosyllable sound with the highest output probability is a sound that is similar in pronunciation to the reference monosyllable sound and the monosyllable sound with the second highest output probability is identical to the reference monosyllable sound, the determination unit 108 determines that the speaker is reading the reference kana character correctly. On the other hand, the determination unit 108 controls the process flow so that even if the monosyllable sound with the highest output probability is a sound that is similar in pronunciation to the reference monosyllable sound, the determination unit 108 does not determine that the speaker is reading the reference kana character correctly if the monosyllable sound with the second highest output probability is not identical to the reference monosyllable sound.
[0095] If the monosyllabic sound with the highest output probability in S507 is a sound whose pronunciation is similar to that of the reference monosyllabic sound, the determination unit 108 may proceed to S1103 instead of S1101.
[0096] In S1103, the determination unit 108 determines whether the output probability of the reference monosyllabic sound is equal to or greater than a predetermined value.
[0097] If the determination result of S1103 is positive, the determination unit 108 proceeds to S510; if the determination result is negative, the determination unit 108 proceeds to S508. That is, the determination unit 108 controls the process flow so that if the monosyllabic sound with the highest output probability is a sound that is similar in pronunciation to the reference monosyllabic sound and the output probability of the reference monosyllabic sound is equal to or greater than a predetermined value, the determination unit 108 determines that the speaker is reading the reference kana character correctly. On the other hand, even if the monosyllabic sound with the highest output probability is a sound that is similar in pronunciation to the reference monosyllabic sound, if the output probability of the reference monosyllabic sound is not equal to or greater than a predetermined value, the determination unit 108 controls the process flow so that it does not determine that the speaker is reading the reference kana character correctly.
[0098] If the monosyllabic sound with the highest output probability in S509 is a sound whose pronunciation is similar to that of a monosyllabic sound represented by any kana character whose shape is similar to that of the reference kana character, the determination unit 108 advances the process to S1102.
[0099] In S1102, the determination unit 108 determines whether the monosyllabic sound with the second highest output probability is the same as the monosyllabic sound represented by any kana character whose shape is similar to the reference kana character.
[0100] If the determination result of S1102 is positive, the determination unit 108 proceeds to S512, and if the determination result is negative, the determination unit 108 proceeds to S514. In other words, the determination unit 108 controls the flow of the process so that if the monosyllabic sound with the highest output probability is a sound whose pronunciation is similar to that of a monosyllabic sound represented by any kana character whose shape is similar to that of the reference kana character, and if the monosyllabic sound with the second highest output probability is the same as that of a monosyllabic sound represented by any kana character whose shape is similar to that of the reference kana character, the determination unit 108 determines that the speaker is reading incorrectly due to the similarity in shape. On the other hand, the judgment unit 108 controls the flow of processing so as not to judge that the speaker is reading incorrectly due to the similarity in shape, even if the monosyllabic sound with the highest output probability is a sound similar in pronunciation to the monosyllabic sound represented by any kana character whose shape is similar to the reference kana character, unless the monosyllabic sound with the second highest output probability is identical to the monosyllabic sound represented by any kana character whose shape is similar to the reference kana character.
[0101] If the monosyllabic sound with the highest output probability in S509 is a sound whose pronunciation is similar to that of a monosyllabic sound represented by any kana character whose shape is similar to that of the reference kana character, the determination unit 108 may proceed to S1104 instead of S1102.
[0102] In S1104, the determination unit 108 determines whether the output probability of the sound of any kana character that is similar in shape to the reference kana character is equal to or greater than a predetermined value.
[0103] If the determination result of S1104 is positive, the determination unit 108 proceeds to S512; if the determination result is negative, the determination unit 108 proceeds to S514. That is, if the monosyllabic sound with the highest output probability is a sound whose pronunciation is similar to that of a monosyllabic sound represented by any kana character whose shape is similar to that of the reference kana character, and the output probability of the sound of any kana character whose shape is similar to that of the reference kana character is equal to or greater than a predetermined value, the determination unit 108 controls the processing flow so as not to determine that the speaker is reading incorrectly due to the similarity of shape. On the other hand, if the monosyllabic sound with the highest output probability is a sound whose pronunciation is similar to that of a monosyllabic sound represented by any kana character whose shape is similar to that of the reference kana character, but the output probability of the sound of any kana character whose shape is similar to that of the reference kana character is not equal to or greater than a predetermined value, the determination unit 108 controls the processing flow so as not to determine that the speaker is reading incorrectly due to the similarity of shape.
[0104] [Second embodiment] Fig. 12 shows an example of a screen displayed by the display device 105 in the second embodiment. As shown in Fig. 12, the display device 105 displays a predetermined number of reference kana characters, each representing a predetermined number of reference monosyllabic sounds, on the display device in a predetermined order. In the example of Fig. 12, the predetermined number is 50. In addition, in the example of Fig. 12, the predetermined order is an order in which the columns are arranged from right to left and the characters are arranged from top to bottom in each column. According to this order, "yo", "zu", "gyu", "ni", "pya", "to", "nya", ... and "cha" are arranged in this order.
[0105] The speaker, who is the test subject, basically reads aloud in this order, following the rule of reading each standard kana character one by one with a pause in between. However, the speaker may skip some standard kana characters. The speech from the reading aloud is input to the speech input unit 106, and speech data indicating the waveform of the speech is input to the speech recognition unit 107. The speech recognition unit 107 divides the speech data into monosyllables in time series. The speech data divided into monosyllables in time series is stored in the storage device 103 by the processor 101. The monosyllable speech data referred to here is different from the recognition results by the speech recognition unit 107 used in the above-mentioned method, and simply indicates the waveform of the monosyllable speech.
[0106] The processor 101, which is operated by a program, executes the following method. However, a functional unit that executes the following method may be connected to the bus 112.
[0107] In S1301, the processor 101 acquires all time-series voice data for each monosyllable from the storage device 103.
[0108] In S1302, the processor 101 determines whether the number of monosyllable fragments in the acquired speech data is the same as the number of displayed reference kana characters. The acquired number of fragments is the number of monosyllable sounds detected in time series based on recognition, and is assumed to be the same as the number of reference kana characters read by the speaker.
[0109] If the number of acquired fragments is the same as the number of displayed reference kana characters, it is assumed that the speaker has read all of the displayed reference kana characters. On the other hand, if the number of acquired fragments is less than the number of displayed reference kana characters, it is assumed that the speaker has skipped some of the displayed reference kana characters. Note that if multiple consecutive fragments with the same recognition result appear, they may be treated as a single fragment.
[0110] If the determination result in S1302 is affirmative, the processor 101 advances the process to S1316; if the determination result is negative, the processor 101 advances the process to repeat the process from S1304 to S1310.
[0111] The repeated process from S1304 to S1310 is repeated for all possible combinations of skipped characters. The number of skipped standard kana characters is the difference obtained by subtracting the number of acquired audio data fragments from the total number of displayed standard kana characters. The number of combinations is n C m Here, n is the number of displayed standard kana characters, and m is the number of skipped standard kana characters. The number of combinations is the same as the number of combinations of selecting the standard kana characters read from all the displayed standard kana characters ( n C m = n C n-m ).
[0112] In step S1305, which is the first step in each iteration from steps S1304 to S1310, the processor 101 assigns the speech data fragments to the reference kana characters that have not been skipped in order, in accordance with the combination of skips in the current iteration.
[0113] Next, the process from S1306 to S1308 is repeated for all the reference kana characters that have not been skipped in the current skip combination.
[0114] The processor 101 controls the repeated processing from S1304 to S1310 and the repeated processing from S1306 to S1308.
[0115] In S1307 in each of the double repetitions, the judgment unit 108 performs a correct / incorrect judgment as to whether the speaker has correctly read aloud the reference kana character corresponding to the current repetition using one of the methods described above.
[0116] When all repetitions from S1306 to S1308 are completed, the processor 101 advances the process to S1309.
[0117] In S1309, the processor 101 calculates the number of correct answers for the combination of skipped readings corresponding to the current iteration from S1304 to S1310. The number of correct answers here refers to the number of reference kana characters determined to have been read correctly in S1307.
[0118] When all repetitions from S1304 to S1310 are completed, the processor 101 advances the process to S1311.
[0119] In S1311, the processor 101 compares the number of correct answers calculated in S1309 for all repetitions from S1304 to S1310, and estimates the skip combination with the highest number of correct answers as the actual skip combination.
[0120] In S1312, the processor 101 assigns the sounds of each of the acquired monosyllabic characters to the reference kana characters that have not been skipped in order, in accordance with the estimated combination.
[0121] In the repetition of steps S1312 to S1315, under the control of the processor 101, the judgment unit 108 judges whether the speaker has correctly read aloud all characters that have not been skipped, using any of the methods described above.
[0122] If the results of the determination in S1307 are saved, the results of the determination for each non-skipped reference kana character in the skip combination with the highest number of correct answers from the saved partial results may be read out. In this case, S1312 and the repetition of S1313 to S1315 do not need to be executed.
[0123] In S1316, the processor 101 assigns the sound of each of the acquired monosyllabic characters to the reference kana characters in turn.
[0124] In the repetition of steps S1317 to S1319, under the control of the processor 101, the determination unit 108 performs a correctness determination on all characters as to whether or not the speaker has read them aloud correctly, using any of the methods described above.
[0125] When the processor 101 has completed the repetition of steps S1313 to S1315, the process proceeds to step S1320. When the processor 101 has completed the repetition of steps S1317 to S1319, the process proceeds to step S1320.
[0126] In S1320, the processor 101 generates a report. The processor 101 stores the generated report in the storage device 103, or outputs the report from an output device (such as a printer) via the input / output unit 109.
[0127] Here, the report includes, for example, the determination result by the determination unit 103 for each reference kana character. Here, the determination result by the determination unit 103 includes, for example, the pairs recorded in S511, S513, S515, or S516 in the case of the first method. A two-dimensional table such as that shown in Fig. 12 may be generated, and the contents of each item of such pairs may be added to each character included in the table.
[0128] The report may also include information on whether each reference kana character was skipped. If skips are intermittent, the report may also include information on the number of skips and the number of consecutive characters for each skip. A two-dimensional table such as that shown in Figure 12 may be generated, and a symbol indicating whether each character in the table was skipped may be added.
[0129] FIG. 14 is a functional block diagram showing the configuration of the speech recognition unit 107 (see FIG. 1) according to the first or second embodiment and the configuration of a model generation device 1401 for generating a trained model for recognition 1414 included in the speech recognition unit 107.
[0130] The model generation device 1401 receives a large amount of training data 1402, each consisting of a set of monosyllabic sound identification information (e.g., character code or character code string) 1403, monosyllabic speech data 1404, and acoustic model 1405, and trains the recognition machine learning model 1406 to generate a recognition trained model 1414. Here, the monosyllabic speech data is obtained by, for example, having a young person read reference kana characters representing monosyllabic sounds. The acoustic model 1405 is obtained from an individual young person who actually reads reference kana characters representing monosyllabic sounds. The training data 1402 can be generated by having a certain number of young people who can read correctly read a large number of various reference kana characters representing monosyllabic sounds. Using this training data 1402, the model generation device 1401 can convert the recognition machine learning model 1406 into the recognition trained model 1414.
[0131] When monosyllabic speech data 1412 is input, speech recognition calculation unit 1411 uses acoustic model 1413 and recognition trained model 1414 to output recognition result 1415. Recognition result 1415 includes multiple pairs of monosyllabic sound identification information 1416 and output probability 1417. Here, monosyllabic speech data is obtained by having a subject speaker read standard kana characters representing monosyllabic sounds. Acoustic model 1413 is preferably that of the subject speaker, but a standard acoustic model for a young person may also be used.
[0132] [Third embodiment] The third embodiment takes into consideration the speaker's language development and palate development.
[0133] 15 , training data 1501 includes monosyllabic sound identification information 1403, monosyllabic speech data 1404, acoustic model 1405, information indicating language development level 1502, and information indicating palatal development level 1503. Using such training data 1501, model generation device 1401 can convert machine learning model for recognition 1504 into trained model for recognition 1505.
[0134] When monosyllabic speech data 1412 is input, speech recognition calculation unit 1411 outputs recognition result 1415 using acoustic model 1413, information 1506 indicating the speaker's language development level, information 1507 indicating the speaker's palatal development level, and trained model for recognition 1505.
[0135] Furthermore, in this embodiment, the fourth monosyllabic reading determination method (see FIGS. 10 and 11) is basically used, and the predetermined value used in S1103 may be adjusted according to the speaker's level of language development and palatal development. In other words, if the speaker's level of language development and palatal development is low, the predetermined value used in S1103 may be lowered to increase the probability of determining that the speaker is reading correctly, and conversely, if the speaker's level of language development and palatal development is high, the predetermined value used in S1103 may be higher to decrease the probability of determining that the speaker is reading correctly.
[0136] Furthermore, in this embodiment, the fourth monosyllabic reading determination method (see FIGS. 10 and 11) is basically used, and the predetermined value used in S1104 may be adjusted according to the speaker's level of language development and palatal development. In other words, if the speaker's level of language development and palatal development is low, the predetermined value used in S1104 may be lowered to increase the probability of determining that the reading is incorrect due to similarity of shape, and conversely, if the speaker's level of language development and palatal development is high, the predetermined value used in S1104 may be higher to decrease the probability of determining that the reading is incorrect due to similarity of shape.
[0137] [Fourth embodiment] The fourth embodiment takes into consideration the age of the speaker.
[0138] 16, training data 1601 includes monosyllabic sound identification information 1403, monosyllabic speech data 1404, an acoustic model 1405, and speaker age information 1602. Using such training data 1601, model generation device 1401 can convert machine learning model for recognition 1603 into trained model for recognition 1604.
[0139] When monosyllabic speech data 1412 is input, speech recognition calculation unit 1411 uses acoustic model 1413, information indicating the speaker's age 1605, and trained model for recognition 1604 to output recognition result 1415.
[0140] Furthermore, in this embodiment, the fourth monosyllabic reading determination method (see FIGS. 10 and 11) is basically used, and the predetermined value used in S1103 may be adjusted according to the age of the speaker. In other words, if the speaker is young, the predetermined value used in S1103 may be lowered to increase the probability of determining that the speaker is reading correctly, and conversely, if the speaker is old, the predetermined value used in S1103 may be higher to decrease the probability of determining that the speaker is reading correctly.
[0141] Furthermore, in this embodiment, the fourth monosyllabic reading determination method (see FIGS. 10 and 11) is basically used, and the predetermined value used in S1104 may be adjusted according to the speaker's level of language development and palate development. In other words, if the speaker is young, the predetermined value used in S1104 may be lowered to increase the probability of determining that the reading is incorrect due to similarity in shape, and conversely, if the speaker is older, the predetermined value used in S1104 may be higher to decrease the probability of determining that the reading is incorrect due to similarity in shape.
[0142] [Fifth embodiment] The fifth embodiment takes into consideration the dialect of the speaker.
[0143] 17, training data 1701 includes monosyllabic sound identification information 1403, monosyllabic speech data 1404, acoustic model 1405, and information indicating region 1702. Using such training data 1701, model generation device 1401 can convert machine learning model for recognition 1703 into trained model for recognition 1604. Here, the region refers to the region where the speaker resides. The region may be represented by prefecture, or may be unique based on the regional distribution of dialects.
[0144] When monosyllabic speech data 1412 is input, speech recognition calculation unit 1411 uses acoustic model 1413, information indicating region 1705, and trained model for recognition 1704 to output recognition result 1415.
[0145] In this embodiment, the table shown in Fig. 4 may be switched depending on the dialect. That is, kana characters that are similar in sound but not in shape to a reference kana character may be switched depending on the dialect. Therefore, the determination results in S507 and S509 are affected by the dialect being spoken and the table that is switched depending on the dialect.
[0146] [Sixth embodiment] In the sixth embodiment, the judgment conditions are changed based on the learning progress and changes of the speaker who is the subject. The fourth monosyllabic reading judgment method shown in Figs. 10 and 11 will be described as an example.
[0147] The sounds adopted as sounds similar in pronunciation to the monosyllabic sounds used as the reference in S507 may be changed according to the speaker's learning progress. For example, if the speaker's learning progress is low, the number of such sounds may be increased, and if the speaker's learning progress is high, the number of such sounds may be reduced.
[0148] Furthermore, sounds whose pronunciation is similar to that of a monosyllabic sound represented by any kana character whose shape is similar to the reference kana character used in S509 may be changed according to the speaker's learning progress. For example, if the speaker's learning progress is low, the number of such sounds may be increased, and if the speaker's learning progress is high, the number of such sounds may be reduced.
[0149] Furthermore, if the learning progress is low, S1101 may be changed from "Is the monosyllabic sound with the second highest output probability identical to the reference monosyllabic sound?" to "Is the monosyllabic sound with the second or third highest output probability identical to the reference monosyllabic sound?"
[0150] Furthermore, if the learning progress level is low, S1102 may be changed from "Is the monosyllabic sound with the second highest output probability the same as the monosyllabic sound represented by any kana character whose shape is similar to the reference kana character?" to "Is the monosyllabic sound with the second or third highest output probability the same as the monosyllabic sound represented by any kana character whose shape is similar to the reference kana character?"
[0151] Furthermore, the predetermined value used in S1103 may be changed depending on the learning progress of the speaker. That is, if the learning progress of the speaker is low, the predetermined value may be set low, and if the learning progress of the speaker is high, the predetermined value may be set high.
[0152] Furthermore, the predetermined value used in S1104 may be changed depending on the learning progress of the speaker. That is, if the learning progress of the speaker is low, the predetermined value may be set low, and if the learning progress of the speaker is high, the predetermined value may be set high.
[0153] The speaker's learning progress may be measured for each monosyllabic sound. For example, if the speaker has trained to pronounce a certain monosyllabic sound correctly, it may be determined that the speaker has made progress in learning that monosyllabic sound.
[0154] [Seventh embodiment] The seventh embodiment is a flexible judgment method, which will be described by taking the first monosyllabic reading judgment method shown in Figs.
[0155] If the determination in S506 is NO and the determination in S507 is YES, it may be interpreted that the reading is correct but the pronunciation is not very good, and this interpretation may be fed back to the speaker by voice or text, which will encourage the speaker to make an effort to improve their pronunciation.
[0156] Thereafter, if the determination in S506 is YES, feedback may be given by voice or text indicating that the pronunciation was previously poor but has improved.
[0157] If the process does not proceed to S512, the speaker may be given feedback by voice or text indicating that the speaker has mistaken the reference character for a character that is similar in shape to the reference character. This will encourage the speaker to try not to make the same mistake.
[0158] If the process does not proceed to S512 after that, the user may be given feedback by voice or text indicating that the user no longer mistakes the reference character for a character that is similar in shape to the reference character, even though the user may have previously mistaken the reference character for a character that is similar in shape to the reference character.
[0159] The monosyllabic reading determination device can be realized by hardware, software, or a combination of these. The monosyllabic reading determination method performed by the monosyllabic reading determination device can also be realized by hardware, software, or a combination of these. "Realized by software" here means that the method is realized by a computer reading and executing a program.
[0160] The program can be stored and supplied to a computer using various types of non-transitory computer-readable media. Non-transitory computer-readable media include various types of tangible storage media. Examples of non-transitory computer-readable media include magnetic recording media (e.g., flexible disks, magnetic tapes, hard disk drives), magneto-optical recording media (e.g., magneto-optical disks), CD-ROMs (Read Only Memory), CD-Rs, CD-R / Ws, and semiconductor memories (e.g., mask ROMs, PROMs (Programmable ROMs), EPROMs (Erasable PROMs), flash ROMs, and RAMs (Random Access Memory)). The program may also be supplied to a computer by various types of transitory computer-readable media. Examples of transitory computer-readable media include electrical signals, optical signals, and electromagnetic waves. The transitory computer-readable media can supply the program to a computer via a wired communication path such as an electric wire or optical fiber, or via a wireless communication path.
[0161] The present invention can be embodied in various other forms without departing from its spirit or main characteristics. Therefore, the above-described embodiments are merely examples and should not be interpreted as limiting. The scope of the present invention is defined by the claims and is not limited to the text of the specification. Furthermore, all modifications and variations within the equivalent range of the claims are within the scope of the present invention.
[0162] A part or all of the above-described embodiments can be described as, but not limited to, the following supplementary notes.
[0163] {Appendix 1} a display control means for displaying on a display device reference kana characters indicating reference monosyllabic sounds; a speech recognition means for recognizing the speech of a speaker who reads the reference kana characters aloud and outputting identification information for identifying each of a plurality of monosyllabic sounds obtained by the recognition together with the output probability of each sound; a determination means for determining whether the speaker is reading the reference kana characters correctly based on identification information for identifying the reference monosyllabic sounds, and the identification information and output probabilities for the plurality of types of monosyllabic sounds; Equipped with Monosyllabic reading judgment device.
[0164] {Appendix 2} The determination means If the monosyllabic sound with the highest output probability is identical to the reference monosyllabic sound, or if the monosyllabic sound with the highest output probability is a sound whose pronunciation is similar to that of the reference monosyllabic sound, it is determined that the speaker is reading the reference kana character correctly. 2. The monosyllabic reading determination device according to claim 1.
[0165] {Appendix 3} The system further comprises means for outputting information to a speaker indicating that the pronunciation is correct but the pronunciation is not very good when the monosyllabic sound with the highest output probability is a sound that is similar in pronunciation to the reference monosyllabic sound. 3. The monosyllabic reading determination device according to claim 2.
[0166] {Appendix 4} The system further comprises means for outputting information indicating that pronunciation has improved when the monosyllabic sound with the highest output probability for a current utterance is a sound similar in pronunciation to the reference monosyllabic sound, and when the monosyllabic sound with the highest output probability for a past utterance is identical to the reference monosyllabic sound. 3. The monosyllabic reading determination device according to claim 2.
[0167] {Appendix 5} The determination means Even if the monosyllabic sound with the highest output probability is similar in pronunciation to the reference monosyllabic sound, if the reference monosyllabic sound is not identical to the monosyllabic sound with the nth highest output probability, it is determined that the speaker is reading the reference kana character incorrectly; The n is a number selected from integers of 2 or more. 3. The monosyllabic reading determination device according to claim 2.
[0168] {Appendix 6} The determination means Even if the monosyllabic sound with the highest output probability is similar in pronunciation to the reference monosyllabic sound, if the output probability of the reference monosyllabic sound is less than a predetermined value, it is determined that the speaker is reading the reference kana character incorrectly. 3. The monosyllabic reading determination device according to claim 2.
[0169] {Appendix 7} The predetermined value is adjusted according to the speaker's language development level and palate development level, or the speaker's age. 7. The monosyllabic reading determination device according to claim 6.
[0170] {Appendix 8} The determination means If the monosyllabic sound with the highest output probability is identical to the monosyllabic sound represented by any kana character whose shape is similar to the reference kana character, or if the monosyllabic sound with the highest output probability is a sound whose pronunciation is similar to the monosyllabic sound represented by any kana character whose shape is similar to the reference kana character, it is determined that the speaker is reading the reference kana character incorrectly. 8. The monosyllabic reading determination device according to any one of appendices 1 to 7.
[0171] {Appendix 9} The determination means If the monosyllabic sound with the highest output probability is identical to the monosyllabic sound represented by any kana character whose shape is similar to the reference kana character, or if the monosyllabic sound with the highest output probability is a sound whose pronunciation is similar to that of the monosyllabic sound represented by any kana character whose shape is similar to the reference kana character, it is determined that the speaker is misreading the reference kana character as any of the kana characters. 9. The monosyllabic reading determination device according to any one of appendices 1 to 8.
[0172] {Appendix 10} The determination means Even if the monosyllabic sound with the highest output probability is a sound similar in pronunciation to the monosyllabic sound represented by any kana character whose shape is similar to the reference kana character, if the monosyllabic sound represented by any kana character whose shape is similar to the reference kana character is not identical to the monosyllabic sound with the nth highest output probability, it is determined that the speaker has not mistaken the reference kana character for any of the kana characters, but has misread the reference kana character, The n is a number selected from integers of 2 or more. 10. The monosyllabic reading determination device according to claim 9.
[0173] {Appendix 11} The determination means Even if the monosyllabic sound with the highest output probability is a sound whose pronunciation is similar to that of a monosyllabic sound represented by any kana character whose shape is similar to the reference kana character, if the output probability of the sound of any of the kana characters is less than a predetermined value, it is determined that the speaker has not mistaken the reference kana character for any of the kana characters, but has misread the reference kana character. 10. The monosyllabic reading determination device according to claim 9.
[0174] {Appendix 12} The determination means If the sound of the monosyllable with the highest output probability is the same as the sound when the reference kana character is read by a person with an underdeveloped palate, or if the sound of the monosyllable with the highest output probability is a sound similar in pronunciation to the sound when the reference kana character is read by a person with an underdeveloped palate, it is determined that the speaker is reading the reference kana character correctly. 12. The monosyllabic reading determination device according to any one of appendices 1 to 11.
[0175] {Appendix 13} The determination means If the speaker is a person with an underdeveloped palate and the sound of the monosyllable with the highest output probability is the same as the sound when the reference kana character is read by a person with an underdeveloped palate, or if the speaker is a person with an underdeveloped palate and the sound of the monosyllable with the highest output probability is a sound similar in pronunciation to the sound when the reference kana character is read by a person with an underdeveloped palate, it is determined that the speaker is reading the reference kana character correctly. 13. The monosyllabic reading determination device according to any one of appendices 1 to 12.
[0176] {Appendix 14} The determination means If the sound of the monosyllable with the highest output probability is the same as the sound when the reference kana character is read ambiguously, or if the sound of the monosyllable with the highest output probability is a sound similar in pronunciation to the sound when the reference kana character is read ambiguously, it is determined that the speaker is reading the reference kana character incorrectly. 14. The monosyllabic reading determination device according to any one of appendices 1 to 13.
[0177] {Appendix 15} the speech recognition means recognizes the speech of a speaker who reads the reference kana characters aloud by referring to at least a recognition trained model and an acoustic model; The trained model for recognition is generated by training a machine learning model for recognition using a large number of pairs of monosyllabic sound identification information, monosyllabic speech data, and acoustic models. 15. The monosyllabic reading determination device according to any one of appendices 1 to 14.
[0178] {Appendix 16} the speech recognition means recognizes the speech of a speaker who reads the reference kana characters aloud by referring to at least a recognition trained model, an acoustic model of the speaker, the language development level of the speaker, and the palatal development level of the speaker; The trained model for recognition is generated by training a machine learning model for recognition using a large number of sets of monosyllabic sound identification information, monosyllabic speech data uttered by a speaker, an acoustic model of the speaker, the speaker's language development level, and the speaker's palatal development level. 15. The monosyllabic reading determination device according to any one of appendices 1 to 14.
[0179] {Appendix 17} the speech recognition means recognizes the speech of a speaker who reads the reference kana characters aloud by referring to at least a recognition trained model, an acoustic model of the speaker, and the age of the speaker; The trained model for recognition is generated by training a machine learning model for recognition using a large number of pairs of monosyllabic sound identification information, monosyllabic speech data uttered by a speaker, an acoustic model of the speaker, and the speaker's age. 15. The monosyllabic reading determination device according to any one of appendices 1 to 14.
[0180] {Appendix 18} the speech recognition means recognizes the speech of a speaker who reads the reference kana characters aloud by referring to at least a recognition trained model, an acoustic model of the speaker, and an area where the speaker lives; The trained model for recognition is generated by training a machine learning model for recognition using a large number of pairs of monosyllabic sound identification information, monosyllabic speech data uttered by a speaker, an acoustic model of the speaker, and an area where the speaker lives. 15. The monosyllabic reading determination device according to any one of appendices 1 to 14.
[0181] {Appendix 19} a display control means for displaying a predetermined number of reference kana characters, each representing a predetermined number of reference monosyllabic sounds, on a display device in a predetermined order; a means for recognizing the speech of a speaker who reads aloud in accordance with a rule that the predetermined number of standard kana characters are read aloud in the predetermined order, and acquiring the number of detected monosyllabic sounds detected in time series based on the recognition; means for estimating, when the number of detected characters is less than the predetermined number, that the number of skipped reference kana characters is equal to the difference obtained by subtracting the number of detected characters from the predetermined number; a means for associating each detected monosyllabic sound with each non-skipped reference kana character for all combinations of the predetermined number of reference kana characters and the difference number of skipped reference kana characters so that the order of the reference kana characters and the order of the detected monosyllabic sounds are aligned, and for each pair of associated reference kana characters and monosyllabic sounds, using a monosyllabic reading determination device, determining whether or not the speaker has correctly read the reference kana character; a means for adopting a result of determination by a monosyllabic reading determination device corresponding to the combination with the largest number of the reference kana characters correctly read by the speaker; Equipped with The monosyllabic reading determination device is a monosyllabic reading determination system that is the monosyllabic reading determination device described in any one of appendices 1 to 18.
[0182] {Appendix 20} a display control means for displaying a predetermined number of reference kana characters, each representing a predetermined number of reference monosyllabic sounds, on a display device in a predetermined order; a means for recognizing the speech of a speaker who reads aloud in accordance with a rule that the predetermined number of standard kana characters are read aloud in the predetermined order, and acquiring the number of detected monosyllabic sounds detected in time series based on the recognition; means for estimating, when the number of detected characters is less than the predetermined number, that the number of skipped reference kana characters is equal to the difference obtained by subtracting the number of detected characters from the predetermined number; a means for associating each detected monosyllabic sound with each non-skipped reference kana character for all combinations of the predetermined number of reference kana characters and the difference number of skipped reference kana characters so that the order of the reference kana characters and the order of the detected monosyllabic sounds are aligned, and for each pair of associated reference kana characters and monosyllabic sounds, using a monosyllabic reading determination device, determining whether or not the speaker has correctly read the reference kana character; a means for adopting a result of determination by a monosyllabic reading determination device corresponding to the combination with the largest number of the reference kana characters correctly read by the speaker; Equipped with Monosyllabic interpretation judgment system.
[0183] {Appendix 21} a display control step of displaying on a display device reference kana characters indicating the reference monosyllabic sounds; a speech recognition step of recognizing the speech of a speaker who reads the reference kana characters aloud and outputting identification information for identifying each of a plurality of types of monosyllabic sounds obtained by the recognition together with the output probability of each sound; a determining step of determining whether the speaker is reading the reference kana characters correctly based on identification information for identifying the reference monosyllabic sound, and the identification information and output probabilities for the plurality of types of monosyllabic sounds; having A method for assessing monosyllabic reading aloud.
[0184] {Appendix 22} A program for causing a computer to function as the monosyllable determination device according to any one of appendices 1 to 18. [Explanation of symbols]
[0185] 104 Display control unit 105 Display device 106 Audio input section 107 Voice Recognition Unit 108 Judgment section
Claims
1. a display control means for displaying on a display device reference kana characters indicating reference monosyllabic sounds; a speech recognition means for recognizing the speech of a speaker who reads the reference kana characters aloud and outputting identification information for identifying each of a plurality of monosyllabic sounds obtained by the recognition together with the output probability of each sound; a determination means for determining whether the speaker is reading the reference kana characters correctly based on identification information for identifying the reference monosyllabic sounds, and the identification information and output probabilities for the plurality of types of monosyllabic sounds; Equipped with Monosyllabic reading judgment device.
2. The determination means If the monosyllabic sound with the highest output probability is identical to the reference monosyllabic sound, or if the monosyllabic sound with the highest output probability is a sound whose pronunciation is similar to that of the reference monosyllabic sound, it is determined that the speaker is reading the reference kana character correctly. The monosyllabic reading determination device according to claim 1 .
3. The system further comprises means for outputting information to a speaker indicating that the pronunciation is correct but the pronunciation is not very good when the monosyllabic sound with the highest output probability is a sound that is similar in pronunciation to the reference monosyllabic sound. The monosyllabic reading determination device according to claim 2.
4. The device further comprises means for outputting information indicating that pronunciation has improved when, for a past utterance, the monosyllabic sound with the highest output probability is a sound whose pronunciation is similar to that of the reference monosyllabic sound, and for a current utterance, the monosyllabic sound with the highest output probability is identical to that of the reference monosyllabic sound. The monosyllabic reading determination device according to claim 2.
5. The determination means Even if the monosyllabic sound with the highest output probability is similar in pronunciation to the reference monosyllabic sound, if the reference monosyllabic sound is not identical to the monosyllabic sound with the nth highest output probability, it is determined that the speaker is reading the reference kana character incorrectly; The n is a number selected from integers of 2 or more. The monosyllabic reading determination device according to claim 2.
6. The determination means Even if the monosyllabic sound with the highest output probability is similar in pronunciation to the reference monosyllabic sound, if the output probability of the reference monosyllabic sound is less than a predetermined value, it is determined that the speaker is reading the reference kana character incorrectly. The monosyllabic reading determination device according to claim 2.
7. The predetermined value is adjusted according to the speaker's language development level and palate development level, or the speaker's age. The monosyllabic pronunciation determination device according to claim 6.
8. The determination means If the monosyllabic sound with the highest output probability is identical to the monosyllabic sound represented by any kana character whose shape is similar to the reference kana character, or if the monosyllabic sound with the highest output probability is a sound whose pronunciation is similar to the monosyllabic sound represented by any kana character whose shape is similar to the reference kana character, it is determined that the speaker is reading the reference kana character incorrectly. The monosyllabic reading determination device according to claim 1 .
9. The determination means If the monosyllabic sound with the highest output probability is identical to the monosyllabic sound represented by any kana character whose shape is similar to the reference kana character, or if the monosyllabic sound with the highest output probability is a sound whose pronunciation is similar to that of the monosyllabic sound represented by any kana character whose shape is similar to the reference kana character, it is determined that the speaker is misreading the reference kana character as any of the kana characters. The monosyllabic reading determination device according to claim 1 .
10. The determination means Even if the monosyllabic sound with the highest output probability is a sound similar in pronunciation to the monosyllabic sound represented by any kana character whose shape is similar to the reference kana character, if the monosyllabic sound represented by any kana character whose shape is similar to the reference kana character is not identical to the monosyllabic sound whose output probability is within the nth highest, it is determined that the speaker has not mistaken the reference kana character for any of the kana characters, but has misread the reference kana character; The n is a number selected from integers of 2 or more. The monosyllabic reading determination device according to claim 9.
11. The determination means Even if the monosyllabic sound with the highest output probability is a sound whose pronunciation is similar to that of a monosyllabic sound represented by any kana character whose shape is similar to the reference kana character, if the output probability of the sound of any of the kana characters is less than a predetermined value, it is determined that the speaker has not mistaken the reference kana character for any of the kana characters, but has misread the reference kana character. The monosyllabic reading determination device according to claim 9.
12. The determination means If the sound of the monosyllable with the highest output probability is the same as the sound when the reference kana character is read by a person with an underdeveloped palate, or if the sound of the monosyllable with the highest output probability is a sound similar in pronunciation to the sound when the reference kana character is read by a person with an underdeveloped palate, it is determined that the speaker is reading the reference kana character correctly. The monosyllabic reading determination device according to claim 1 .
13. The determination means If the speaker is a person with an underdeveloped palate and the sound of the monosyllable with the highest output probability is the same as the sound when the reference kana character is read by a person with an underdeveloped palate, or if the speaker is a person with an underdeveloped palate and the sound of the monosyllable with the highest output probability is a sound similar in pronunciation to the sound when the reference kana character is read by a person with an underdeveloped palate, it is determined that the speaker is reading the reference kana character correctly. The monosyllabic reading determination device according to claim 1 .
14. The determination means If the sound of the monosyllable with the highest output probability is the same as the sound when the reference kana character is read ambiguously, or if the sound of the monosyllable with the highest output probability is a sound similar in pronunciation to the sound when the reference kana character is read ambiguously, it is determined that the speaker is reading the reference kana character incorrectly. The monosyllabic reading determination device according to claim 1 .
15. the speech recognition means recognizes the speech of a speaker who reads the reference kana characters aloud by referring to at least a recognition trained model and an acoustic model; The trained model for recognition is generated by training a machine learning model for recognition using a large number of pairs of monosyllabic sound identification information, monosyllabic speech data, and acoustic models. The monosyllabic reading determination device according to claim 1 .
16. the speech recognition means recognizes the speech of a speaker who reads the reference kana characters aloud by referring to at least a recognition trained model, an acoustic model of the speaker, the language development level of the speaker, and the palatal development level of the speaker; The trained model for recognition is generated by training a machine learning model for recognition using a large number of sets of monosyllabic sound identification information, monosyllabic speech data uttered by a speaker, an acoustic model of the speaker, the speaker's language development level, and the speaker's palatal development level. The monosyllabic reading determination device according to claim 1 .
17. the speech recognition means recognizes the speech of a speaker who reads the reference kana characters aloud by referring to at least a recognition trained model, an acoustic model of the speaker, and the age of the speaker; The trained model for recognition is generated by training a machine learning model for recognition using a large number of pairs of monosyllabic sound identification information, monosyllabic speech data uttered by a speaker, an acoustic model of the speaker, and the speaker's age. The monosyllabic reading determination device according to claim 1 .
18. the speech recognition means recognizes the speech of a speaker who reads the reference kana characters aloud by referring to at least a recognition trained model, an acoustic model of the speaker, and an area where the speaker lives; The trained model for recognition is generated by training a machine learning model for recognition using a large number of pairs of monosyllabic sound identification information, monosyllabic speech data uttered by a speaker, an acoustic model of the speaker, and an area where the speaker lives. The monosyllabic reading determination device according to claim 1 .
19. a display control means for displaying a predetermined number of reference kana characters, each representing a predetermined number of reference monosyllabic sounds, on a display device in a predetermined order; a means for recognizing the speech of a speaker who reads aloud in accordance with a rule that the predetermined number of standard kana characters are read aloud in the predetermined order, and acquiring the number of detected monosyllabic sounds detected in time series based on the recognition; means for estimating, when the number of detected characters is less than the predetermined number, that the number of skipped reference kana characters is equal to the difference obtained by subtracting the number of detected characters from the predetermined number; a means for associating each detected monosyllabic sound with each non-skipped reference kana character for all combinations of the predetermined number of reference kana characters and the difference number of skipped reference kana characters so that the order of the reference kana characters and the order of the detected monosyllabic sounds are aligned, and for each pair of associated reference kana characters and monosyllabic sounds, using a monosyllabic reading determination device, determining whether or not the speaker has correctly read the reference kana character; a means for adopting a result of determination by a monosyllabic reading determination device corresponding to the combination with the largest number of the reference kana characters correctly read by the speaker; Equipped with The monosyllabic reading determination device is a monosyllabic reading determination device according to any one of claims 1 to 18.
20. a display control means for displaying a predetermined number of reference kana characters, each representing a predetermined number of reference monosyllabic sounds, on a display device in a predetermined order; a means for recognizing the speech of a speaker who reads aloud in accordance with a rule that the predetermined number of standard kana characters are read aloud in the predetermined order, and acquiring the number of detected monosyllabic sounds detected in time series based on the recognition; means for estimating, when the number of detected characters is less than the predetermined number, that the number of skipped reference kana characters is equal to the difference obtained by subtracting the number of detected characters from the predetermined number; a means for associating each detected monosyllabic sound with each non-skipped reference kana character for all combinations of the predetermined number of reference kana characters and the difference number of skipped reference kana characters so that the order of the reference kana characters and the order of the detected monosyllabic sounds are aligned, and for each pair of associated reference kana characters and monosyllabic sounds, using a monosyllabic reading determination device, determining whether or not the speaker has correctly read the reference kana character; a means for adopting a result of determination by a monosyllabic reading determination device corresponding to the combination with the largest number of the reference kana characters correctly read by the speaker; Equipped with Monosyllabic interpretation judgment system.
21. A display control step in which the computer displays on a display device reference kana characters indicating the reference monosyllabic sounds; a speech recognition step in which a computer recognizes the speech of a speaker who reads the reference kana characters aloud and outputs identification information for identifying each of a plurality of types of monosyllabic sounds obtained by the recognition together with the output probability of each sound; a determination step in which the computer determines whether the speaker is reading the reference kana characters correctly based on identification information for identifying the reference monosyllabic sounds, and the identification information and output probabilities for the plurality of types of monosyllabic sounds; having A method for assessing monosyllabic reading aloud.
22. A program for causing a computer to function as the monosyllabic reading determination device according to any one of claims 1 to 18.
Citation Information
Patent Citations
The character-reading device
JP1983133168U
Electronic apparatus and program
JP2016157042A
Monosyllable voice recognition equipment
JP1987269198A