Program, system and method for supporting exercise of speaking sentence
The system addresses the inadequacy of conventional speech practice systems by evaluating and updating speech ability values, enhancing speech proficiency through targeted feedback and structured practice sessions.
Patent Information
- Application Number
- JP2024004699
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-01-16
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2044-01-16
AI Technical Summary
Conventional speech practice systems fail to adequately improve speech ability, particularly in languages like English, due to a lack of support for procedural knowledge in using vocabulary in context.
A system and method that evaluates the speech of individual phrases and sentences, updating a speech ability value based on these evaluations to enhance the learning user's speech practice, including features like time limits, model voices, and native language translations to facilitate effective speech practice.
Enhances speech ability by providing targeted feedback and updating speech ability values, promoting the acquisition of procedural knowledge and improving speech proficiency through structured practice sessions.
Smart Images

Figure 2025110712000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a program, a system, and a method for assisting in speech practice of sentences.
Background Art
[0002] Conventionally, various techniques for assisting in speech practice of sentences have been proposed. For example, Patent Document 1 below discloses a foreign language learning system for performing foreign language conversation learning through dialogue with a virtual speaker. In this system, by applying a conversation level calculation algorithm to the speech voice information of the learner, an evaluation of the conversation level for the speech voice information of each sentence spoken by the learner is performed. In such a system, since learning is performed through free dialogue with a virtual speaker, it may be possible to efficiently perform foreign language conversation learning without time and space restrictions.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] However, although the conventional systems as described above can support efficient speech practice, it cannot be said that a sufficient mechanism for improving the speech ability has been realized. For example, in English learning, it is said that many learners have difficulty with listening and speaking compared to writing and reading because they lack procedural knowledge for extracting the vocabulary they have memorized according to the scenes where English is used. Therefore, for example, support for acquiring such procedural knowledge can be considered important for improving the speech ability.
[0005] One of the objectives of the embodiments of the present invention is to assist in improving the speech ability of sentences.
Means for Solving the Problem
[0006] A program according to an embodiment of the present invention is a program for assisting in speech practice of sentences, and causes one or more computers to perform steps of: acquiring a first speech audio by a learning user of a first test sentence having a first group of sentences including a first target sentence; evaluating the speech of each sentence included in the first group of sentences based on the first speech audio; evaluating the speech of the entire first test sentence based on the first speech audio based on the evaluation of the speech of each sentence included in the first group of sentences; and updating the speech ability value of the learning user for the first test sentence based on the evaluation of the speech of the entire first test sentence based on the first speech audio and the evaluation of the speech of the first target sentence based on the first speech audio, or the speech ability value of the learning user for the first target sentence updated based on the evaluation of the speech of the first target sentence based on the first speech audio.
[0007] A system according to an embodiment of the present invention includes one or more computer processors and is a system for assisting in speech practice of sentences. The one or more computer processors perform steps of: acquiring a first speech audio by a learning user of a first test sentence having a first group of sentences including a first target sentence; evaluating the speech of each sentence included in the first group of sentences based on the first speech audio; evaluating the speech of the entire first test sentence based on the first speech audio based on the evaluation of the speech of each sentence included in the first group of sentences; and updating the speech ability value of the learning user for the first test sentence based on the evaluation of the speech of the entire first test sentence based on the first speech audio and the evaluation of the speech of the first target sentence based on the first speech audio, or the speech ability value of the learning user for the first target sentence updated based on the evaluation of the speech of the first target sentence based on the first speech audio.
[0008] A method according to an embodiment of the present invention is a method for assisting in speech practice of a sentence, which is executed by one or more computers, and includes steps of: acquiring a first speech voice of a learning user for a first question sentence having a first group of phrases including a first target phrase; evaluating the speech of each phrase included in the first group of phrases based on the first speech voice; evaluating the speech of the entire first question sentence based on the first speech voice based on the evaluation of the speech of each phrase included in the first group of phrases; and updating a speech ability value of the learning user for the first question sentence based on the evaluation of the speech of the entire first question sentence based on the first speech voice and the evaluation of the speech of the first target phrase based on the first speech voice, or the speech ability value of the learning user for the first target phrase updated based on the evaluation of the speech of the first target phrase based on the first speech voice.
Advantages of the Invention
[0009] Various embodiments of the present invention assist in improving the speech ability of sentences.
Brief Description of the Drawings
[0010]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17
Mode for Carrying Out the Invention
[0011] Hereinafter, embodiments of the present invention will be described with reference to the drawings. In each drawing, the same or similar components may be assigned the same reference numerals.
[0012] FIG. 1 is a configuration diagram schematically showing the configuration of a network including a learning support server 10 according to an embodiment of the present invention. As shown in the figure, the learning support server 10 is communicably connected to a learning user terminal 30 via a communication network NW such as the Internet. In FIG. 1, only one learning user terminal 30 is shown, but the learning support server 10 can be communicably connected to a plurality of learning user terminals 30. The learning support server 10 provides a learning support service to a learning user who operates the learning user terminal 30. The learning support service has a function for supporting speech practice of sentences. The learning support server 10 is an example of a device that implements all or part of the system of the present invention.
[0013] First, the hardware configuration of the learning support server 10 will be described. The learning support server 10 is configured as a general computer and includes, as shown in FIG. 1, a computer processor 11, a main memory 12, an input / output I / F 13, a communication I / F 14, and a storage (storage device) 15, and these components are electrically connected via a bus or the like (not shown).
[0014] The computer processor 11 is composed of a CPU, GPU, FPGA, etc., or a circuit including them, reads various programs stored in the storage 15 into the main memory 12, and executes various instructions included in the programs. The main memory 12 is composed of, for example, DRAM or the like.
[0015] The input / output I / F 13 includes various input / output devices for exchanging information with an operator or the like. The input / output I / F 13 includes, for example, information input devices such as a keyboard and a pointing device (e.g., a mouse, a touch panel, etc.), a voice input device such as a microphone, and an image input device such as a camera. The input / output I / F 13 also includes an image output device such as a display and a voice output device such as a speaker.
[0016] The communication I / F 14 is implemented as hardware such as a network adapter, various communication softwares, or a combination thereof, and is configured to enable wired or wireless communication via a communication network NW or the like.
[0017] The storage 15 is constituted by, for example, a magnetic disk or a flash memory. The storage 15 stores various programs including an operating system and various data. For example, as shown in FIG. 1, the storage 15 has various tables 151 that store various information. Further, for example, the storage 15 stores a server-side program 40 according to an embodiment of the present invention. The program 40 is a program for causing the learning support server 10 to function as all or part of a system for providing a learning support service. At least a part of the server-side program 40 may be configured to be executed on the learning user terminal 30 side via a learning user terminal-side program 44 described later.
[0018] In the present embodiment, the learning support server 10 may be configured using a plurality of computers each having the hardware configuration described above. For example, the learning support server 10 may be constituted by one or more server devices.
[0019] The learning support server 10 configured in this way has functions as a web server and an application server, executes various processes in response to requests from the learning user terminal-side program 44 installed in the learning user terminal 30, and sends screen data (e.g., HTML data) and control data according to the results of the processes to the learning user terminal 30. On the learning user terminal 30, a web page or other screen based on the received data is output.
[0020] Next, the hardware configuration of the learning user terminal 30 will be described. The learning user terminal 30 is configured as a general computer and, as shown in FIG. 2, includes a computer processor 31, a main memory 32, an input / output I / F 33, a communication I / F 34, and a storage (memory device) 35, and these components are electrically connected via a bus or the like (not shown).
[0021] The computer processor 31 is constituted by a CPU, a GPU, an FPGA, etc., or a circuit including them, reads various programs stored in the storage 35 etc. into the main memory 32, and executes various instructions included in the program. The main memory 32 is constituted by, for example, DRAM or the like.
[0022] The input / output I / F 33 includes various input / output devices for exchanging information with an operator or the like. The input / output I / F 33 includes, for example, information input devices such as a keyboard and a pointing device (e.g., a mouse, a touch panel, etc.), a voice input device such as a microphone, and an image input device such as a camera. Further, the input / output I / F 33 includes an image output device such as a display and a voice output device such as a speaker.
[0023] The communication I / F 34 is implemented as hardware such as a network adapter, various communication softwares, and combinations thereof, and is configured to enable wired or wireless communication via a communication network NW or the like.
[0024] The storage 35 is composed of, for example, a magnetic disk or a flash memory. The storage 35 stores various programs including an operating system and various data. The programs stored in the storage 35 can be downloaded and installed from an application market or the like. Further, the storage 35 stores the above-described learning user terminal-side program 44. The program 44 is configured as a web browser or other application (for example, an application for a learning support service) and can be configured to execute at least a part of the server-side program 40 as described above.
[0025] In this embodiment, the learning user terminal 30 can be configured as a smartphone, a tablet terminal, a personal computer, or the like.
[0026] The learning user who operates the learning user terminal 30 configured as described above can use the learning support service provided by the learning support server 10 by executing communication with the learning support server 10 via the learning user terminal-side program 44 installed in the storage 35.
[0027] Next, the functions of the learning support server 10 configured as described above will be described. As shown in FIG. 1, the computer processor 11 of the learning support server 10 is configured to function as a management function control unit 112 and a learning management unit 114 by executing instructions included in a program (for example, at least a part of the server-side program 40) read into the main memory 12.
[0028] The management function control unit 112 is configured to execute various processes related to the control of the management function of the learning support service. For example, the management function control unit 112 transmits screen data, control data, etc. of various screens related to the management function to the learning user terminal 30, executes various processes in response to the operation input by the user via the screen output on the learning user terminal 30, and transmits screen data, control data, etc. corresponding to the result of the process to the learning user terminal 30. The management functions controlled by the management function control unit 112 include, for example, login processing (user authentication), billing control, and management of user information, etc.
[0029] The learning management unit 114 is configured to execute various processes related to the management of learning performed by the learning user. For example, the learning management unit 114 transmits screen data, control data, etc. of various screens related to the management of learning to the learning user terminal 30, executes various processes in response to the operation input by the user via the screen output on the learning user terminal 30, and transmits screen data, control data, etc. corresponding to the result of the process to the learning user terminal 30.
[0030] In the present embodiment, the learning management unit 114 is configured to execute various processes related to the control of the speech practice of sentences. For example, the learning management unit 114 is configured to acquire the speech voice (first speech voice) of the question sentence by the learning user. For example, the learning management unit 114 receives the speech voice data input via the microphone of the learning user terminal 30 from the learning user terminal 30.
[0031] In the present embodiment, the question sentence (first question sentence) has a group of sentences (first group of sentences) composed of a plurality of clauses, and the group of sentences includes a target clause (first target clause) that is the main target of the speech practice. A clause includes at least one of a word and a phrase. Also, the group of sentences may include a plurality of target clauses.
[0032] In addition, the learning management unit 114 is configured to evaluate the utterance of each phrase included in the question text based on the acquired speech audio of the utterance.
[0033] In addition, the learning management unit 114 is configured to evaluate the utterance of the entire question text based on the evaluation of the utterance of each phrase.
[0034] In addition, the learning management unit 114 is configured to update the speech ability value of the learning user for the question text based on the evaluation of the utterance of the entire question text, the evaluation of the utterance of the target phrase, or the speech ability value of the learning user for the target phrase updated based on the evaluation of the utterance of the target phrase. That is, the speech ability value for the question text is updated based on the evaluation of the current utterance of the entire question text and the evaluation of the current utterance of the target phrase, or the speech ability value for the target phrase (updated based on the evaluation). The speech ability value can also be said to be a parameter indicating the degree of speech ability. For example, it is managed in a table included in various tables 151. The speech ability value is set based on the evaluation of multiple utterances based on multiple speech audios from the past to the present.
[0035] In this way, the learning support server 10 in the present embodiment evaluates the utterance of each phrase included in the phrase group included in the question text by the learning user of the question text, evaluates the utterance of the entire question text based on the evaluation of the utterance of each phrase, and evaluates the utterance of the entire question text and the evaluation of the utterance of the target phrase, or the speech ability value for the target phrase updated based on the evaluation of the utterance of the target phrase. Based on this, the speech ability value for the question text is updated. Therefore, it is possible to support speech practice based on both the speech ability for the target phrase and the speech ability for the entire text. In this way, the learning support server 10 supports the improvement of the speech ability of the text.
[0036] In this embodiment, the learning management unit 114 acquires other utterance voices (second utterance voices) of a learning user for other question-presenting texts (second question-presenting texts) having other groups of sentences (second groups of sentences) including the same target sentence, evaluates the utterance of each sentence based on the other utterance voices, evaluates the utterance of the entire question-presenting text based on the evaluation of the utterance of each sentence, updates the utterance ability value for the target sentence based on the evaluation of the utterance of the target sentence, and is configured to update the utterance ability value of the learning user for the other question-presenting text based on the evaluation of the utterance of the entire other question-presenting text and the evaluation of the utterance of the target sentence, or the utterance ability value for the target sentence. That is, in this embodiment, it is possible to include a specific target sentence in a plurality of question-presenting texts. In this case, through the utterance practice of each question-presenting text (acquisition of utterance voices and evaluation of utterances, and update of utterance ability values based on the evaluation), the utterance ability value for the specific target sentence is repeatedly updated. Such a configuration enables efficient update of the utterance ability value for the target sentence through the utterance practice of a plurality of question-presenting texts.
[0037] Further, the learning management unit 114 may be configured to update the utterance ability value for the general sentence based on the evaluation of the utterance of the general sentence different from the target sentence among the groups of sentences included in the question-presenting text, and update the utterance ability value for the question-presenting text based on the utterance ability value for the general sentence. In this way, the group of sentences included in the question-presenting text may be composed of one or more general sentences in addition to one or more target sentences that are the main targets of the utterance practice. Such a configuration enables support for the utterance practice based on the utterance ability for the general sentence.
[0038] In this case, when updating the utterance ability value for the question-presenting text, the utterance ability value for the target sentence may have a greater influence on the utterance ability value for the question-presenting text than the utterance ability value for the general sentence. Such a configuration enables updating the utterance ability value for the question-presenting text so that the utterance ability value for the target sentence is prioritized.
[0039] In addition, the learning management unit 114 acquires other utterance voices (third utterance voices) of a learning user for other test questions (third test questions) having other target phrases (second target phrases) and other phrase groups (third phrase groups) including the same general phrase, evaluates the utterance of each phrase based on the other utterance voices, evaluates the utterance of the entire test question based on the evaluation of the utterance of each phrase, updates the utterance ability value for the other target phrase based on the evaluation of the utterance of the other target phrase, updates the utterance ability value for the general phrase based on the evaluation of the utterance of the general phrase, and updates the utterance ability value of the learning user for the other test question based on the evaluation of the utterance of the entire other test question, the evaluation of the utterance of the other target phrase, or the utterance ability value for the other target phrase, and the evaluation of the utterance of the general phrase, or the utterance ability value for the general phrase. That is, in the present embodiment, a specific general phrase can be included in a plurality of test questions. In this case, the utterance ability value for the specific general phrase is repeatedly updated through the utterance practice (acquisition of utterance voices and evaluation of utterances, and update of utterance ability values based on the evaluation) of each test question. Such a configuration enables efficient update of the utterance ability value for the general phrase through the utterance practice of a plurality of test questions.
[0040] In addition, when the number of favorably evaluated phrases with good utterance evaluations among each phrase is equal to or greater than a threshold value and the utterance evaluation of each target phrase is good, the learning management unit 114 sets the utterance evaluation of the entire test question as a positive first evaluation. On the other hand, when the number of favorably evaluated phrases is less than the threshold value or the utterance evaluation of at least one target phrase is not good, the learning management unit 114 sets the utterance evaluation of the entire test question as a negative second evaluation. Such a configuration enables the evaluation of the utterance of the entire test question based on the number of phrases with good utterance evaluations and the utterance evaluation of the target phrase.
[0041] In this embodiment, a time limit may be set for the speech of the learning user. For example, the learning management unit 114 is configured to acquire the speech audio uttered by the learning user within the time limit. For example, on the learning user terminal 30, when the elapsed time since the start of the input of the speech audio reaches a threshold value, the input of the speech audio automatically ends. In this case, the time limit may vary according to the length of the question text (for example, the longer the number of words in the question text, the longer the time). Such a configuration promotes smooth speech within the time limit by the learning user.
[0042] In this embodiment, before acquiring the speech audio of the question text, instruction information for instructing the speech of the question text may be provided to the learning user. For example, the learning management unit 114 transmits data including the instruction information to the learning user terminal 30, and the instruction information is displayed on the screen output on the learning user terminal 30. In this case, the instruction information may include a question text in which the target sentence is in an unrecognizable state (for example, blank). Such a configuration promotes the learning user to speak while extracting the target sentence from their own memory, and as a result, supports the acquisition of procedural knowledge for extracting the target sentence.
[0043] In this embodiment, the question texts for speech practice include various types of texts. For example, the question text includes a foreign language text. In this case, the instruction information may include a translation of the question text into the native language of the user (for example, a Japanese translation). Such a configuration enables the learning user to speak the foreign language question text while referring to the native language translation. Note that the words of the native language translation corresponding to the target sentence may be emphasized by making them bold or changing their color.
[0044] In this embodiment, an exam question passage may be selected from among a plurality of passages (for example, those managed in tables included in various tables 151). In this case, various rules may be applied as rules for selecting the exam question passage. For example, the learning management unit 114 is configured to select an exam question passage from among these plurality of passages based on the speaking ability value for each of the plurality of passages (for example, passages with lower speaking ability are preferentially selected). Also, for example, the learning management unit 114 is configured to select an exam question passage from among these plurality of passages based on the speaking ability value for the target phrase (for example, passages having target phrases with lower speaking ability are preferentially selected). Also, for example, the learning management unit 114 is configured to select an exam question passage from among the plurality of passages according to a predetermined learning plan.
[0045] In this embodiment, the difficulty level of the speaking exercise may be determined from among a plurality of difficulty levels. For example, the learning management unit 114 is configured to determine the difficulty level of the speaking exercise based on the speaking ability value for the exam question passage (for example, the higher the speaking ability, the higher the difficulty level). Also, for example, the learning management unit 114 is configured to determine the difficulty level of the speaking exercise based on the speaking ability value for the target phrase included in the exam question passage (for example, the higher the speaking ability, the higher the difficulty level).
[0046] In this embodiment, the mode of providing the instruction information for instructing the speaking of the exam question passage to the learning user may be changed according to the difficulty level of the speaking exercise. For example, when the difficulty level of the speaking exercise is the first difficulty level, a model voice of the exam question passage may be provided, while when the difficulty level is the second difficulty level which is higher than the first difficulty level, the model voice may not be provided. The provided model voice is reproduced (output) on the learning user terminal 30. Such a configuration enables control of the presence or absence of providing the model voice according to the difficulty level of the speaking exercise.
[0047] Also, for example, when the difficulty level of the speech practice is the first difficulty level, in response to an instruction from the learning user, the timing of the speech time limit starts. On the other hand, when the difficulty level is a second difficulty level that is higher than the first difficulty level, the timing of the time limit starts automatically (e.g., at the timing when a predetermined time has elapsed after providing the instruction information) regardless of the instruction from the learning user. Such a configuration enables controlling the timing at which the timing of the speech time limit starts according to the difficulty level of the speech practice.
[0048] Further, the learning management unit 114 may be configured to determine the progress of the speech practice for the entire plurality of sentences based on the speech ability values for each of the plurality of sentences. Also, the learning management unit 114 may be configured to determine the overall speech ability value of the learning user based on the speech ability values for each of the plurality of sentences. Such a configuration enables determining the progress of the speech practice and the overall speech ability value based on the speech ability values for each sentence.
[0049] Next, a specific example as an aspect of the learning support server 10 of the present embodiment having such a function will be described. First, in this example, the information managed by each table included in the various tables 151 will be described.
[0050] FIG. 3 illustrates the information managed by the learning user information table 1511 in this example. The learning user information table 1511 in this example manages information regarding the learning user, and as shown in the figure, in association with the "learning user ID" that identifies an individual learning user, it manages information such as "basic information" including the name, etc., and "available problem set information" which is information regarding the available problem sets. The available problem set information includes a problem set ID that identifies each of the plurality of available problem sets.
[0051] Figure 4 illustrates the information managed by the problem set information table 1512 in this example. The problem set information table 1512 in this example manages information regarding problem sets provided in the learning support service, and as illustrated, manages information such as "name" in association with a "problem set ID" that identifies an individual problem set.
[0052] The problem sets provided in the learning support service in this example include problem sets for the purpose of practicing speaking of foreign language sentences. For example, a problem set for practicing speaking of English sentences includes a plurality of English texts as problems. Also, problem sets for the purpose of practicing speaking of sentences to be memorized by speaking in the native language (e.g., sales talk, etc.) may be included other than for the purpose of practicing speaking of foreign language sentences.
[0053] Figure 5 illustrates the information managed by the sentence information table 1513 in this example. The sentence information table 1513 in this example manages information regarding foreign language sentences that are the subject of speaking practice, and as illustrated, manages information such as a "problem set ID" that identifies the problem set in which this sentence is included, "sentence content" that is the sentence itself, and "native language translation" that is the native language translation of this foreign language sentence, in association with a "sentence ID" that identifies an individual sentence.
[0054] Figure 6 illustrates the information managed by the in - sentence word information table 1514 in this example. The in - sentence word information table 1514 in this example manages information regarding words (phrases) included in the sentences that are the subject of speaking practice, and as illustrated, manages information such as a "word ID" that identifies a word, and a "target word flag" that indicates whether or not that word is a target word (target phrase) that is the main target of speaking practice, in association with a combination of a "sentence ID" that identifies an individual sentence and "word order" that is the word order in that sentence.
[0055] FIG. 7 illustrates the information managed by the word information table 1515 in this example. The word information table 1515 in this example manages information related to words, and as shown in the figure, manages information such as "word content" indicating the word itself in association with a "word ID" that identifies an individual word.
[0056] FIG. 8 illustrates the information managed by the sentence speaking ability value management table 1516 in this example. The sentence speaking ability value management table 1516 in this example manages information related to the speaking ability value of a learning user for sentences, and as shown in the figure, manages information such as "speaking ability value" in association with a combination of a "learning user ID" that identifies an individual learning user and a "sentence ID" that identifies an individual sentence. In this example, the speaking ability value is a numerical value from 0 to 100 points, and the larger the value, the higher the speaking ability.
[0057] FIG. 9 illustrates the information managed by the word speaking ability value management table 1517 in this example. The word speaking ability value management table 1517 in this example manages information related to the speaking ability value of a learning user for words, and as shown in the figure, manages information such as "speaking ability value" in association with a combination of a "learning user ID" that identifies an individual learning user, a "word ID" that identifies an individual word, and a "target word flag" that indicates whether the word is a target word. Thus, in this example, the speaking ability value for a word is managed separately for the case where the word is a target word, which is the main target of speaking practice, and the case where the word is a general word (general phrase) that is not a target word.
[0058] FIG. 10 illustrates the information managed by the learning plan information table 1518 in this example. The learning plan information table 1518 in this example manages information related to the learning plan set for a learning user, and as shown in the figure, manages information such as "learning period" in association with a combination of a "learning user ID" that identifies an individual learning user and a "problem set ID" that identifies an individual problem set. The problem set is selected from a plurality of problem sets available to the corresponding learning user.
[0059] In this example, the learning plan is set by each learning user. Specifically, through a screen (not shown) output on the learning user terminal 30, a problem set and the corresponding learning period (learning start date and learning end date) are set.
[0060] Above, in this example, the information managed by each table has been described. Next, in this example, the processes executed by the learning support server 10 and the screens output on the learning user terminal 30 and the like will be described.
[0061] FIG. 11 illustrates a learning top screen 50 output on the learning user terminal 30. As shown in the figure, the screen 50 has a problem set list display area 52 that displays a list of problem sets (identified by referring to the learning plan information table 1518) available (included in the learning period) at that time by the corresponding learning user, and a back button 54. In the problem set list display area 52, a plurality of individual display areas 521 each corresponding to an individual problem set are arranged side by side in the vertical direction.
[0062] As shown in the figure, each individual display area 521 displays the image and name of the corresponding problem set, and has a progress indicator 522 indicating the progress of the learning (including speech practice) of the problem set and a learning button 524. In this example, the progress of the speech practice is calculated based on the speech ability value (managed in the sentence speech ability value management table 1516) for each of the plurality of sentences included in the problem set. For example, the greater the speech ability value for each sentence, the more the progress advances.
[0063] The learning button 524 is an object for starting the learning of the corresponding problem set. FIG. 12 is a flowchart illustrating the processing executed by the learning support server 10 in response to the selection of the learning button 524 corresponding to a problem set for the purpose of practicing the pronunciation of foreign language sentences. First, as shown in the figure, the learning support server 10 determines an exam question sentence from among a plurality of sentences included in the corresponding problem set (step S100). For example, each of the plurality of sentences included in the corresponding problem set is determined as the exam question sentence according to a predetermined order. Note that the exam question sentence may be determined according to other rules. For example, the determination of the exam question sentence may be performed based on the speaking ability value. For example, a sentence whose speaking ability value does not exceed the threshold value may be determined as the exam question sentence.
[0064] Next, the learning support server 10 determines the difficulty level of the exercise (step S110). Specifically, the difficulty level of the exercise is determined based on the speaking ability value for the exam question sentence. In this example, the difficulty level has three levels: easy mode, normal mode, and difficult mode. For example, when the speaking ability value of the learning user for the exam question sentence is less than the first threshold value (for example, 20 points), the difficulty level is determined to be the easy mode; when the speaking ability value is equal to or greater than the first threshold value and less than the second threshold value (for example, 80 points), the difficulty level is determined to be the normal mode; and when the speaking ability value is equal to or greater than the second threshold value, the difficulty level is determined to be the difficult mode.
[0065] Subsequently, the learning support server 10 transmits exercise information to the learning user terminal 30 (step S120). Specifically, the information including the sentence content, the native language translation, and the determined difficulty level managed in the sentence information table 1513 is transmitted as the exercise information.
[0066] In the learning user terminal 30, a speaking exercise is performed based on the received exercise information. FIG. 13 illustrates an exercise screen 60 output on the learning user terminal 30 that has received the exercise information. As shown in the figure, the exercise screen 60 includes a native language translation display area 62 that displays the native language translation of the question text (in the example of FIG. 13, the Japanese translation), a question text display area 64 that displays the question text itself (in the example of FIG. 13, the English text), and a recording control button 66 with an icon of a microphone added thereto.
[0067] The exercise screen 60 illustrated in FIG. 13 corresponds to the case where the difficulty level of the exercise is in the easy mode. In this case, in the question text display area 64, the entire question text is displayed, and for two target words (in the example of FIG. 13, "translate" and "languages") among the words included in the question text, square brackets ([]) are added. Note that the number of target words included in the text may be one, or three or more. Also, the number of target words included in each text may be the same or different from each other.
[0068] Also, when the difficulty level of the exercise is in the easy mode, on the learning user terminal 30, after the output of the exercise screen 60, a model voice, which is a model speaking voice of the question text, is played (output). As illustrated in FIG. 13, while the model voice is being output, the text "Let's listen to the example" is displayed below the recording control button 66 of the exercise screen 60. This text is changed to the text "Tap to speak" when the playback of the model voice is completed. After the completion of the playback of the model voice, when the recording control button 66 is selected by the learning user, recording (input of the speaking voice via the microphone) is started on the learning user terminal 30. The learning user will perform the speaking of the question text while referring to the played model voice and looking at the entire question text displayed in the question text display area 64.
[0069] FIG. 14 illustrates the recording control button 66 during recording of the spoken voice. As shown, inside the contour line of the recording control button 66 during recording, a remaining time indicator 662 having a donut shape is arranged. The indicator 662 indicates the remaining time of the speaking limit time.
[0070] In this example, according to the number of words included in the question text, the speaking limit time changes. For example, the longer the number of words, the longer the limit time. When the remaining time of the limit time runs out, the recording on the learning user terminal 30 ends.
[0071] FIG. 15 illustrates the exercise screen 60 when the difficulty level of the exercise is in the normal mode. In the exercise screen 60 in this state, as shown, in the question text display area 64, only the general words excluding the target word among each word of the question text are displayed, the target word is not displayed (becomes blank), and square brackets are added at the position of the target word.
[0072] Also, when the difficulty level of the exercise is in the normal mode, after the output of the exercise screen 60, the playback of the model voice is not performed, and immediately below the recording control button 66, the text "Tap to speak" is displayed. When the learning user selects the recording control button 66, recording starts on the learning user terminal 30. The learning user will speak the question text while retrieving the target word from memory while looking at the question text in which the target word is blank and is displayed in the question text display area 64.
[0073] FIG. 16 illustrates the exercise screen 60 when the difficulty level of the exercise is in the difficult mode. In the exercise screen 60 in this state, as shown, in the question text display area 64, as in the case of the normal mode, only the general words excluding the target word among each word of the question text are displayed, the target word is not displayed, and square brackets are added at the position of the target word.
[0074] Also, when the difficulty level of the exercise is in the difficult mode, after the output of the exercise screen 60, the model voice is not played, and the text "3 seconds until recording starts" is displayed below the recording control button 66. In the difficult mode, after the output of the exercise screen 60, when a predetermined time (specifically 3 seconds) has elapsed, on the learning user terminal 30, recording is automatically started (recording will not start even if the recording control button 66 is selected). The learning user will speak the question text while retrieving the target word from memory while looking at the question text displayed in the question text display area 64 where the part of the target word is blank.
[0075] Returning to the flowchart of FIG. 12, when transmitting the exercise information, next, the learning support server 10 receives the spoken voice of the learning user recorded on the learning user terminal 30 from the learning user terminal 30 (step S130).
[0076] Subsequently, the learning support server 10 evaluates the speech of each phrase in the question text based on the received spoken voice, and transmits the evaluation to the learning user terminal 30 (step S140). Various methods can be applied to evaluate the speech of each phrase. In this example, through the speech recognition technology, the spoken text corresponding to the spoken voice is obtained, and through the comparison between the spoken text and the question text, the evaluation (pass or fail) of the speech of each word in the question text is performed. Note that when evaluating the speech of each word, linking (reason) may be considered. Also, the evaluation of the speech may be internally performed in multiple stages of 3 or more. When performing multiple-stage evaluation, an evaluation of a predetermined stage or higher is considered a pass, and an evaluation below the predetermined stage is considered a fail.
[0077] FIG. 17 illustrates an exercise screen 60 output on the learning user terminal 30 that has received the evaluation of the utterance of each phrase. In the problem statement display area 64 of the exercise screen 60 in this state, as shown in the figure, among the words included in the problem statement, for the words whose utterance evaluation is unacceptable (in the example of FIG. 17, "languages" and "English"), the word itself is highlighted, and below it, a pronunciation symbol object 641 with the pronunciation symbol of the corresponding word added is arranged.
[0078] Returning to the flowchart of FIG. 12, subsequently, the learning support server 10 evaluates the utterance of the entire problem statement (step S150). The evaluation of the utterance of the entire problem statement is performed by calculating an overall utterance evaluation value. In this example, the overall utterance evaluation value becomes 100 points when both of the following two conditions are satisfied, and becomes 0 points when at least one of them is not satisfied. Condition 1: Among the words included in the problem statement, the number of words whose utterance evaluation is acceptable is 80% or more of the total number of all words Condition 2: The utterance evaluation of all target words is acceptable For example, in the example of FIG. 17, since the number of words whose utterance evaluation is acceptable (13) is 80% or more of the total number of all words (15), Condition 1 is satisfied. However, since the utterance evaluation of "languages", which is one of the target words, is unacceptable, Condition 2 is not satisfied. As a result, the overall utterance evaluation value becomes 0 points. Regarding Condition 2, when two or more predetermined numbers of target words are consecutive in the problem statement, if the difference or ratio between the number of target words whose utterance evaluation is acceptable and the number of target words whose utterance evaluation is unacceptable among the predetermined number of target words exceeds a predetermined value determined in advance, it may be determined that Condition 2 is satisfied. At this time, it may be determined that Condition 2 is satisfied, and the utterance evaluation of all target words may be treated as acceptable, and the utterance ability value may be updated assuming that the target words that were unacceptable were also acceptable. Thereby, even in a situation where there is a high possibility of being acceptable even if there is an incorrect scoring due to insufficient recognition accuracy of speech recognition, it can be considered acceptable.
[0079] Next, the learning support server 10 updates the speech ability value of the learning user for the words (step S160). Specifically, among the words included in the question text, 10 points are added to the speech ability value for the words whose speech evaluation is qualified, and 5 points are subtracted from the speech ability value for the words whose speech evaluation is unqualified. As described above, the speech ability value for a word is managed separately depending on whether it is a target word or a general word. For example, in the example of FIG. 17, 10 points are added to the speech ability value as a general word for "asked" whose speech evaluation is qualified, 10 points are added to the speech ability value as a target word for "translate" which is also qualified, 5 points are subtracted from the speech ability value as a target word for "languages" whose speech evaluation is unqualified, and 5 points are subtracted from the speech ability value as a general word for "English" which is also unqualified. As described above, the speech ability value for a word is managed in the word speech ability value management table 1517.
[0080] Subsequently, the learning support server 10 updates the speech ability value of the learning user for the question text (step S170). In this example, the speech ability value for the text is calculated using the following formula. Text speech ability value = (the current overall speech evaluation value + the average value of the speech ability values of the general words included in the text + the average value of the speech ability values of the target words included in the text) / 3 The value calculated in this way is set as the updated text speech ability value. As described above, the speech ability value for the text is managed in the text speech ability value management table 1516. Note that instead of the speech ability values of the general words and the target words, the current speech evaluations of the general words and the target words may be used.
[0081] Subsequently, the learning support server 10 determines whether the speech ability value for the updated question text exceeds the threshold (step S180). If the speech ability value exceeds the threshold (YES in step S180), the speech drill for this text ends. After the end of the speech drill, another text from the same question set may be determined as the question text, or the learning user may be allowed to select the question text. On the other hand, if the speech ability value for the updated question text does not exceed the threshold (NO in step S180), the process returns to step S110, and the processes of steps S110 to S170 are repeated until the speech ability value for the question text exceeds the threshold (or until the learning user instructs to end).
[0082] In the above example, when the difficulty level of the drill is the easy mode, the speech restriction time may not be set (in this case, for example, the recording ends in response to the selection of the recording control button 66 by the learning user). Also, when the difficulty level of the drill is the difficult mode, the question text may be made non - visible in the question text display area 64 of the drill screen 60.
[0083] In the above example, the overall speech evaluation value is calculated based on the evaluation of the speech of each word included in the question text, but other criteria can also be applied. For example, based on the speech speed and pause frequency in the speech audio, the "fluency" of the speech may be evaluated, and the overall speech evaluation value may be calculated based on the "fluency".
[0084] In the above example, the question text may be automatically generated. For example, when one or more target words are specified, a text including these target words may be automatically generated. Such automatic generation of text can be realized, for example, using a large - scale language model.
[0085] In the above example, based on the speech ability values for the texts included in multiple question sets, the overall speech ability value of the learning user may be determined. Such an overall speech ability value is displayed, for example, on a screen such as the learning top screen 50.
[0086] The learning support server 10 according to the present embodiment described above evaluates the utterance of each word included in the word group of the question text based on the spoken voice of the learning user of the question text, and evaluates the utterance of the entire question text based on the evaluation of the utterance of each word. Based on the evaluation of the utterance of the entire question text and the evaluation of the utterance of the target word, or the utterance ability value for the target word updated based on the evaluation of the utterance of the target word, the utterance ability value for the question text is updated. Therefore, it is possible to support speaking practice based on both the speaking ability for the target word and the speaking ability for the entire text. In this way, the learning support server 10 supports the improvement of the speaking ability of the text.
[0087] In another embodiment of the present invention, some or all of the functions of the learning support server 10 in the above-described embodiment can be realized by the cooperation of the learning support server 10 and the learning user terminal 30, or can be realized by the learning user terminal 30. For example, some or all of the functions of the learning management unit 114 of the learning support server 10 can be realized by the learning user terminal 30. For example, the determination of the question text, the evaluation of the utterance based on the spoken voice, and the update of the utterance ability value may also be performed on the learning user terminal 30. That is, the system of the present invention can be configured by the learning support server 10, or can be configured by the learning support server 10 and the learning user terminal 30, or can be configured by the learning user terminal 30.
[0088] The processes and procedures described in this specification can be realized by software, hardware, or any combination thereof, in addition to those explicitly described. For example, the processes and procedures described in this specification can be realized by implementing the logic corresponding to the processes and procedures in a medium such as an integrated circuit, volatile memory, non-volatile memory, or magnetic disk. Also, the processes and procedures described in this specification can be implemented as a computer program corresponding to the processes and procedures and can be executed on various computers.
[0089] Even if it is described that the processes and procedures described in this specification are executed by a single device, software, component, or module, such processes or procedures can be executed by a plurality of devices, a plurality of software, a plurality of components, and / or a plurality of modules. Also, the software and hardware elements described in this specification can be realized by integrating them into fewer components or decomposing them into more components.
[0090] In this specification, when a component of the invention is described as either singular or plural, or is described without being limited to either singular or plural, unless the context indicates otherwise, the component can be either singular or plural.
Explanation of Signs
[0091] 10 Learning Support Server 11 Computer Processor 112 Management Function Control Unit 114 Learning Management Unit 15 Storage 151 Various Tables 1511 Learning User Information Table 1512 Problem Set Information Table 1513 Text Information Table 1514 Intra-Text Word Information Table 1515 Word Information Table 1516 Text Speech Ability Value Management Table 1517 Word Speech Ability Value Management Table 1518 Learning Plan Information Table 30 Learning User Terminal 40 Server-Side Program 44 Learning User Terminal-Side Program 50 Learning Top Screen 60 Drill Screen 66 Recording Control Button 662 Remaining Time Indicator
Claims
1. A program for supporting speech practice of an article, which causes one or more computers to acquire a first speech audio by a learning user of a first test article having a first group of sentences including a first target sentence; evaluate the speech of each sentence included in the first group of sentences based on the first speech audio; evaluate the speech of the entire first test article based on the first speech audio based on the evaluation of the speech of each sentence included in the first group of sentences; update the speech ability value of the learning user for the first test article based on the evaluation of the speech of the entire first test article based on the first speech audio, the evaluation of the speech of the first target sentence based on the first speech audio, or the speech ability value of the learning user for the first target sentence updated based on the evaluation of the speech of the first target sentence based on the first speech audio; program.
2. The one or more computers further acquire a second speech audio by the learning user of a second test article having a second group of sentences including the first target sentence; evaluate the speech of each sentence included in the second group of sentences based on the second speech audio; evaluate the speech of the entire second test article based on the second speech audio based on the evaluation of the speech of each sentence included in the second group of sentences; update the speech ability value of the learning user for the first target sentence based on the evaluation of the speech of the first target sentence based on the second speech audio; update the speech ability value of the learning user for the second test article based on the evaluation of the speech of the entire second test article based on the second speech audio, the evaluation of the speech of the first target sentence based on the second speech audio, or the speech ability value for the first target sentence; The program according to Claim 1.
3. The one or more computers further execute a step of updating the speech ability value for the general sentence based on the evaluation of the speech based on the first speech audio of a general sentence different from the first target sentence among the first group of sentences, The step of updating the speech ability value for the first test article updates the speech ability value for the first test article further based on the evaluation of the speech of the general sentence based on the first speech audio or the speech ability value for the general sentence. The program according to Claim 1.
4. The step of updating the speaking ability value for the first question sentence is such that the speaking ability value for the first target sentence has a greater influence on the speaking ability value for the first question sentence than the speaking ability value for the general sentence. The program of claim 3.
5. The one or more computers are further caused to obtain a third spoken voice of the learning user for a third question sentence having a third sentence group including a second target sentence and the general sentence; evaluate the speaking of each sentence included in the third sentence group based on the third spoken voice; evaluate the speaking of the entire third question sentence based on the third spoken voice based on the evaluation of the speaking of each sentence included in the third sentence group; update the speaking ability value of the learning user for the second target sentence based on the evaluation of the speaking of the second target sentence based on the third spoken voice; update the speaking ability value of the learning user for the general sentence based on the evaluation of the speaking of the general sentence based on the third spoken voice; update the speaking ability value of the learning user for the third question sentence based on the evaluation of the speaking of the entire third question sentence based on the third spoken voice, the evaluation of the speaking of the second target sentence based on the third spoken voice, or the speaking ability value for the second target sentence, and the evaluation of the speaking of the general sentence based on the third spoken voice, or the speaking ability value for the general sentence. The program of claim 3.
6. The step of evaluating the speaking of the entire first question sentence based on the first spoken voice is such that when the number of well-evaluated sentences with good speaking evaluations among each sentence included in the first sentence group is equal to or greater than a threshold value, and the speaking evaluation of the first target sentence is good, the evaluation of the speaking of the entire first question sentence based on the first spoken voice is a positive first evaluation. On the other hand, when the number of well-evaluated sentences among each sentence included in the first sentence group is less than the threshold value, or the speaking evaluation of the first target sentence is not good, the evaluation of the speaking of the first question sentence based on the first spoken voice is a negative second evaluation. The program of claim 1.
7. The step of obtaining the first spoken voice is to obtain the first spoken voice made by the learning user within a limited time. The program of claim 1.
8. The limited time varies according to the length of the first question sentence. The program according to claim 1.
9. Cause the one or more computers to further execute a step of providing the learning user with instruction information for instructing the utterance of the first question sentence before acquiring the first uttered voice, The instruction information includes the first question sentence in a state where the first target phrase is not visible, The program according to claim 1.
10. Cause the one or more computers to further execute a step of providing the learning user with instruction information for instructing the utterance of the first question sentence before acquiring the first uttered voice, The first question sentence is a sentence in a foreign language, The instruction information includes a translation of the first question sentence into the native language of the first question sentence, The program according to claim 1.
11. Cause the one or more computers to further execute a step of selecting the first question sentence from among the plurality of sentences based on the utterance ability values for each of the plurality of sentences, The program according to claim 1.
12. Cause the one or more computers to further determine, based on the utterance ability value for the first question sentence, the difficulty level of the utterance practice of the first question sentence from among a plurality of difficulty levels, The program according to claim 1.
13. Cause the one or more computers to further execute a step of determining the difficulty level of the utterance practice of the first question sentence from among a plurality of difficulty levels and a step of providing the learning user with instruction information for instructing the utterance of the first question sentence before acquiring the first uttered voice, In the step of providing the instruction information, when the difficulty level is the first difficulty level, a model voice of the first question sentence is provided, while when the difficulty level is a second difficulty level higher than the first difficulty level, the model voice of the first question sentence is not provided, The program according to claim 1.
14. Cause the one or more computers to further execute a step of determining the difficulty level of the utterance practice of the first question sentence from among a plurality of difficulty levels and a step of providing the learning user with instruction information for instructing the utterance of the first question sentence before acquiring the first uttered voice, The step of acquiring the first uttered voice acquires the first uttered voice made by the learning user within a limited time, When the time limit is being measured, when the difficulty level is the first difficulty level, it is started in response to an instruction from the learning user, while when the difficulty level is a second difficulty level that is higher in difficulty than the first difficulty level, it is automatically started regardless of an instruction from the learning user. The program according to claim 1.
15. Causing the one or more computers to further execute a step of determining a progress degree of the speech practice for the entire plurality of sentences based on the speech ability value for each of the plurality of sentences. The program according to claim 1.
16. Causing the one or more computers to further execute a step of determining a comprehensive speech ability value of the learning user based on the speech ability value for each of the plurality of sentences. The program according to claim 1.
17. A system including one or more computer processors for assisting in speech practice of sentences, wherein the one or more computer processors A step of obtaining a first speech audio of a learning user for a first question sentence having a first sentence group including a first target phrase; A step of evaluating the speech of each phrase included in the first sentence group based on the first speech audio; A step of evaluating the speech of the entire first question sentence based on the first speech audio based on the evaluation of the speech of each phrase included in the first sentence group; Updating the speech ability value of the learning user for the first question sentence based on the evaluation of the speech of the entire first question sentence based on the first speech audio and the evaluation of the speech of the first target phrase based on the first speech audio, or the speech ability value of the learning user for the first target phrase updated based on the evaluation of the speech of the first target phrase based on the first speech audio. System.
18. A method executed by one or more computers for assisting in speech practice of sentences, comprising: A step of obtaining a first speech audio of a learning user for a first question sentence having a first sentence group including a first target phrase; A step of evaluating the speech of each phrase included in the first sentence group based on the first speech audio; A step of evaluating the speech of the entire first question sentence based on the first speech audio based on the evaluation of the speech of each phrase included in the first sentence group; Updating the speaking ability value of the learning user for the first question sentence based on the evaluation of the speaking of the entire first question sentence based on the first spoken voice, the evaluation of the speaking of the first target phrase based on the first spoken voice, or the speaking ability value of the learning user for the first target phrase updated based on the evaluation of the speaking of the first target phrase based on the first spoken voice. Method.
Citation Information
Patent Citations
Device and method for studying foreign language, and medium therefor
JP2001265211A
Device and method of assisting language learning
JP2004258231A
Speech evaluation device, speech evaluation method, method for producing teacher change information, and program
JP2017198790A
Method for reinforcing internalization of foreign language sentence pattern according to consecutive and simultaneous interpretation tests using voice recognition engine and providing language training service
KR1020140004533A
Method and apparatus for analysing sentence stress
KR1020160146356A