Programs, devices, and methods for assisting users with mouth and / or tongue exercises.

The program and device enhance mouth and tongue exercises by using speech recognition and interactive feedback to improve oral frailty through engaging and accurate exercise assistance.

JP2026074706AActive Publication Date: 2026-05-07CAT CORP
View PDF 8 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
CAT CORP
Filing Date
2024-10-21
Publication Date
2026-05-07

AI Technical Summary

Technical Problem

Conventional mouth and tongue exercises for improving oral frailty are monotonous and lack feedback on correct performance, especially for the elderly, who may struggle to execute these exercises effectively.

Method used

A program and device that utilize speech recognition technology to assist users in performing mouth and tongue exercises by recognizing monosyllabic words, adjusting distance and volume, and providing game-like feedback to ensure correct execution.

Benefits of technology

Enhances user engagement and accuracy in mouth and tongue exercises by providing real-time feedback and game-like interaction, ensuring correct performance and improving oral frailty through interactive and engaging methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026074706000001_ABST
    Figure 2026074706000001_ABST
Patent Text Reader

Abstract

To provide a program to assist users with mouth and / or tongue exercises. [Solution] When the program of the present invention is executed by the processor unit of the device, it causes the processor unit to at least receive the user's voice input, recognize at least one monosyllabic word contained in the user's voice input, and display content associated with the recognized at least one monosyllabic word on the device. In one embodiment, recognizing at least one monosyllabic word may include recognizing at least one monosyllabic word using a speech recognition model suitable for recognizing multiple different monosyllabic words spoken in succession in a short period of time and / or the same monosyllabic word spoken multiple times in succession in a short period of time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a program, a device, and a method for assisting gymnastics of a user's mouth and / or tongue.

Background Art

[0002] Conventionally, the decline around the mouth (so-called oral frailty) has been attracting attention (for example, see Non-Patent Document 1).

Prior Art Documents

Non-Patent Documents

[0003]

Non-Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] Although various gymnastics of the mouth and / or tongue for improving oral frailty are known, the conventional gymnastics are only monotonous and boring (that is, do not last long), and it is not known whether the elderly can perform such gymnastics correctly when they perform them.

[0005] An object of the present invention is to provide a program, a device, and a method for assisting gymnastics of a user's mouth and / or tongue, which enable it to be known from the content displayed when the user performs gymnastics of the mouth and / or tongue in a game-like feeling that the user can perform the gymnastics of the mouth and / or tongue correctly.

Means for Solving the Problems

[0006] In one aspect of the present invention, the program of the present invention is a program for assisting a user's mouth and / or tongue exercises, and when executed by the processor unit of a device, the program causes the processor unit to at least receive the user's voice input, recognize at least one monosyllabic word contained in the user's voice input, and display content associated with the recognized at least one monosyllabic word on the device.

[0007] In one embodiment of the present invention, recognizing the at least one monosyllabic word may include recognizing the at least one monosyllabic word using a speech recognition model suitable for recognizing multiple different monosyllabic words spoken in succession in a short period of time and / or the same monosyllabic word spoken multiple times in succession in a short period of time.

[0008] In one embodiment of the present invention, the training data for the speech recognition model may be a correspondence between monosyllabic words and the speech data of those monosyllabic words.

[0009] In one embodiment of the present invention, when the program is executed by the processor unit of the device, the processor unit may further cause the device to display an indicator for adjusting a first distance between the user's face and the device.

[0010] In one embodiment of the present invention, the indicator may include a display area for displaying the user as captured by the camera of the device, a bar representing the difference between the first distance and a predetermined appropriate distance, a numerical value representing the difference between the first distance and the predetermined appropriate distance, or guide lines for guiding the position of the user's face as displayed on the device by being captured by the camera of the device.

[0011] In one embodiment of the present invention, the size of the display area may be adjusted such that when the user's face within the display area is of a desired size, the distance between the user's face and the device is within a desired range.

[0012] In one embodiment of the present invention, when the program is executed by the processor unit of the device, the program may further cause the processor unit to determine the volume of the user's voice input, determine whether the volume of the user's voice input exceeds a first threshold, and, if it is determined that the volume of the user's voice input exceeds the first threshold, to increase the magnification of the user's face displayed in the display area, and / or, when the program is executed by the processor unit of the device, the program may further cause the processor unit to determine whether the volume of the user's voice input falls below a second threshold, and, if it is determined that the volume of the user's voice input falls below the second threshold, to decrease the magnification of the user's face displayed in the display area.

[0013] In one embodiment of the present invention, when the program is executed by the processor unit of the device, the program may further cause the processor unit to: identify the volume of the user's voice input; determine whether the volume of the user's voice input exceeds a first threshold; and, if it is determined that the volume of the user's voice input exceeds the first threshold, to perform a process to cause the user to lower the volume of the user's voice input; and / or, when the program is executed by the processor unit of the device, the program may further cause the processor unit to: determine whether the volume of the user's voice input falls below a second threshold; and, if it is determined that the volume of the user's voice input falls below the second threshold, to perform a process to cause the user to raise the volume of the user's voice input.

[0014] In one aspect of the present invention, the device of the present invention is a device for assisting a user in mouth and / or tongue exercises, the device comprising a processor unit, the processor unit configured to at least receive the user's voice input, recognize at least one monosyllabic word contained in the user's voice input, and display content associated with the recognized at least one monosyllabic word on the device.

[0015] In one aspect of the present invention, the method of the present invention is a method performed in a device for assisting a user's mouth and / or tongue exercises, the device comprising a processor unit, the method comprising the processor unit receiving a user's voice input, the processor unit recognizing at least one monosyllabic word contained in the user's voice input, and the processor unit displaying content associated with the recognized at least one monosyllabic word on the device. [Brief explanation of the drawing]

[0016] [Figure 1A] A diagram showing an example of a screen displayed on a device operated by a user. [Figure 1B] This diagram shows another example of a screen displayed on a device operated by a user. [Figure 2] This diagram shows another example of a screen displayed on a device operated by a user. [Figure 3] A diagram showing an example configuration of device 300 for assisting the user with mouth and / or tongue exercises. [Figure 4] This figure shows an example of the processing performed in device 300. [Figure 5] This figure shows another example of the processing performed on device 300. [Modes for carrying out the invention]

[0017] The following definitions are used in this specification.

[0018] A "monosyllabic word" refers to a word having a single syllable consisting of a single vowel, or a word having a single syllable consisting of a combination of a single consonant and a single vowel. "Monosyllabic words" include, for example, but are not limited to, "pa", "ta", "ka", "ra", etc.

[0019] "Exercises for the mouth and / or tongue" refer to exercises for stimulating the muscles necessary for moving the mouth and / or the muscles necessary for moving the tongue. Therefore, "exercises for the mouth and / or tongue" include both exercises performed with vocalization and exercises performed without vocalization.

[0020] Hereinafter, embodiments of the present invention will be described while referring to the drawings.

[0021] 1. Assisting the user with mouth and / or tongue exercises. The applicant considered that in order to improve oral frailty, it is necessary to integrate the examination and training of oral frailty. Also, in view of the fact that the elderly have been interested in games in recent years and it is expected that the number of elderly people who can use smartphones in the future will increase, the applicant considered that it is best to attempt to improve oral frailty in game applications that can be launched on elderly devices.

[0022] One exercise that contributes to improving oral frailty is to repeatedly pronounce a specific monosyllabic word in quick succession. However, conventional speech recognition models have been unable to recognize that a specific monosyllabic word has been pronounced multiple times in quick succession, even when the user pronounces it multiple times in quick succession. This is because conventional speech recognition models cannot predict that the same monosyllabic word will be received as speech input multiple times in quick succession. For example, if a user pronounces the specific monosyllabic word "pa" as "papapapapa" in quick succession, a conventional speech recognition model, after receiving the initial speech input of "papa," cannot predict that "pa" will be pronounced again immediately after "papa" has been pronounced in quick succession, resulting in a failure and an inability to recognize what is being input. Therefore, the applicant has developed a new speech recognition model suitable for recognizing multiple monosyllabic words pronounced in quick succession and the same monosyllabic word pronounced multiple times in quick succession.

[0023] Furthermore, the applicant found that stabilizing the volume of the user's voice input is also important in recognizing multiple monosyllabic words spoken in quick succession and the same monosyllabic word spoken multiple times in quick succession. If the volume of the user's voice input is unstable, even if the user speaks a monosyllabic word multiple times in quick succession (e.g., "pa-pa-pa-pa"), the device will only recognize that the user has spoken the monosyllabic word once (e.g., "pa"). This is because, for example, if the volume of the user's voice input continues to exceed a predetermined threshold, a monosyllabic word spoken multiple times in quick succession (e.g., "pa-pa-pa-pa") will be recognized as a single long monosyllabic word (e.g., "paaaa"), and a single long monosyllabic word (e.g., "paaaa") will be considered a single short monosyllabic word (e.g., "pa") as a monosyllabic voice input. Therefore, stabilizing the volume of the user's voice input is necessary for the device to more accurately recognize the monosyllabic words spoken by the user.

[0024] 2. Screen displayed on a device operated by the user. Figure 1A shows an example of a screen displayed on a user-operated device. Screen 110 shown in Figure 1A is a game screen that operates in response to the user's voice input to assist the user with mouth and / or tongue exercises. Screen 110 may appear on the device when a program to assist the user with mouth and / or tongue exercises is launched on the device and a game for mouth and / or tongue exercises is started.

[0025] In the embodiment shown in Figure 1A, the screen 110 includes a display area 111 for displaying an object being captured by the device's camera (e.g., the user). The display area 111 may display video of the object being captured in real time, or it may display a still image of the object being captured. This allows the user to check their own mouth area displayed in the display area 111 (i.e., self-monitoring) and see whether they are moving their mouth properly. The display area 111 can also display an illustration or animation of the user or a specific person's portrait. For example, the portrait animation may be an animation that moves in accordance with the movement of the object being captured in real time by the camera (e.g., the user's mouth movements).

[0026] Furthermore, the size of the display area 111 is adjusted so that the distance between the user's face and the device is within a desired range when the user's face within the display area is of a desired size (for example, when the user's forehead and chin are positioned near the outline of the display area 111, or when the ratio of the area of ​​the user's face within the display area to the total area of ​​the display area is within a predetermined range). By setting the distance between the user's face and the device within a desired range, the volume of the user's voice input is stabilized, and it becomes possible to easily recognize the difference in volume between one monosyllabic word and another monosyllabic word included in the user's voice input. As a result, the device can distinguish and recognize each monosyllabic word included in the user's voice input, and therefore can more accurately recognize multiple monosyllabic words spoken in quick succession and the same monosyllabic word spoken multiple times in quick succession. The desired range is, for example, a range suitable for stabilizing the volume of the user's voice input. Furthermore, "short time" refers to, for example, less than 0.1 seconds, less than 0.5 seconds, less than 1 second, less than 2 seconds, less than 3 seconds, less than 4 seconds, less than 5 seconds, less than 6 seconds, less than 7 seconds, less than 8 seconds, less than 9 seconds, less than 10 seconds, etc., but is not limited to these.

[0027] Figure 1B shows another example of a screen displayed on a user-operated device. Screen 120 in Figure 1B is the screen that transitioned from screen 110 in Figure 1A in response to the user's voice input being input and recognized.

[0028] As shown in screen 120, when one of the one or more monosyllabic words included in the user's voice input is recognized, content 121 associated with the recognized monosyllabic word is displayed. This allows the user to recognize that they are pronouncing the monosyllabic word correctly by the display of content 121 associated with the recognized monosyllabic word. Monosyllabic words include, for example, words produced by closing and releasing the lips to produce a plosive sound, words produced by bringing the tip of the tongue close to or in contact with the front of the palate, words produced by bringing the base of the tongue close to or in contact with the back of the palate, and / or words produced by pressing the tongue against the palate with the tongue curled so that the tip of the tongue is toward the base of the tongue, and using the entire tongue. Words produced by closing and releasing the lips to produce a plosive sound may include, for example, words consisting of sounds such as the "m", "b", and "p" sounds. Words produced by bringing the tip of the tongue close to or in contact with the front of the palate may include, for example, words consisting of sounds such as the "t", "z", and "d" sounds. Words produced by bringing the base of the tongue close to or in contact with the back of the palate may include, for example, words consisting of sounds from the "k" row and the "g" row. Words produced by curling the tongue so that the tip of the tongue is towards the base of the tongue and pressing the tongue against the palate, using the entire tongue, may include, for example, words consisting of sounds from the "r" row. Monosyllabic words may include words produced by opening the mouth vertically, words produced by closing the teeth, words produced by pursing the lips, and words produced by opening the mouth horizontally. Words produced by opening the mouth vertically may include, for example, words consisting of sounds from the "a" row. Words produced by closing the teeth may include, for example, words consisting of sounds from the "i" row. Words produced by pursing the lips may include, for example, words consisting of sounds from the "u" row and the "o" row. Words produced by opening the mouth horizontally may include, for example, words consisting of sounds from the "e" row. In the embodiment shown in Figure 1B, when the user's voice input is recognized as containing the monosyllabic word "pa," the content "pa" (in katakana) corresponding to the recognized monosyllabic word "pa" is displayed on screen 110.

[0029] The displayed content 121 moves toward the right side of the screen and may, for example, collide with an item moving from the right side of the screen and disappear. This game-like element allows users to enjoy exercising their mouth and / or tongue.

[0030] Furthermore, if a specific monosyllabic word included in the user's voice input is recognized, content 121 associated with that specific monosyllabic word may be displayed on screen 110, and if a monosyllabic word other than that specific one is recognized, content 122 associated with the other monosyllabic word may be displayed on screen 110. This allows the user to recognize that they are pronouncing the specific monosyllabic word correctly by the display of content 121 associated with that specific monosyllabic word, and to recognize that they are not pronouncing the specific monosyllabic word correctly by the display of content 122 associated with the other monosyllabic word.

[0031] Furthermore, multiple speech modes may exist, and the specific monosyllabic word may differ for each speech mode. For example, if the monosyllabic word "pa" is the specific monosyllabic word in the first speech mode, then when the user pronounces "pa," content 121 will be displayed, and when the user pronounces "ta," content 122 will be displayed. Similarly, if the monosyllabic word "ta" is the specific monosyllabic word in a second speech mode different from the first speech mode, then when the user pronounces "pa," content 122 will be displayed, and when the user pronounces "ta," content 121 will be displayed.

[0032] In the embodiment shown in Figure 1B, the katakana character "パ" was described as the content 121 corresponding to the recognized specific monosyllabic word "パ," but the present invention is not limited to this. The content 121 corresponding to the recognized specific monosyllabic word "パ" may be, for example, the hiragana character "ぱ," the alphabet character "Pa" including uppercase letters, or the alphabet character "pa" with only lowercase letters.

[0033] In the embodiments shown in Figures 1A and 1B, examples were described in which screens 110 and 120 have a display area 111, but the present invention is not limited thereto. Screens 110 and 120 may have a complete illustration or animation of a portrait, or a character, instead of the display area 111.

[0034] Furthermore, although the embodiments shown in Figures 1A and 1B describe a display area 111 having a perfect circle, the present invention is not limited thereto. The display area 111 can have any shape as long as it can display the object being imaged by the camera. For example, the shape of the display area 111 may be elliptical, polygonal such as a square or rectangle, or complex shapes such as a star or gourd.

[0035] Furthermore, the backgrounds and items shown on screens 110 and 120 can be changed to alternative backgrounds and items depending on the situation or concept of performing mouth and / or tongue exercises.

[0036] Figure 2 shows another example of a screen displayed on a user-operated device. Screen 210 shown in Figure 2 is a game screen that operates in response to the user's voice input to assist the user with mouth and / or tongue exercises. Screen 210 may appear on the device when a program to assist the user with mouth and / or tongue exercises is launched on the device and a game for mouth and / or tongue exercises is started.

[0037] In the embodiment shown in Figure 2, the screen 210 includes a display area 211 for displaying an object being captured by the device's camera. The display area 211 in Figure 2 has the same configuration as the display area 111 in Figure 1A, so a detailed explanation is omitted here.

[0038] As shown in screen 210, when a monosyllabic word is recognized from among several monosyllabic words included in the user's voice input, content 212 associated with the recognized monosyllabic word is displayed. In the embodiment shown in Figure 2, when the user's voice input is recognized to include the monosyllabic word "pa", a numerical value related to the recognized monosyllabic word "pa" is displayed as content 212 associated with the recognized monosyllabic word. The more times the monosyllabic word "pa" is recognized, the more the numerical value associated with the recognized monosyllabic word "pa" may change. This allows the user to recognize that they are pronouncing the monosyllabic word correctly by the display and / or change of content 211 associated with the recognized monosyllabic word. The numerical value associated with the recognized monosyllabic word may be, for example, the count of recognized monosyllabic words, or a score representing the accuracy indicating the degree to which the recognized monosyllabic word matches a specific monosyllabic word.

[0039] Furthermore, if a specific monosyllabic word included in the user's voice input is recognized, a numerical value associated with that specific monosyllabic word will be reflected on screen 210. However, if a monosyllabic word other than that specific one is recognized, the numerical value associated with that other monosyllabic word does not need to be reflected on screen 210. For example, if a specific monosyllabic word included in the user's voice input is recognized, the numerical value associated with that specific monosyllabic word will increase on screen 210. If a monosyllabic word other than that specific one is recognized, the numerical value associated with that other monosyllabic word will not be reflected on screen 210, and the numerical value associated with that specific monosyllabic word will not change on screen 210. This allows the user to recognize that they are correctly pronouncing a specific monosyllabic word by seeing the numerical value associated with that specific monosyllabic word reflected on screen 210, and to recognize that they are not correctly pronouncing a specific monosyllabic word by seeing the numerical value associated with that specific monosyllabic word not reflected on screen 210.

[0040] Although the embodiment shown in Figure 2 describes an example in which the screen 210 has a display area 211, the present invention is not limited thereto. The screen 210 may have a complete illustration or animation of a portrait, or a character, instead of the display area 211.

[0041] Furthermore, although the embodiment shown in Figure 2 describes a display area 211 having a perfect circle, the present invention is not limited thereto. The display area 211 can have any shape as long as it can display the object being imaged by the camera. For example, the shape of the display area 211 may be elliptical, polygonal such as a square or rectangle, or complex shapes such as a star or gourd.

[0042] Furthermore, the background and items shown on screen 210 can be changed to alternative backgrounds and items depending on the situation or concept of performing mouth and / or tongue exercises.

[0043] 3. Configuration of a device to assist the user with mouth and / or tongue exercises. Figure 3 shows an example configuration of device 300 for assisting the user with mouth and / or tongue exercises.

[0044] Device 300 is an electronic device operated by a user. In the embodiment shown in Figure 3, device 300 includes an interface unit 311, a processor unit 312 including one or more CPUs (Central Processing Units), a memory unit 313, an input unit 314, a display unit 315, and a camera 316. Device 300 may be able to communicate via a network with a server device (not shown) for recording and managing logs of mouth and / or tongue exercises performed by the user. This makes it possible to perform an examination of the user's mouth and / or tongue based on the logs of the mouth and / or tongue exercises performed by the user. For example, device 300 may be a portable wireless terminal such as a mobile phone, smartphone, or tablet terminal, or a personal computer such as a laptop PC or notebook PC.

[0045] The interface unit 311 controls communication with, for example, a server device (not shown).

[0046] The memory unit 313 stores the program required to execute the process and the data required to execute that program. The method by which the program is stored in the memory unit 313 is not specified. For example, the program may be pre-installed in the memory unit 313. Alternatively, the program may be installed in the memory unit 313 by being downloaded via a network, or it may be installed in the memory unit 313 via a storage medium such as an optical disc or USB drive.

[0047] The processor unit 312 controls the operation of the entire device 300. The processor unit 312 reads the program stored in the memory unit 313 and executes the program. As a result, the device 300 can function as a device that performs desired steps, and the processor unit 212 of the device 300 can operate as a means to achieve the desired function.

[0048] The input unit 314 is capable of receiving voice input from the user and can also receive input corresponding to the user's selected actions (e.g., click, tap).

[0049] The display unit 315 can display, for example, screen 110 in Figure 1A, screen 120 in Figure 1B, screen 210 in Figure 2, and so on.

[0050] Camera 316 is capable of capturing images of the object to be imaged. The object captured by camera 316 can be displayed in the display area 111 on screen 110 in Figure 1A and screen 120 in Figure 1B, and in the display area 211 on screen 210 in Figure 2.

[0051] 4. Processes performed on the device Figure 4 shows an example of the processing performed in device 300. Each step shown in Figure 4 is performed, for example, by the processor unit 312 of device 300. It is assumed that the program to assist the user's mouth and / or tongue exercises is pre-stored in the memory unit 313 of device 300. The steps shown in Figure 4 will now be explained.

[0052] Step S401: An indicator for adjusting a first distance between the user's face and the device 300 is displayed on the device 300 (in particular, on the display unit 315 of the device 300). This process is performed, for example, in response to the activation of a program on the device 300 to assist the user in mouth and / or tongue exercises. The indicator includes, for example, a display area for displaying an object being imaged by the camera 316 of the device 300, a bar representing the difference between the first distance between the user's face and the device 300 and a predetermined appropriate distance, a numerical value representing the difference between the first distance between the user's face and the device 300 and a predetermined appropriate distance, or guide lines for guiding the position of the user's face as it is imaged by the camera 316 of the device 300 and displayed on the device 300. This display area may be, for example, the display area 111 displayed on screen 110 in Figure 1A and screen 120 in Figure 1B, or the display area 211 displayed on screen 210 in Figure 2. This bar may be, for example, a bar with a slider that indicates a first distance between the user's face and the device 300 (or the difference between the first distance between the user's face and the device 300 and a predetermined appropriate distance), or it may be a progress bar. This guide line may be, for example, a line that indicates the approximate width of the user's face.

[0053] The size of the display area is adjusted so that the distance between the user's face and the device is within a desired range when the user's face within the display area is of a desired size (for example, when the user's forehead and chin are positioned near the outline of the display area, or when the ratio of the area of ​​the user's face within the display area to the total area of ​​the display area is within a predetermined range). This makes it possible to stabilize the distance between the user's face and the device 300, and therefore, to stabilize the volume of the user's voice input.

[0054] Furthermore, multiple indicators for adjusting the first distance between the user's face and the device 300 may be displayed on the device 300 simultaneously. For example, a display area and guide lines may be displayed on the device 300 simultaneously, or a display area and a bar may be displayed on the device 300 simultaneously. Also, if no indicators are displayed on the device 300 (for example, if a complete illustration or animation of a portrait is displayed, or if a character is displayed), step S401 is skipped.

[0055] Step S402: User voice input is received. The user voice input is received, for example, via the input unit 314 of device 300. The user voice input includes at least one monosyllabic word. The at least one monosyllabic word included in the user voice input may include, for example, several different monosyllabic words uttered in succession (e.g., in a short period of time). The at least one monosyllabic word included in the user voice input may include, for example, the same monosyllabic word uttered multiple times in succession (e.g., in a short period of time). The at least one monosyllabic word may include a word produced by closing and then releasing the lips to produce a plosive sound, a sound produced by bringing the tip of the tongue close to or in contact with the front of the palate, a sound produced by bringing the base of the tongue close to or in contact with the back of the palate, and / or a sound produced by pressing the tongue against the palate with the tongue curled so that the tip of the tongue is toward the base of the tongue, using the entire tongue. At least one monosyllabic word is one that is pronounced by opening the mouth vertically, closing the teeth, pursing the lips, or opening the mouth horizontally.

[0056] Step S403: At least one monosyllabic word included in the user's voice input received in step S402 is recognized. This process is performed using a speech recognition model suitable for recognizing, for example, multiple different monosyllabic words spoken in succession in a short period of time and / or the same monosyllabic word spoken multiple times in succession in a short period of time. The speech recognition model may be configured to be machine-learnable and may, for example, be artificial intelligence (AI). The speech recognition model may be stored, for example, in the memory section 313 of device 300, or in storage on a cloud connected via a network. If the speech recognition model is stored in storage on a cloud connected via a network, device 300 can access the speech recognition model via the network.

[0057] The training data for a speech recognition model may include, for example, the correspondence between monosyllabic words and the speech data of those monosyllabic words. Alternatively, the training data for a speech recognition model may consist only of the correspondence between monosyllabic words and the speech data of those monosyllabic words. With such training data, device 300 can recognize monosyllabic words that correspond to speech data of monosyllabic words identical or similar to the received speech data of a monosyllabic word as monosyllabic words spoken by the user. By limiting the training data for a speech recognition model to monosyllabic words in this way, it is possible to improve the recognition ability of multiple monosyllabic words spoken consecutively in a short period of time and / or the same monosyllabic word spoken multiple times consecutively in a short period of time (in particular, the same monosyllabic word spoken multiple times consecutively in a short period of time).

[0058] Step S404: Content associated with at least one monosyllabic word recognized in step S403 is displayed. This content may include, for example, content 121 displayed on screen 120 in Figure 1B, or content 212 displayed on screen 210 in Figure 2. This content may further include, for example, content 122 displayed on screen 120 in Figure 1B.

[0059] Figure 5 shows another example of processing performed in device 300. Each step shown in Figure 5 is performed, for example, by the processor unit 312 of device 300. Each step shown in Figure 5 is performed after receiving the user's voice input in step S402 of Figure 4. The steps shown in Figure 5 will be described below.

[0060] Step S501: The volume of the user's voice input received in step S402 of Figure 4 is identified. For example, the volume of the user's voice input is identified by referring to the volume parameter included in the user's voice input.

[0061] Step S502: It is determined whether the volume of the user's voice input exceeds the first threshold. If the result is "Yes", the process proceeds to step S503; if the result is "No", the process proceeds to step S504.

[0062] Step S503: Processing is performed to instruct the user to lower the volume of their voice input. Processing to instruct the user to lower the volume of their voice input includes, for example, displaying a warning on device 300 indicating that the user's voice is too loud, increasing the magnification of the user's face displayed in display area 111 or display area 211 on device 300, and decreasing the size / area of ​​display area 111 or display area 211 on device 300. By displaying a warning on device 300 indicating that the user's voice is too loud, it is possible to induce the user to speak more softly, thereby reducing the volume of the user's voice input received by device 300. Furthermore, the process of increasing the magnification of the user's face displayed in display area 111 or display area 211 on device 300 can be achieved, for example, by zooming in using the zoom function of the camera 316 of device 300. By increasing the magnification of the user's face displayed within the display area 111 or display area 211 on device 300, the user's face becomes larger relative to the display area. This allows the user to move their face away from device 300 to adjust their position so that their face fits within the display area, thereby reducing the volume of the user's voice input received by device 300 without requiring the user to speak more softly. Alternatively, by reducing the size and area of ​​the display area 111 or display area 211 on device 300, the user's face becomes larger relative to the display area. This allows the user to move their face away from device 300 to adjust their position so that their face fits within the display area, thereby reducing the volume of the user's voice input received by device 300 without requiring the user to speak more softly.

[0063] If the device 300 determines that the volume of the user's voice input exceeds a first threshold, it may, in addition to or instead of requiring the user to lower the volume of their voice input, perform a process to reduce the volume of the user's voice input. This makes it possible for the device 300 to adjust the volume of the user's voice input that it has already received to an appropriate level.

[0064] Step S504: It is determined whether the volume of the user's voice input is below the second threshold. The second threshold is smaller than the first threshold. If the result is "Yes", the process proceeds to step S505; if the result is "No", the process terminates (successful termination).

[0065] Step S505: A process is executed to instruct the user to increase the volume of their voice input. This process may include, for example, displaying a warning on device 300 indicating that the user's voice is too quiet. By displaying a warning on device 300 indicating that the user's voice is too quiet, it is possible to encourage the user to speak louder, thereby increasing the volume of the user's voice input received by device 300. The process may also include, for example, reducing the magnification of the user's face displayed in display area 111 or display area 211 on device 300. This can be achieved, for example, by zooming out using the zoom function of camera 316 of device 300. By reducing the magnification of the user's face displayed within the display area 111 or display area 211 on device 300, the user's face becomes smaller relative to the display area. This makes it possible to guide the user to move their face closer to device 300 to adjust their face within the display area, thereby increasing the volume of the user's voice input received by device 300 without requiring the user to speak louder. Furthermore, the process of prompting the user to increase the volume of their voice input may include, for example, a process of increasing the size or area of ​​the display area 111 or display area 211 displayed on device 300. By increasing the size or area of ​​the display area 111 or display area 211 displayed on device 300, the user's face becomes smaller relative to the display area. This makes it possible to guide the user to move their face closer to device 300 to adjust their face within the display area, thereby increasing the volume of the user's voice input received by device 300 without requiring the user to speak louder.

[0066] If the device 300 determines that the volume of the user's voice input is below a second threshold, it may, in addition to or instead of the process of prompting the user to increase the volume of their voice input, perform a process to increase the volume of the user's voice input. This makes it possible for the device 300 to adjust the volume of the user's voice input that it has already received to an appropriate level.

[0067] In the embodiments shown in Figures 4 and 5, an example was described in which the processing of each step shown in Figures 4 and 5 is realized by the processor executing a program stored in the memory unit. However, the present invention is not limited to this. At least some of the processing of each step shown in Figures 4 and 5 may be realized by hardware configurations such as control circuits.

[0068] As described above, the present invention has been illustrated using preferred embodiments, but the present invention should not be construed as being limited to these embodiments. It should be understood that the scope of the present invention should be interpreted solely by the claims. Those skilled in the art will understand that, based on the description of the specific preferred embodiments of the present invention and common technical knowledge, an equivalent scope can be implemented. [Industrial applicability]

[0069] The present invention is useful in providing a program, device, and method for assisting a user in performing mouth and / or tongue exercises, which allows the system to determine whether the user is performing the exercises correctly based on the content displayed when the user performs the exercises in a game-like manner. [Explanation of symbols]

[0070] 110 screens 111 Display area 120 screens 121 Contents 122 contents 210 screens 211 Display area 212 Contents 300 devices 311 Interface section 312 Processor section 313 Memory section 314 Input section 315 Display section 316 Camera

Claims

1. A program for assisting a user with mouth and / or tongue exercises, wherein the program is executed by the device's processor unit, Receiving the voice input of the aforementioned user, Recognizing at least one monosyllabic word included in the voice input of the user, To display the content associated with the recognized at least one monosyllabic word on the device. A program that causes the processor unit to perform at least the above.

2. The program according to claim 1, wherein the recognition of the at least one monosyllabic word includes recognizing the at least one monosyllabic word using a speech recognition model suitable for recognizing multiple different monosyllabic words spoken in succession in a short period of time and / or the same monosyllabic word spoken multiple times in succession in a short period of time.

3. The program according to claim 2, wherein the training data for the speech recognition model is the correspondence between monosyllabic words and the speech data of those monosyllabic words.

4. When the program is executed by the processor unit of the device, Displaying an indicator on the device for adjusting the first distance between the user's face and the device. The program according to claim 1, further causing the processor unit to perform the above.

5. The program according to claim 4, wherein the indicator includes a display area for displaying the user as captured by the camera of the device, a bar representing the difference between the first distance and a predetermined appropriate distance, a numerical value representing the difference between the first distance and the predetermined appropriate distance, or guide lines for guiding the position of the user's face as displayed on the device by being captured by the camera of the device.

6. The program according to claim 5, wherein the size of the display area is adjusted such that the distance between the user's face and the device is within a desired range when the user's face within the display area is of a desired size.

7. When the program is executed by the processor unit of the device, To identify the volume of the user's voice input, Determining whether the volume of the user's voice input exceeds a first threshold, If it is determined that the volume of the user's voice input exceeds the first threshold, the magnification of the user's face displayed in the display area is increased. The processor unit is further made to perform the above, and / or When the program is executed by the processor unit of the device, Determining whether the volume of the user's voice input is below a second threshold, If it is determined that the volume of the user's voice input is below the second threshold, the magnification of the user's face displayed in the display area is reduced. The program according to claim 6, further causing the processor unit to perform the above.

8. When the program is executed by the processor unit of the device, To identify the volume of the user's voice input, Determining whether the volume of the user's voice input exceeds a first threshold, If it is determined that the volume of the user's voice input exceeds the first threshold, the process is executed to cause the user to lower the volume of the user's voice input. The processor unit is further made to perform the above, and / or When the program is executed by the processor unit of the device, Determining whether the volume of the user's voice input is below a second threshold, If it is determined that the volume of the user's voice input is below the second threshold, the process is executed to cause the user to increase the volume of the user's voice input. The program according to claim 1, further causing the processor unit to perform the above.

9. A device for assisting a user in mouth and / or tongue exercises, wherein the device comprises a processor unit, and the processor unit is Receiving the voice input of the aforementioned user, Recognizing at least one monosyllabic word included in the voice input of the user, To display the content associated with the recognized at least one monosyllabic word on the device. A device configured to perform at least one of the following actions.

10. A method performed in a device for assisting a user's mouth and / or tongue exercises, wherein the device comprises a processor unit, and the method is The processor unit receives the user's voice input, The processor unit recognizes at least one monosyllabic word included in the user's voice input, The processor unit displays the content associated with the recognized at least one monosyllabic word on the device. Methods that include...

Citation Information

Patent Citations

  • Enunciation training method for hard-hearing person

    JP1986102684A

  • Foreign language learning device, foreign language learning method and medium

    JP2001282098A

  • Speech recognizing device

    JP2006030899A

  • Language learning apparatus

    JP2006163269A

  • Pronunciation training device

    JP2010185967A