Computer program for terminal device, computer program for voice recognition server, and communication system

The terminal device employs dual voice recognition models to enhance string recognition accuracy for users with and without speech disorders, addressing the limitations of existing technologies.

WO2026083743A1PCT designated stage Publication Date: 2026-04-23BROTHER KOGYO KK
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
BROTHER KOGYO KK
Filing Date
2025-09-12
Publication Date
2026-04-23

AI Technical Summary

Technical Problem

Existing voice recognition technologies fail to distinguish between voices of users with and without speech disorders, leading to inaccurate character string recognition.

Method used

A terminal device equipped with a computer program that utilizes two voice recognition models, one for users with speech disorders and another for users without, generating and displaying appropriate character strings based on these models.

Benefits of technology

Enables accurate character string display for both users with and without speech disorders by using separate recognition models, improving communication effectiveness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2025032304_23042026_PF_FP_ABST
    Figure JP2025032304_23042026_PF_FP_ABST
Patent Text Reader

Abstract

The purpose of the present invention is to provide a feature with which it is possible to display an appropriate character string in a situation in which each of the voice of a user having a speech sound disorder and the voice of a user having no speech sound disorder may be inputted. A computer program causes a computer of a terminal device to function as: a supply unit that supplies voice data to a speech recognition engine; an acquisition unit that acquires, from the speech recognition engine, a text string set including first text string data and second text string data; and a display control unit that causes a display unit to display a text string screen including a text string corresponding to the first text string data and / or the second text string data. The first character string data is data generated on the basis of a first speech recognition model for a user having a speech sound disorder, and the second character string data is data generated on the basis of a second speech recognition model for a user not having a speech sound disorder.
Need to check novelty before this filing date? Find Prior Art

Description

Computer program for a terminal device, computer program for a voice recognition server, and communication system

[0001] This specification discloses a technology related to voice recognition.

[0002] Patent Document 1 discloses an application that recognizes voice and displays a character string.

[0003] Japanese Patent Application Laid-Open No. 2019-016206

[0004] In the above technology, nothing is disclosed about distinguishing between a user with speech disorder and a user without speech disorder and recognizing voice. This specification provides a technology that enables a terminal device to display an appropriate character string in a situation where voices of a user with speech disorder and a user without speech disorder can be input respectively.

[0005] This specification discloses a computer program for a terminal device. The computer program causes a computer of the terminal device to function as the following parts: a supply part that supplies voice data corresponding to the voice to a voice recognition engine when the voice is input; an acquisition part that acquires a set of character strings including first character string data corresponding to the voice data and second character string data corresponding to the voice data from the voice recognition engine when the voice data is supplied to the voice recognition engine, where the first character string data is data generated by the voice recognition engine based on a first voice recognition model for a user with speech disorder, and the second character string data is data generated by the voice recognition engine based on a second voice recognition model for a user without speech disorder; and a display control part that causes a display part of the terminal device to display a character string screen including at least one of a first character string corresponding to the first character string data and a second character string corresponding to the second character string data when the set of character strings is acquired from the voice recognition engine.

[0006] According to the above configuration, the terminal device obtains from the speech recognition engine a first string data generated based on a first speech recognition model for users with speech impairments, and a second string data generated based on a second speech recognition model for users without speech impairments. The terminal device then displays a string screen containing at least one of the first string and the second string. To this end, the terminal device can display an appropriate string in situations where both the voice of a user with a speech impairment and the voice of a user without a speech impairment may be input.

[0007] A computer-readable recording medium for storing the above-mentioned computer program is also novel and useful. Furthermore, the terminal device itself and the method executed by the terminal device are also novel and useful. In addition, a communication system comprising the terminal device and a voice recognition server is also novel and useful.

[0008] This specification further discloses a computer program for a speech recognition server. The speech recognition server may include a computer and a memory that stores a first speech recognition model for a user with a speech disorder and a second speech recognition model for a user without a speech disorder. The computer program may cause the computer to function as follows: a generation unit that, when speech data is received from a terminal device, generates first string data corresponding to the speech data based on the first speech recognition model and generates second string data corresponding to the speech data based on the second speech recognition model; a proofreading unit that performs proofreading on a first string corresponding to the first string data and a second string corresponding to the second string data; and a transmission unit that transmits string data corresponding to one of the strings determined to be an accurate string by the proofreading to the terminal device.

[0009] According to the above configuration, the speech recognition server generates first string data based on a first speech recognition model for users with speech disorders and second string data based on a second speech recognition model for users without speech disorders. The speech recognition server then performs proofreading on the first and second strings and transmits the string data corresponding to the string that is determined to be accurate through the proofreading to the terminal device. As a result, the terminal device can display the appropriate string in situations where both the voice of a user with a speech disorder and the voice of a user without a speech disorder may be input.

[0010] A computer-readable recording medium for storing the above-mentioned computer program is also novel and useful. Furthermore, the above-mentioned speech recognition server itself and the method executed by the above-mentioned speech recognition server are also novel and useful. In addition, a communication system comprising the above-mentioned speech recognition server and terminal equipment is also novel and useful.

[0011] The configuration of the communication system is shown. Sequence diagrams for each process in the first embodiment are shown. Sequence diagrams for each process in the second embodiment are shown. Sequence diagrams for each process in the third embodiment are shown. Sequence diagrams for each process in the fourth embodiment are shown. Sequence diagrams for each process in the fifth embodiment are shown.

[0012] (First Embodiment) (Configuration of Communication System 2: Figure 1) As shown in Figure 1, the communication system 2 comprises a mobile terminal 10 and a voice recognition server 100. In this embodiment, a technology is disclosed for the mobile terminal 10 to display strings corresponding to the voices of a user without articulation disorders and a user with articulation disorders.

[0013] (Configuration of the mobile terminal 10) The mobile terminal 10 is a portable user terminal device such as a smartphone, tablet PC, or mobile phone. In the modified example, a stationary user terminal device may be used instead of the mobile terminal 10. The mobile terminal 10 comprises an operation unit 12, a display unit 14, a microphone 16, a communication interface 18, and a control unit 30.

[0014] The operation unit 12 is a user interface for the user to input various information, and includes, for example, hardware keys. The display unit 14 is a display for displaying various information. The display unit 14 functions as a touch panel. That is, the display unit 14 also functions as the operation unit 12. The microphone 16 is a device for inputting the user's voice. The communication interface 18 is a wireless interface in this embodiment, but in modified versions it may be a wired interface. The communication interface 18 may be an interface for connecting to a wireless Local Area Network (LAN) 4, or it may be an interface for performing wireless communication such as 4G or 5G. The mobile terminal 10 can communicate with the voice recognition server 100 via the communication interface 18.

[0015] The control unit 30 comprises a CPU 32 and a memory 34. The memory 34 comprises volatile memory and non-volatile memory. The volatile memory includes RAM and cache memory. The non-volatile memory may be ROM, flash memory, Solid State Drive (SSD), Hard Disk Drive (HDD), or a combination thereof. The non-volatile memory stores programs 40 and 42. Programs 40 and 42 stored in the non-volatile memory are loaded into the volatile memory, and various processes are realized by the CPU 32 executing programs 40 and 42.

[0016] The OS program 40 is a program that enables the basic operation of the mobile terminal 10. The voice display application 42 is installed on the mobile terminal 10 from, for example, a server on the internet (not shown). Hereinafter, the voice display application 42 will be simply referred to as "app 42". When app 42 detects the user's voice input to the microphone 16, it sends voice data corresponding to the voice to the voice recognition server 100 and receives string data corresponding to the voice data from the voice recognition server 100. Then, app 42 displays a screen containing the string corresponding to the string data on the display unit 14.

[0017] (Configuration of the speech recognition server 100) The speech recognition server 100 is located on the Internet 6. The speech recognition server 100 functions as a speech recognition engine. The speech recognition server 100 receives voice data from an external device (e.g., a mobile terminal 10), generates string data corresponding to the voice data, and transmits the string data to the external device. The speech recognition server 100 may be a single server or a collection of multiple servers. Hereinafter, the speech recognition server 100 will be simply referred to as "server 100".

[0018] The server 100 comprises a communication interface 118 and a control unit 130. The communication interface 118 is an interface for communicating with an external device (e.g., a mobile terminal 10). The control unit 130 comprises a CPU 132 and a memory 134. The CPU 132 performs various processes according to a program 140 stored in the memory 134. The memory 134 comprises volatile memory and non-volatile memory. Various processes are realized when the program 140 stored in the non-volatile memory is loaded into the volatile memory and the program 140 is executed by the CPU 132. The memory 134 also stores speech recognition model data 150.

[0019] The speech recognition model data 150 is data for converting speech data into string data. For each of several users with articulation disorders, the speech recognition model data 150 includes a user ID that identifies the user and a speech recognition model for that user. For example, user ID "001" is associated with speech recognition model 152, and user ID "002" is associated with speech recognition model 154. The speech recognition model data 150 further includes a disability-free speech recognition model 160, which is a general-purpose model for users without articulation disorders. Hereafter, speech recognition model 152 (or 154) will be simply referred to as "model 152 (or 154)". Similarly, disability-free speech recognition model 160 will be simply referred to as "model 160".

[0020] A user with a speech impediment, such as the owner of the mobile device 10, accesses the server 100 in advance using the application 42 and registers the user ID "001" with the server 100. The user then inputs various voices of their own into the mobile device 10 and provides the server 100 with voice data corresponding to those voices. Based on this, the server 100 generates a model 152 for the user and stores the generated model 152 in association with the user ID "001".

[0021] (Processing performed by each device 10, 100: Figure 2) Next, referring to Figure 2, the processing performed by each device 10, 100 will be explained. In the following explanation, for ease of understanding, the operations realized by the CPUs 32, 132 of each device will be described from the perspective of the mobile terminal 10 and the server 100, rather than from the perspective of the CPUs.

[0022] Mobile device 10 is owned by a user with a speech impediment (hereinafter referred to as "disabled user"). The disabled user is with a person without a speech impediment (hereinafter referred to as "non-disabled user") and wishes to communicate with the non-disabled user. In this case, the disabled user launches application 42 at T10. As a result, mobile device 10 performs the following processes according to application 42.

[0023] At T12, the mobile terminal 10 displays the login screen on the display unit 14. At T14, the mobile terminal 10 receives an input operation from the disabled user for login information, including user ID "001". The login information may also include a password. At T14, when the mobile terminal 10 receives the input operation for login information, at T16, it sends the login information, including user ID "001", to the server 100.

[0024] When server 100 receives login information from mobile terminal 10 at T16, it identifies model 152 associated with user ID "001" included in the login information. Based on this, server 100 can perform speech recognition on voice data received from mobile terminal 10 thereafter, using model 152, and generate string data corresponding to the voice data. Note that "based on" here can be rephrased as "using" or "consequently".

[0025] Subsequently, the disabled user speaks. In this case, the mobile terminal 10 detects the voice input via the microphone 16 at T20. After the disabled user finishes speaking, if a predetermined period of silence (for example, 3 seconds) has elapsed, the mobile terminal 10 determines that the speaking has ended and, at T22, transmits the voice data corresponding to the voice detected at T20 to the server 100.

[0026] When server 100 receives voice data from mobile terminal 10 at T22, it performs speech recognition on the voice data at T24. Specifically, server 100 performs speech recognition based on the identified model 152 to generate string data corresponding to the voice data, and also performs speech recognition based on model 160 to generate string data corresponding to the voice data. Hereinafter, the former string data will be referred to as "model 152 string data," and the latter string data will be referred to as "model 160 string data." Then, at T30, server 100 sends a string set containing the model 152 string data and the model 160 string data to mobile terminal 10.

[0027] When the mobile terminal 10 receives a string set from the server 100 at T30, at T40, it displays an updated talk screen SC11 on the display unit 14, which contains two strings corresponding to the two string data included in the string set. The updated talk screen SC11 includes one box. This box contains the string "How are you" corresponding to the model 152 string data and the string "Ho are yo" corresponding to the model 160 string data. In this way, in response to a single utterance, one box is displayed that contains both the speech recognition result based on model 152 for disabled users (i.e., "How are you") and the speech recognition result based on model 160 for non-disabled users (i.e., "Ho are yo"). Therefore, the user can recognize the correct string from the two strings contained in the box. In T20, since the voice of a disabled user was input, the speech recognition result "How are you" based on Model 152 for that disabled user is an accurate string (i.e., a string with recognizable meaning). On the other hand, the speech recognition result "Ho are yo" based on Model 160 for a non-disabled user is an inaccurate string (i.e., a string with unrecognizable meaning). Therefore, a non-disabled user can appropriately recognize what the disabled user said by looking at the speech recognition result based on Model 152.

[0028] In particular, within a single box corresponding to a single utterance, the string corresponding to the Model 152 string data and the string corresponding to the Model 160 string data are displayed vertically side by side, with a line break between the former and the latter strings. Therefore, compared to a configuration where the former and the latter strings are displayed without a line break, non-disabled users can easily recognize the former string. In this embodiment, within a single box, the speech recognition result of Model 152 for disabled users is displayed at the top, and the speech recognition result of Model 160 for non-disabled users is displayed at the bottom. However, in a modified version, the speech recognition result of Model 160 for non-disabled users may be displayed at the top, and the speech recognition result of Model 160 for non-disabled users may be displayed at the bottom.

[0029] Next, a non-faulty user speaks. In this case, the mobile terminal 10 detects the voice input via the microphone 16 at T50. After a predetermined period of silence has elapsed since the non-faulty user finished speaking, the mobile terminal 10 determines that the speaking has ended and, at T52, transmits voice data corresponding to the voice detected at T50 to the server 100.

[0030] When server 100 receives voice data from mobile terminal 10 at T52, at T54, similar to T24, it performs voice recognition based on two models 152 and 160 to generate model 152 string data and model 160 string data. Then, at T60, server 100 sends a string set containing model 152 string data and model 160 string data to mobile terminal 10.

[0031] When the mobile terminal 10 receives a string set from the server 100 at T60, at T70, it displays an updated chat screen SC12 on the display unit 14, which contains two strings corresponding to the two string data included in the string set. The updated chat screen SC12 has one box added below the one box displayed at T40. The box added at T70 contains the string "I fin thank you" corresponding to the model 152 string data and the string "I'm fine thank you" corresponding to the model 160 string data. At T50, since the voice of a non-defective user was input, the result of speech recognition based on model 160 for the non-defective user, "I'm fine thank you," is an accurate string. On the other hand, the result of speech recognition based on model 152 for the disabled user, "I fin thank you," is an inaccurate string. Therefore, by viewing the speech recognition results based on Model 160, disabled users can appropriately recognize what a non-disabled user has said.

[0032] Subsequently, each time a disabled or non-disabled user speaks, a box containing the speech recognition results from each of the two models 152 and 160 is displayed. This allows both disabled and non-disabled users to appropriately recognize what the other person has said.

[0033] (Effects of this embodiment) According to this embodiment, the mobile terminal 10 receives from the voice recognition server 100 string data generated based on model 152 for a disabled user and string data generated based on model 160 for a non-disabled user (T30, T60). The mobile terminal 10 then displays updated talk screens SC11 and SC12 containing two strings corresponding to the two string data (T40, T70). In this way, the mobile terminal 10 can display the appropriate string in situations where both the voice of a disabled user and the voice of a non-disabled user may be input.

[0034] (Correspondence) Mobile terminal 10 and server 100 are examples of "terminal device" and "speech recognition engine," respectively. Model 152 and Model 160 are examples of "first speech recognition model" and "second speech recognition model," respectively. Model 152 string data and Model 160 string data are examples of "first string data" and "second string data," respectively. Screens SC11 and SC12 are examples of "string screens." Furthermore, a state of silence continuing for a predetermined period of time after voice input is detected is an example of "a predetermined condition being met."

[0035] T22 and T52 are examples of processes performed by the "supply unit" and "terminal-side transmission unit" of the "terminal device". T30 and T60 are examples of processes performed by the "acquisition unit" and "receiving unit" of the "terminal device". T40 and T70 are examples of processes performed by the "display control unit" of the "terminal device". T24 and T54 are examples of processes performed by the "generation unit" of the "speech recognition server". T30 and T60 are examples of processes performed by the "server-side transmission unit" of the "speech recognition server".

[0036] (Second Embodiment: Figure 3) Next, the second embodiment will be described with reference to Figure 3. In the first embodiment, two strings corresponding to one utterance are displayed in one box. In this embodiment, one of the two strings corresponding to one utterance is displayed inside the box, and the other string is displayed outside the box. That is, the two strings are displayed in different display modes.

[0037] First, the same process as T10 to T16 in Figure 2 is performed. T120 to T130 is the same as T20 to T30 in Figure 2. At T130, when the mobile terminal 10 receives a string set from the server 100, at T140, it displays the updated talk screen SC21 on the display unit 14, which contains two strings corresponding to the two string data included in the string set. The updated talk screen SC21 includes one box containing a string corresponding to the model 152 string data. Below this box, the updated talk screen SC21 also contains a string corresponding to the model 160 string data. The latter string is displayed outside the box and has a smaller font size than the former string. In this way, the two strings are displayed in different ways, so the user can easily understand that the two strings are the result of speech recognition based on different speech recognition models. T150 to T160 is the same as T50 to T60 in Figure 2. T170 is the same as T140 except that the string of characters is different. As a result, the updated chat screen SC22 is displayed in T170.

[0038] In the second embodiment, as an example of displaying two strings in different ways, one string is placed outside the box, and the font sizes of the two strings are different. Alternatively, any of the following embodiments may be adopted: (1) One string and the other string are placed in the same box, and the font sizes of the two strings are different. (2) One string and the other string are placed in the same box, or these strings are placed in separate boxes, and the font colors of the two strings are different. (3) One string and the other string are placed in the same box, or these strings are placed in separate boxes, and the background colors of the two strings are different.

[0039] (Third Embodiment: Figure 4) Next, the third embodiment will be described with reference to Figure 4. In the first and second embodiments, two strings corresponding to a single utterance are displayed simultaneously. In this embodiment, only one of the two strings corresponding to a single utterance is displayed initially, and when a predetermined operation is performed in that state, the other string is displayed.

[0040] First, the same processing as T10 to T16 in Figure 2 is performed. T220 to T230 is the same as T20 to T30 in Figure 2. When the mobile terminal 10 receives a string set from the server 100 at T230, at T232 it performs proofreading on the two strings corresponding to the two string data contained in the string set. Specifically, the mobile terminal 10 identifies the number of characters in each of the two strings and determines that the string with the larger number of characters is the correct string. In particular, at T232, the string "How are you" corresponding to the model 152 string data has 9 characters, and the string "Ho are yo" corresponding to the model 160 string data has 7 characters. Therefore, the mobile terminal 10 determines that the string "How are you" is the correct string. When speech recognition is performed on the same audio data using two different speech recognition models, and two strings are generated, the string with more characters is usually more likely to be the accurate string corresponding to the audio. In this embodiment, the mobile terminal 10 performs proofreading based on the number of characters, so it can appropriately identify the accurate string.

[0041] The mobile terminal 10 displays the updated talk screen SC31 on the display unit 14, which contains boxes with only the strings that were deemed accurate in the proofreading process of T232 in T240. In this way, only one string corresponding to each utterance is displayed, so the user can easily understand the results of the speech recognition.

[0042] The mobile terminal 10 accepts an operation in T242 to select a box within the updated chat screen SC31. In this case, the mobile terminal 10 displays the updated chat screen SC32 on the display unit 14, which contains the string that was not deemed correct during the proofreading in T232, instead of the string contained in the box within the updated chat screen SC31. In this way, the user can check the results of speech recognition based on the other model 160 by performing a predetermined operation while the results of speech recognition based on model 152 of the two models 152, 160 are displayed.

[0043] T250 to T270 are the same as T220 to T240 except that the input voice is different. In particular, in T262, the character string "I fin thank you" corresponding to the model 152 character string data is 12 characters, and the character string "I'm fine thank you" corresponding to the model 160 character string data is 14 characters. For this reason, the mobile terminal 10 determines that the character string "I fin thank you" is the correct character string, and in T240, causes the display unit 14 to display the updated talk screen SC33 in which a box containing only the character string is arranged.

[0044] (Fourth Embodiment: FIG. 5) Subsequently, referring to FIG. 5, the fourth embodiment will be described. In the fourth embodiment, similar to the third embodiment, only one of the two character strings corresponding to one utterance is displayed. However, different from the third embodiment, the other character string is not displayed instead of the one character string.

[0045] First, the same processing as T10 to T16 in FIG. 2 is executed. T320 to T340 are the same as T220 to T240 in FIG. 4. The mobile terminal 10 causes the display unit 14 to display the updated talk screen SC41 in which a box containing only the character string determined to be correct in the proofreading of T332 is arranged in T340. In this embodiment, different from the third embodiment, even if any operation is executed in the state where the updated talk screen SC41 is displayed, instead of the character string included in the box in the updated talk screen SC41, the character string not determined to be correct in the proofreading of T332 is not displayed.

[0046] T350 to T370 are the same as T320 to T340 except that the input voice is different. As a result, the updated talk screen SC42 is displayed. In the updated talk screen SC42, even if any operation is executed, the character string not determined to be correct in the proofreading of T362 is not displayed.

[0047] (Fifth Embodiment: FIG. 6) Subsequently, referring to FIG. 6, the fifth embodiment will be described. In the fifth embodiment, the voice recognition server 100 executes proofreading.

[0048] First, the same processing as that of T10 to T16 in FIG. 2 is executed. T420 to T424 are the same as T20 to T24 in FIG. 2. At T426, the voice recognition server 100 performs proofreading on the two character strings corresponding to the two character string data generated at 424. The method of this proofreading is the same as T232 in FIG. 4. That is, the voice recognition server 100 identifies the number of characters of each of the two character strings, and determines that the one character string having the larger number of characters is the accurate character string. Then, at T430, the voice recognition server 100 transmits one character string data corresponding to the one character string determined to be the accurate character string in the proofreading of T424 to the mobile terminal 10. Here, the voice recognition server 100 does not transmit one character string data corresponding to the one character string not determined to be the accurate character string in the proofreading of T424 to the mobile terminal 10.

[0049] When the mobile terminal 10 receives the character string data from the voice recognition server 100 at T430, at T440, it causes the display unit 14 to display the updated talk screen SC51 in which a box including the character string corresponding to the character string data is arranged. In this way, since one character string determined to be the accurate character string in the proofreading is displayed, the user can appropriately understand the result of the voice recognition. T450 to T470 are the same as T420 to T440 except that the input voice is different. As a result, the updated talk screen SC52 is displayed.

[0050] (Effect of this embodiment) According to this embodiment, the voice recognition server 100 generates character string data based on the model 152 for disabled users and generates character string data based on the model 160 for non-disabled users (T424, T454). Then, the voice recognition server 100 performs proofreading on the two character strings corresponding to the two character string data (T426, T456), and transmits the character string data corresponding to the one character string determined to be the accurate character string by the proofreading to the mobile terminal 10 (T430, T460). For this reason, the mobile terminal 10 can display an appropriate character string in a situation where the voice of a user with speech disorder and the voice of a user without speech disorder can be input respectively.

[0051] (Correspondence) T422 and T452 are examples of processes executed by the "terminal-side transmitting unit" of the "terminal device". T430 and T460 are examples of processes executed by the "receiving unit" of the "terminal device". T440 and T470 are examples of processes executed by the "display control unit" of the "terminal device". T424 and T454 are examples of processes executed by the "generation unit" of the "speech recognition server". T426 and T456 are examples of processes executed by the "proofreading unit" of the "speech recognition server". T430 and T460 are examples of processes executed by the "server-side transmitting unit" of the "speech recognition server".

[0052] Although specific examples of the present invention have been described in detail above, these are merely illustrative and do not limit the scope of the claims. The technology described in the claims includes various modifications and changes to the specific examples illustrated above. Modifications of the above embodiments are listed below.

[0053] (Modification 1) Instead of providing a speech recognition server 100, the application 42 of the mobile terminal 10 may perform speech recognition. In this modification, the application 42 may use the CPU 32 as a detection module for detecting voice input, a speech recognition engine, and a display module. When the detection module detects voice input, it supplies voice data to the speech recognition engine. The display module obtains a string set from the speech recognition engine and, similar to the first embodiment described above, displays an updated talk screen containing each string corresponding to each string data included in the string set on the display unit 14. Note that by having the application 42 perform speech recognition, an updated talk screen similar to that of the second to fifth embodiments may be displayed. In this modification, the detection module is an example of a "supply unit," and the display module is an example of an "acquisition unit" and a "display control unit." In particular, in this modification, the speech recognition engine does not need to have multiple speech recognition models corresponding to multiple disabled users, and may have only one speech recognition model corresponding to one disabled user who is a user of the mobile terminal 10. Generally speaking, a "speech recognition engine" does not need to have multiple first speech recognition models to accommodate multiple users with articulation disorders.

[0054] (Modification 2) In the third embodiment described above, the proofreading of T232 and T262 in Figure 4 may be omitted. In this case, the mobile terminal 10 may first display a box containing the string corresponding to one of the predetermined string data from the Model 152 string data and Model 160 string data included in the string set, and if that box is selected, it may display a box containing the string corresponding to the other string data. Generally speaking, the "display control unit" does not need to perform proofreading.

[0055] (Modification 3) Proofreading is not limited to determining whether a string of characters is accurate. For example, proofreading may include determining whether each word in a string is included in the dictionary based on dictionary data, and determining whether a string has many words included in the dictionary. Alternatively, proofreading may include identifying the number of typographical errors in a string and determining whether a string has few typographical errors.

[0056] (Modification 4) In the first embodiment described above, one box containing two vertically aligned strings is displayed in response to one voice input. Alternatively, various display formats are conceivable. For example, one box containing two horizontally aligned strings may be displayed in response to one voice input. Another box containing one string corresponding to the Model 152 string data and another box containing one string corresponding to the Model 160 string data may be displayed in response to one voice input. That is, two boxes may be displayed in response to one voice input. Furthermore, no boxes may be displayed. For example, two vertically aligned strings may be displayed without a box in response to one voice input. The interval between these two strings is the first interval. In addition, two more vertically aligned strings may be displayed without a box in response to the next voice input. The interval between these two strings is the first interval. Here, the interval between the first two strings displayed and the next two strings displayed is a second interval that is larger than the first interval. Even in this format, users can easily understand that the first two strings of text are displayed in response to one voice input, and the next two strings of text are displayed in response to the next voice input.

[0057] (Modification 5) In each of the above embodiments, the mobile terminal 10 determines that speech has ended when voice is input and silence continues for a predetermined period of time, and sends the voice data to the voice recognition server 100. Alternatively, the mobile terminal 10 may display a screen that includes a start button to be selected before voice input begins and an end button to be selected when voice input is completed. In this case, the mobile terminal 10 detects voice input after the start button is selected, and then determines that speech has ended when the end button is selected, and sends the voice data to the voice recognition server 100. In this modification, the selection of the end button is an example of "a predetermined condition being met".

[0058] (Modification 6) In the above embodiment, each process in Figures 2 to 6 is implemented by software, but at least one of these processes may be implemented by hardware such as a logic circuit.

[0059] Furthermore, the technical elements described herein or in the drawings demonstrate technical usefulness individually or in various combinations, and are not limited to the combinations described in the claims at the time of filing. In addition, the technologies illustrated herein or in the drawings achieve multiple objectives simultaneously, and achieving even one of these objectives constitutes technical usefulness in itself.

[0060] Even if, in the claims of this patent application, each claim depends on only some of the claims, it is not limited to the claim being dependent only on those specific claims. To the extent that it is not technically contradictory, each claim may be dependent on other claims that were not dependent at the time of application. That is, the technologies of each claim can be combined in various ways as follows: (Item 1) A computer program for a terminal device, which causes the computer of the terminal device to function as follows: a supply unit that supplies voice data corresponding to voice input to a voice recognition engine when voice is input; an acquisition unit that, when voice data is supplied to the voice recognition engine, acquires a string set from the voice recognition engine, which includes a first string data corresponding to the voice data and a second string data corresponding to the voice data, wherein the first string data is data generated by the voice recognition engine based on a first voice recognition model for a user with a speech disorder, and the second string data is data generated by the voice recognition engine based on a second voice recognition model for a user without a speech disorder; and a display control unit that, when the string set is acquired from the voice recognition engine, causes the display unit of the terminal device to display a string screen including at least one of a first string corresponding to the first string data and a second string corresponding to the second string data. (Item 2) The computer program according to Item 1, wherein both the first string and the second string are displayed on the string screen. (Item 3) The computer program according to Item 2, wherein the first string and the second string are displayed side by side vertically on the string screen, with a line break between the first string and the second string.(Item 4) The computer program according to Item 2 or 3, wherein the supply unit supplies one voice data corresponding to the voice to the voice recognition engine each time a voice is input and predetermined conditions are met; the acquisition unit acquires one string set from the voice recognition engine each time one voice data is supplied to the voice recognition engine, which includes one first string data corresponding to the voice data and one second string data corresponding to the voice data; and the display control unit places one box corresponding to the string set on the string screen each time one string set is acquired from the voice recognition engine, and the box includes both a first string corresponding to one first string data included in the string set corresponding to the box and a second string corresponding to one second string data included in the string set. (Item 5) The computer program according to any one of Items 2 to 4, wherein the first string and the second string are displayed in different display modes on the string screen. (Item 6) The computer program according to Item 2, wherein the display control unit, when the string set is obtained from the speech recognition engine, places only one of the first string and the second string on the string screen, and when a predetermined operation is performed while one of the strings is placed on the string screen, places the other of the first string and the second string on the string screen in place of the one string. (Item 7) The computer program according to Item 6, wherein the display control unit, when the string set is obtained from the speech recognition engine, performs proofreading on the first string and the second string, and places one of the strings that is determined to be an accurate string by the proofreading on the string screen.(Item 8) The computer program according to Item 1, wherein on the string screen, only one of the first string and the second string is displayed, and the display control unit, when the string set is obtained from the speech recognition engine, performs proofreading on the first string and the second string, and places the one of the strings that is determined to be the correct string by the proofreading on the string screen. (Item 9) The computer program according to Item 7 or 8, wherein the proofreading includes determining that the string with the more characters among the first string and the second string is the correct string. (Item 10) A communication system comprising a terminal device and a speech recognition server, wherein the terminal device comprises a display unit, a terminal-side transmission unit that transmits speech data corresponding to the speech to the speech recognition server when speech is input, a receiving unit that receives a string set from the speech recognition server, which includes a first string data corresponding to the speech data and a second string data corresponding to the speech data, when the speech data is transmitted to the speech recognition server, and a display control unit that causes the display unit to display a string screen containing at least one of a first string corresponding to the first string data and a second string corresponding to the second string data when the string set is received from the speech recognition server, wherein the speech recognition server comprises a memory that stores a first speech recognition model for a user with articulation disorders and a second speech recognition model for a user without articulation disorders, A communication system comprising: a generation unit that generates first string data based on a first speech recognition model and generates second string data based on a second speech recognition model when the voice data is received from the terminal device; and a server-side transmission unit that transmits the string set, including the first string data and the second string data, to the terminal device.(Item 11) A computer program for a speech recognition server, wherein the speech recognition server comprises a computer, a memory for storing a first speech recognition model for a user with a speech disorder, and a second speech recognition model for a user without a speech disorder, and the computer program causes the computer to function as follows: a generation unit that, when speech data is received from a terminal device, generates first string data corresponding to the speech data based on the first speech recognition model and generates second string data corresponding to the speech data based on the second speech recognition model; a proofreading unit that performs proofreading on a first string corresponding to the first string data and a second string corresponding to the second string data; and a transmission unit that transmits string data corresponding to one of the strings that is determined to be an accurate string by the proofreading to the terminal device. (Item 12) The computer program according to Item 11, wherein the proofreading includes determining that the string with more characters than the first string and the second string is the accurate string.(Item 13) A communication system comprising a terminal device and a speech recognition server, wherein the terminal device comprises a display unit, a terminal-side transmission unit that transmits voice data corresponding to the voice to the speech recognition server when voice is input, a receiving unit that receives string data corresponding to the voice data from the speech recognition server when the voice data is transmitted to the speech recognition server, and a display control unit that causes the display unit to display a string screen including a string corresponding to the string data when the string data is received from the speech recognition server, wherein the speech recognition server comprises a memory that stores a first speech recognition model for a user with articulation disorders and a second speech recognition model for a user without articulation disorders, a generation unit that generates the first string data based on the first speech recognition model and generates the second string data based on the second speech recognition model when the voice data is received from the terminal device, and a proofreading unit that performs proofreading on a first string corresponding to the first string data and a second string corresponding to the second string data, A communication system comprising: a server-side transmission unit that transmits the string data corresponding to one of the strings determined to be an accurate string by the aforementioned proofreading to the terminal device.

[0061] 2: Communication system, 10: Mobile terminal, 12: Operation unit, 14: Display unit, 16: Microphone, 18: Communication interface, 30: Control unit, 32: CPU, 34: Memory, 40: OS program, 42: Voice display application, 100: Voice recognition server, 118: Communication interface, 130: Control unit, 132: CPU, 134: Memory, 140: Program, 150: Voice recognition model data, 152, 154: User-specific voice recognition model, 160: Fault-free voice recognition model

Claims

1. A computer program for a terminal device, which causes the computer of the terminal device to function as follows: a supply unit that supplies voice data corresponding to voice input to a voice recognition engine when voice is input; an acquisition unit that, when voice data is supplied to the voice recognition engine, acquires a string set from the voice recognition engine, which includes a first string data corresponding to the voice data and a second string data corresponding to the voice data, wherein the first string data is data generated by the voice recognition engine based on a first voice recognition model for a user with articulation disorders, and the second string data is data generated by the voice recognition engine based on a second voice recognition model different from the first voice recognition model; and a display control unit that, when the string set is acquired from the voice recognition engine, causes the display unit of the terminal device to display a string screen including at least one of a first string corresponding to the first string data and a second string corresponding to the second string data.

2. The computer program according to claim 1, wherein both the first string and the second string are displayed on the string screen.

3. The computer program according to claim 2, wherein the first string and the second string are displayed side by side vertically on the string screen, and a line break is inserted between the first string and the second string.

4. The computer program according to claim 2, wherein the supply unit supplies one audio data corresponding to the audio to the speech recognition engine each time an audio input is received and predetermined conditions are met; the acquisition unit acquires from the speech recognition engine, each time one audio data is supplied to the speech recognition engine, one string set including one first string data corresponding to the audio data and one second string data corresponding to the audio data; and the display control unit places one box corresponding to the string set on the string screen each time a string set is acquired from the speech recognition engine, and the box includes both a first string corresponding to one first string data included in the string set corresponding to the box and a second string corresponding to one second string data included in the string set.

5. The computer program according to claim 2, wherein the first string and the second string are displayed in different display modes on the string screen.

6. The computer program according to claim 2, wherein the display control unit, when the set of strings is obtained from the speech recognition engine, places only one of the first string and the second string on the string screen, and when a predetermined operation is performed while one of the strings is placed on the string screen, places the other of the first string and the second string on the string screen in place of the one string.

7. The computer program according to claim 6, wherein the display control unit, when the set of strings is obtained from the speech recognition engine, performs proofreading on the first string and the second string, and places the string that is determined to be an accurate string by the proofreading on the string screen.

8. The computer program according to claim 1, wherein on the string screen, only one of the first string and the second string is displayed, and the display control unit, when the string set is obtained from the speech recognition engine, performs proofreading on the first string and the second string, and places the one of the strings that is determined to be an accurate string by the proofreading on the string screen.

9. The computer program according to claim 7 or 8, wherein the proofreading includes determining that the string with the greater number of characters among the first string and the second string is the correct string.

10. A communication system comprising a terminal device and a speech recognition server, wherein the terminal device comprises a display unit, a terminal-side transmission unit that transmits voice data corresponding to the voice to the speech recognition server when voice is input, a receiving unit that receives a string set from the speech recognition server, which includes a first string data corresponding to the voice data and a second string data corresponding to the voice data, when the voice data is transmitted to the speech recognition server, a display control unit that causes the display unit to display a string screen containing at least one of a first string corresponding to the first string data and a second string corresponding to the second string data when the string set is received from the speech recognition server, wherein the speech recognition server comprises a memory that stores a first speech recognition model for a user with articulation disorders and a second speech recognition model different from the first speech recognition model, and a generation unit that generates the first string data based on the first speech recognition model and generates the second string data based on the second speech recognition model when the voice data is received from the terminal device, A communication system comprising: a server-side transmitting unit that transmits the string set, which includes the first string data and the second string data, to the terminal device.

11. A computer program for a speech recognition server, wherein the speech recognition server comprises a computer, a memory for storing a first speech recognition model for a user having a speech disorder, and a second speech recognition model different from the first speech recognition model, and the computer program causes the computer to function as follows: a generation unit that, when speech data is received from a terminal device, generates first string data corresponding to the speech data based on the first speech recognition model and generates second string data corresponding to the speech data based on the second speech recognition model; a proofreading unit that performs proofreading on a first string corresponding to the first string data and a second string corresponding to the second string data; and a transmission unit that transmits string data corresponding to one of the strings determined to be an accurate string by the proofreading to the terminal device.

12. The computer program according to claim 11, wherein the proofreading includes determining that the string with the greater number of characters among the first string and the second string is the correct string.

13. A communication system comprising a terminal device and a speech recognition server, wherein the terminal device comprises a display unit, a terminal-side transmission unit that transmits voice data corresponding to the voice to the speech recognition server when voice is input, a receiving unit that receives string data corresponding to the voice data from the speech recognition server when the voice data is transmitted to the speech recognition server, and a display control unit that causes the display unit to display a string screen including a string corresponding to the string data when the string data is received from the speech recognition server, wherein the speech recognition server comprises a memory that stores a first speech recognition model for a user with articulation disorders and a second speech recognition model different from the first speech recognition model, a generation unit that generates first string data based on the first speech recognition model and generates second string data based on the second speech recognition model when the voice data is received from the terminal device, and a proofreading unit that performs proofreading on a first string corresponding to the first string data and a second string corresponding to the second string data, A communication system comprising: a server-side transmission unit that transmits the string data corresponding to one of the strings determined to be an accurate string by the aforementioned proofreading to the terminal device.

Citation Information

Patent Citations

  • Voice Recognition

    JP2023503718A