Information processing device, information processing method and information processing program

The information processing device facilitates accurate voice input of characters with the same pronunciation by displaying conversion candidates for user selection, addressing the challenge of distinguishing between similar characters.

JP2025187440APending Publication Date: 2025-12-25KYOCERA DOCUMENT SOLUTIONS INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2024096238
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-06-13
Publication Date
2025-12-25

AI Technical Summary

Technical Problem

Existing voice input systems struggle to distinguish between multiple different characters with the same pronunciation, such as uppercase and lowercase letters, making accurate input difficult.

Method used

An information processing device and method that includes a speech recognition unit to recognize character-unit texts, a candidate presentation unit to display conversion candidates for characters with the same pronunciation, and a character string determination unit to determine the final string based on user selection.

Benefits of technology

Enables accurate and easy input of multiple different characters with the same pronunciation using voice input.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025187440000001_ABST
    Figure 2025187440000001_ABST
Patent Text Reader

Abstract

To provide an information processing device, information processing method, and information processing program that can accurately and easily distinguish and input multiple different characters with the same pronunciation while using voice input.SOLUTION: An information processing device 10, in a character string input mode, comprises a control circuit that includes: a speech recognition unit that recognizes voice data input as an utterance from a user via a voice input device as multiple character-unit texts; a candidate presentation unit that presents, for the multiple character-unit texts, one or more different characters with the same pronunciation corresponding to the pronunciation of each character-unit text as selectable conversion candidates; and a string determination unit that determines a string based on characters selected by the user for each of the multiple character-unit texts.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to an information processing device, an information processing method, and an information processing program that generate a character string from voice data uttered and input by a user via a voice input device 105. [Background technology]

[0002] Systems for inputting characters by voice are known (Patent Documents 1 and 2). [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Patent Publication No. 2021-96493 [Patent Document 2] Japanese Patent Application Publication No. 2020-136972 Summary of the Invention [Problem to be solved by the invention]

[0004] The order of characters in email addresses, passwords, and file names can be arbitrary. Therefore, when using voice input, each character must be pronounced individually. In this case, multiple different characters with the same pronunciation (for example, uppercase and lowercase letters of the alphabet) cannot be distinguished using voice input alone.

[0005] In view of the above circumstances, an object of the present disclosure is to provide an information processing device, an information processing method, and an information processing program that are capable of distinguishing between multiple different characters with the same pronunciation and inputting them accurately and easily using voice input. [Means for solving the problem]

[0006] An information processing device according to an embodiment of the present disclosure includes: a speech recognition unit that recognizes speech data input by a user via a speech input device as a plurality of character-unit texts in a character string input mode; a candidate presentation unit that presents, for each of the plurality of character unit texts, one or more different characters having the same pronunciation as that of each of the character unit texts, as conversion candidates in a selectable manner; a character string determination unit that determines a character string based on characters selected by a user for each of the plurality of character unit texts; It is equipped with:

[0007] An information processing method according to an embodiment of the present disclosure includes: The computer of the information processing device In a character string input mode, speech data input by a user via a speech input device is recognized as a plurality of character unit texts; For the plurality of character unit texts, one or more different characters having the same pronunciation corresponding to the pronunciation of each character unit text are presented as conversion candidates in a selectable manner; A character string is determined for each of the plurality of character unit texts based on characters selected by the user.

[0008] An information processing program according to an embodiment of the present disclosure includes: The computer of the information processing device, a speech recognition unit that recognizes speech data input by a user via a speech input device as a plurality of character-unit texts in a character string input mode; a candidate presentation unit that presents, for each of the plurality of character unit texts, one or more different characters having the same pronunciation as that of each of the character unit texts, as conversion candidates in a selectable manner; a character string determination unit that determines a character string based on a character selected by a user for each of the plurality of character unit texts; Operate as. [Effects of the Invention]

[0009] According to the present disclosure, it is possible to easily input multiple different characters with the same pronunciation with high accuracy while using voice input.

[0010] The effects described here are not necessarily limited to those described herein, and may be any of the effects described in this disclosure. [Brief explanation of the drawings]

[0011] [Figure 1] 1 illustrates a hardware configuration of an image forming apparatus when the information processing apparatus according to an embodiment of the present disclosure is an image forming apparatus. [Figure 2] 1 shows a functional configuration of an information processing device. [Figure 3] 1 shows an operation flow of the information processing device. [Figure 4] The pronunciation table is shown below. [Figure 5] An example of the default display of the selection screen is shown below. [Figure 6] 10 shows an example of the display after selection on the selection screen. [Figure 7] Indicates a string field. DETAILED DESCRIPTION OF THE INVENTION

[0012] Hereinafter, embodiments of the present disclosure will be described with reference to the drawings.

[0013] 1. Hardware configuration of information processing device

[0014] The information processing device 10 may be any information processing device such as an image forming device such as an MFP, a smartphone, a tablet computer, or a personal computer.

[0015] FIG. 1 illustrates a hardware configuration of an image forming apparatus when the information processing apparatus according to an embodiment of the present disclosure is an image forming apparatus.

[0016] The image forming apparatus 10 includes a control circuit 100 that constitutes a computer. The control circuit 100 is composed of a processor, such as a CPU 11a (Central Processing Unit), a RAM 11b (Random Access Memory), a ROM 11c (Read Only Memory), and dedicated hardware circuits, and is responsible for overall operational control of the image forming apparatus 10. The CPU 11a loads an information processing program stored in the ROM 11c into the RAM 11b and executes it to perform the operations described in the operational flow below and control the display and operational input of the touch panel 17. The ROM 11c permanently stores the programs and data executed by the CPU 11a. The ROM 11c is an example of a non-transitory computer-readable recording medium.

[0017] The control circuit 100 is connected to an image reading unit 12 (image scanner), an image processing unit 14 (including a GPU (Graphics Processing Unit)), an image memory 15, an image forming unit 16 (printer), a touch panel (front panel) 17 which is an operation unit having a display unit 17a, a large-capacity non-volatile storage device 18 such as an HDD (Hard Disk Drive) or an SSD (Solid State Drive), a facsimile communication unit 19, and a network communication interface 13 (communication unit). The control circuit 100 controls the operation of each of the above-mentioned connected units and transmits and receives signals or data to and from each unit. The operation unit of the touch panel 17 is one form of input device, and a voice input device 105 including a microphone may be provided as an input device.

[0018] 2. Functional configuration of information processing device

[0019] FIG. 2 shows the functional configuration of the information processing device.

[0020] The computer of the information processing device 10 operates as a mode switching unit 101, a voice recognition unit 102, a candidate presentation unit 103, and a character string determination unit 104 by executing an information processing program.

[0021] 3. Operation flow of information processing device

[0022] FIG. 3 shows the operation flow of the information processing device.

[0023] The mode switching unit 101 starts the character string input mode (step S1). The character string input mode is a mode for inputting a character string (email address, password, file name, etc.) consisting of an arbitrary sequence of characters, rather than a normal sentence. The mode switching unit 101 may automatically switch to the character string input mode when a field for inputting a character string (email address input field, etc.) is selected. The mode switching unit 101 may also switch to the character string input mode when the user selects to switch modes by pressing an input mode switching button, etc.

[0024] The voice input device 105 accepts a user's speech and transmits voice data to the voice recognition unit 102. The voice input device 105 may be built into or external to the information processing device 10, or may be a device (such as a smartphone, a wearable device, a tablet computer, or a personal computer) communicably connected via a network such as the Internet. For example, when the user utters "P, A, S, S, Q," the voice input device 105 transmits this voice data to the voice recognition unit 102.

[0025] In the character string input mode, the speech recognition unit 102 recognizes speech data input by the user via the speech input device 105 as a plurality of character-unit texts (step S2). For example, the speech recognition unit 102 divides the speech data into character-unit speech data based on silent sections (periods during which the silent time is equal to or greater than a threshold). For example, the speech recognition unit 102 divides the speech data "P, A, S, S, Q" into character-unit speech data "P", "A", "S", "S", and "Q". The speech recognition unit 102 converts the character-unit speech data into text, thereby recognizing it as a plurality of character-unit texts "P", "A", "S", "S", and "Q".

[0026] FIG. 4 shows the pronunciation table.

[0027] The pronunciation table 110 associates and registers one or more different pronunciations 112 (character strings representing pronunciation) and character types 113 (types such as alphabets, numbers, symbols, etc.) with each character 111. For characters 111 whose character type 113 is alphabets, at least one of uppercase and lowercase letters may be registered, and both may be registered. For example, one or more different pronunciations 112 "nine" and "que" and character type 113 "numbers" are associated and registered with the character 111 "9." One or more different pronunciations 112 "que" and character type 113 "alphabet" are associated and registered with the character 111 "q" and / or "Q." In this way, for example, "que," one of the pronunciations of the number "9" in a non-alphabetic language (Japanese), is common to the pronunciation "que" of any of the alphabets "q" and / or "Q." In this case, the common pronunciation 112 "Q" is registered in the pronunciation table 110 in association with the different characters 111 "9" and "q" and / or "Q".

[0028] The candidate presentation unit 103 refers to the pronunciation table 110 stored in the storage device 18, and lists, for each of the multiple character-unit texts, one or more different characters with the same pronunciation that corresponds to the pronunciation of each character-unit text as conversion candidates (step S3). The one or more different characters each having the pronunciation of each character-unit text include at least one of uppercase letters, lowercase letters, numbers, and symbols. In this example, the candidate presentation unit 103 creates the following list, which lists one or more different characters with the same pronunciation that correspond to the pronunciation of each of the character-unit texts "P", "A", "S", "S", and "Q".

[0029] "P": [p,P] "A": [a,A] "S": [s,S] "S": [s,S] "Queue": [q,Q,9]

[0030] FIG. 5 shows an example of the default display of the selection screen.

[0031] The candidate presentation unit 103 presents (displays on the display unit 17a) a selection screen 200 that allows the user to select a character from the list of conversion candidates (step S4). The selection screen 200 displays one or more different characters 202 with the same pronunciation corresponding to the pronunciation 201 in parallel, corresponding to the pronunciation 201 of each character unit text (a character string representing the pronunciation). A check box 203 for selecting the character 202 is displayed corresponding to each character 202. When multiple different characters 202 are displayed for one pronunciation 201, a check mark 204 is placed on one of the characters 202 (for example, the character with the highest frequency of use) by default.

[0032] FIG. 6 shows an example of the display on the selection screen after selection.

[0033] The user selects, for example, via the touch panel 17, one character 202 having the same pronunciation as the pronunciation 201 corresponding to the pronunciation 201 of each character unit text displayed on the selection screen 200. In this example, the user selects (checks 204) the character 202 "p" for the pronunciation 201 "pee", the character 202 "A" for the pronunciation 201 "ae", the character 202 "S" for the pronunciation 201 "es", the character 202 "s" for the pronunciation 201 "es", and the character 202 "9" for the pronunciation 201 "kyu".

[0034] Figure 7 shows a string field.

[0035] The character string determination unit 104 determines a character string 301 based on the characters 202 selected by the user on the selection screen 200 for each of the multiple character unit texts (step S5). In this example, the character string determination unit 104 determines the character string 301 "pASs9" based on the characters 202 "p", "A", "S", "s", and "9" selected by the user on the selection screen 200 for each of the multiple character unit texts "P", "A", "S", "S", and "Q". The character string determination unit 104 inputs the determined character string 301 "pASs9" into the character string field 300 and presents it to the user (displays it on the display unit 17a).

[0036] In this embodiment, the information processing device 10 executes the information processing locally, but as a modification, the information processing may be executed by a server.

[0037] 4. Conclusion

[0038] When using voice input, it is possible to determine uppercase and lowercase letters from the context of normal text. Also, using all lowercase letters often doesn't cause any problems. However, the order of characters in email addresses, passwords, and file names can be arbitrary. Therefore, when using voice input, each character must be pronounced individually. In this case, multiple different characters with the same pronunciation (for example, uppercase and lowercase letters of the alphabet) cannot be distinguished using voice input alone. For email addresses and passwords, it is necessary to clearly distinguish between uppercase and lowercase letters, and voice input makes it impossible to guess the case of the characters.

[0039] For example, when you input "P, A, S, S" by voice, there are multiple possible combinations of uppercase and lowercase letters for the same alphabet, such as "pass," "Pass," and "paSs," making it impossible to guess the correct answer. In addition, there are other characters besides alphabets that have similar pronunciations but different characters, making it difficult to distinguish between them. For example, in Japanese, the alphabet "Q" and "q" and the number "9" are almost identical in pronunciation, making it impossible to distinguish between them by voice input.

[0040] Therefore, according to this embodiment, different characters with the same pronunciation are displayed to allow the user to select one. In a system for inputting characters by voice, when there are multiple characters that correspond to the input voice, candidate characters are displayed on a panel to allow the user to select the correct character. This makes it easy to distinguish between multiple different characters with the same pronunciation and input them accurately and easily using voice input.

[0041] Although the embodiments and modified examples of the present technology have been described above, the present technology is not limited to the above-described embodiments, and it goes without saying that various modifications can be made within the scope of the gist of the present technology. [Explanation of symbols]

[0042] 10. Information processing equipment 100 control circuit 101 Mode switching section 102 Voice Recognition Unit 103 Candidate presentation section 104 String Determination Unit 105 Voice input device

Claims

1. a speech recognition unit that recognizes speech data input by a user via a speech input device as a plurality of character-unit texts in a character string input mode; a candidate presentation unit that presents, for each of the plurality of character unit texts, one or more different characters having the same pronunciation as that of each of the character unit texts, as conversion candidates in a selectable manner; a character string determination unit that determines a character string based on characters selected by a user for each of the plurality of character unit texts; An information processing device comprising:

2. 2. The information processing device according to claim 1, The one or more different characters each having a pronunciation of each character unit text include at least one of uppercase alphabetic characters, lowercase alphabetic characters, numbers, and symbols. Information processing device.

3. 3. The information processing device according to claim 2, The pronunciation of the numbers in a non-alphabetic language is common to the pronunciation of any alphabet. Information processing device.

4. 2. The information processing device according to claim 1, a mode switching section that starts the character string input mode when a field for inputting a character string is selected or when mode switching is selected; An information processing device further comprising:

5. The computer of the information processing device In a character string input mode, speech data input by a user via a speech input device is recognized as a plurality of character unit texts; For the plurality of character unit texts, one or more different characters having the same pronunciation corresponding to the pronunciation of each character unit text are presented as conversion candidates in a selectable manner; A character string is determined based on characters selected by the user for each of the plurality of character unit texts. Information processing methods.

6. The computer of the information processing device, a speech recognition unit that recognizes speech data input by a user via a speech input device as a plurality of character-unit texts in a character string input mode; a candidate presentation unit that presents, for each of the plurality of character unit texts, one or more different characters having the same pronunciation as that of each of the character unit texts, as conversion candidates in a selectable manner; a character string determination unit that determines a character string based on a character selected by a user for each of the plurality of character unit texts; An information processing program that operates as a

Citation Information

Patent Citations

  • Setting input device and image forming apparatus

    JP2020136972A

  • Control device, control system and control program

    JP2021096493A