Voice input device and program
The voice input device facilitates voice input into business software by using a start and end instruction mechanism with a voice recognition dictionary, enhancing efficiency and supporting multiple software types.
Patent Information
- Application Number
- JP2024035045
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-03-07
- Publication Date
- 2025-09-19
AI Technical Summary
Business software primarily requires keyboard input, making voice input difficult.
A voice input device with a start and end instruction mechanism, an input information acquisition unit, and an input unit that allows voice input into multiple areas using a voice recognition dictionary, enabling efficient voice data input into business software.
Enables voice input into business software, improving efficiency by allowing simultaneous data input while listening, and supporting various software types with customized voice recognition dictionaries.
Smart Images

Figure 2025136454000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to a voice input device and a program. [Background technology]
[0002] Patent Document 1 describes a transcription system. The transcription system has a management server and multiple information terminals. The multiple information terminals output character strings transcribed from conversations related to input audio data to the management server. The management server transmits audio data for each divided audio section to the multiple information terminals, combines the individual character strings received from each information terminal, and generates transcribed text data related to the entire conversation of the original audio data. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Application Laid-Open No. 2008-107624 Summary of the Invention [Problem to be solved by the invention]
[0004] Business software mainly requires keyboard input, making voice input difficult.
[0005] An object of the present invention is to enable voice input into an input area. [Means for solving the problem]
[0006] The voice input device has a start instruction means for instructing one of a plurality of input areas displayed to start voice input, an end instruction means for instructing to end voice input, an input information acquisition means for inputting voice via a microphone between the instruction to start voice input and the instruction to end voice input and acquiring input information converted from the voice, and an input means for inputting input information into the plurality of input areas based on the input information acquired by the input information acquisition means, and inputting a control code for moving to the next input area after inputting the input information into the input area. [Effects of the Invention]
[0007] According to the present invention, voice input can be made into the input area. [Brief explanation of the drawings]
[0008] [Figure 1] FIG. 1 is a diagram illustrating an example of the configuration of a voice input system. [Figure 2] FIG. 10 is a diagram illustrating an example of a pop-up menu. [Figure 3] FIG. 10 is a diagram illustrating an example of a voice recognition dictionary registration window. [Figure 4] FIG. 10 is a diagram illustrating an example of a setting screen. [Figure 5] FIG. 10 is a diagram illustrating an example of a sales input screen. [Figure 6] FIG. 10 is a diagram illustrating an example of a journal entry input screen of accounting software. [Figure 7] FIG. 10 is a diagram showing an example of a journal entry input screen of blue return software. [Figure 8] FIG. 10 is a diagram illustrating an example of a sales input screen of the sales software. [Figure 9] FIG. 10 is a diagram showing an example of a product name selection screen. [Figure 10] 10 is a flowchart illustrating a processing method of the voice input device. [Figure 11] 10 is a flowchart illustrating a processing method of the voice input device. DETAILED DESCRIPTION OF THE INVENTION
[0009] 1A is a diagram showing an example of the functional configuration of a voice input system 100 according to this embodiment. A voice input device 101 is, for example, a personal computer, and is capable of inputting voice data into an input area of business software.
[0010] The voice input device 101 can install business software therein and use the business software. In this case, the voice input system 100 includes the voice input device 101, a voice recognition server 102, and a network 104. The network 104 is, for example, the Internet.
[0011] Furthermore, the voice input device 101 can use business software on the business software server 103 without installing business software. In this case, the voice input system 100 includes the voice input device 101, the voice recognition server 102, the business software server 103, and the network 104.
[0012] The speech input device 101 includes a dictionary setting unit 111 , a start instruction unit 112 , an end instruction unit 113 , an input information acquisition unit 114 , and an input unit 115 .
[0013] The dictionary setting unit 111 can generate a voice recognition dictionary 137 (FIG. 1B) for converting voice into input information in response to a user operation. The dictionary setting unit 111 then sets the voice recognition dictionary 137 of FIG. 1B in the voice recognition server 102 via the network 104.
[0014] In response to the operation of the mouse 132 in FIG. 1(B), the start instruction unit 112 instructs the start of voice input into one of the multiple input areas of the business software displayed on the display 134 (FIG. 1(B)) of the voice input device 101.
[0015] The end instruction unit 113 instructs the end of voice input in response to the operation of the mouse 132 in FIG. 1(B).
[0016] 1B via the microphone 133 between the instruction to start voice input and the instruction to end voice input, and requests the voice recognition server 102 to convert the voice using the voice recognition dictionary 137 set by the dictionary setting unit 111. The voice recognition server 102 converts the voice into input information using the voice recognition dictionary 137. The input information acquisition unit 114 acquires the input information into which the voice has been converted from the voice recognition server 102.
[0017] The input unit 115 inputs input information into a plurality of input areas of the business software based on the input information acquired by the input information acquisition unit 114.
[0018] Fig. 1(B) is a diagram showing an example of the hardware configuration of the voice input device 101 of Fig. 1(A). The voice input device 101 is, for example, a personal computer, and includes a CPU 121, a ROM 122, a RAM 123, a communication interface 124, an input device 125, an output device 126, a storage device 127, and a bus 128.
[0019] To the bus 128, a CPU 121, a ROM 122, a RAM 123, a communication interface 124, an input device 125, an output device 126, and a storage device 127 are connected.
[0020] The CPU 121 processes data or performs calculations, and controls various components connected via a bus 128 , and executes the processing of the voice input device 101 .
[0021] The ROM 122 stores a computer program for the CPU 121 in advance, and the computer program is started by the CPU 121 executing it.
[0022] The storage device 127 stores programs such as business software 135 and voice input software 136. The CPU 121 loads the programs stored in the storage device 127 into the RAM 123 and executes them.
[0023] The RAM 123 is used as a working memory for data and temporary storage for controlling each component element.
[0024] The communication interface 124 is an interface for communication, and communicates with the voice recognition server 102 or the business software server 103 via the network 104 in FIG. 1(A).
[0025] The input device 125 has a keyboard 131, a mouse 132, and a microphone 133, and can perform various designations or inputs. The microphone 133 inputs voice and generates voice data. The mouse 132 is an example of a pointing device.
[0026] The output device 126 includes a display 134, a speaker, etc. The display 134 displays the screen of the business software, etc.
[0027] The storage device 127 is a hard disk drive (HDD), a solid state drive (SSD), a memory card, a semiconductor storage device, a disk storage device, etc., and the stored contents are not erased even when the power is turned off. The storage device 127 stores business software 135, voice input software 136, a voice recognition dictionary 137, an item master 138, a customer master 139, a product master 140, etc.
[0028] The business software 135 includes sales software, accounting software, and blue return filing software.
[0029] The voice input software 136 is software for inputting voice data into the input area of the business software 135, and can improve the efficiency of business processes.
[0030] The voice recognition dictionary 137 is a dictionary dedicated to voice recognition, and stores pairs of "conversion results" and "Hiragana readings of the voice." By using the voice recognition dictionary 137, the voice input software 136 can perform voice input suitable for business processing.
[0031] The account master 138 stores account information for each account code. The customer master 139 stores customer information for each customer code. The product master 140 stores product information for each product code.
[0032] The voice input software 136 can perform voice input suited to the characteristics of various types of business software 135. Furthermore, the voice input software 136 enables voice input for any software of the business software 135.
[0033] The voice input software 136 allows efficient voice input of data into business software 135 such as sales management software, accounting software, quotation software, and payroll software.
[0034] By expanding and executing the voice input software (program) 136 in the RAM 123, the CPU 121 can realize the functions of the dictionary setting unit 111, the start instruction unit 112, the end instruction unit 113, the input information acquisition unit 114, and the input unit 115 in Figure 1 (A).
[0035] By storing five voice recognition dictionaries 137 in the storage device 127, one voice input device 101 can be used for five pieces of business software 135.
[0036] By using the voice input device 101, the user can simultaneously input the order while listening to the order details over the telephone.
[0037] FIG. 2 is a diagram showing an example of a pop-up menu 200 of the voice input software 136 that the CPU 121 of the voice input device 101 displays on the display 134. As shown in FIG.
[0038] In the voice recognition dictionary switching 201, the CPU 121 can switch and set one of the five voice recognition dictionaries 137 by operating the mouse 132 by the user.
[0039] In the operation instructions 202, the CPU 121 can display an operation manual for the voice input software 136 on the display 134 in response to the user's operation of the mouse 132.
[0040] When the conversion mode 203 is "Hiragana mode," the CPU 121 always returns the conversion result of the voice recognition in hiragana. The "Hiragana mode" is used when registering the voice recognition dictionary 137 in the voice recognition dictionary registration window 301 in FIG.
[0041] In the voice recognition time 204, the CPU 121 displays the monthly cumulative recognition time and the previous month cumulative recognition time.
[0042] In the version information 205, the CPU 121 displays the version information window 303 in FIG.
[0043] FIG. 3 shows examples of a voice recognition dictionary registration window 301, a voice recognition result display window 302, and a version information window 303 that are displayed on the display 134 by the CPU 121 of the voice input device 101.
[0044] In the voice recognition dictionary registration window 301, the CPU 121 registers a pair of a "conversion result" and a "Hiragana reading of the voice" in the voice recognition dictionary 137 through the user's operation of the keyboard 131 and mouse 132. For example, the CPU 121 registers "quantity, sum ryo" in the voice recognition dictionary 137 and sets it in the voice recognition dictionary 137. When the user inputs the voice of "sum ryo", the CPU 121 requests the voice recognition server 102 to recognize the voice information of "sum ryo". Then, the voice recognition server 102 converts the voice information of "sum ryo" into text information of "quantity" using the voice recognition dictionary 137. The CPU 121 acquires the text information of "quantity" from the voice recognition server 102.
[0045] In addition, in the voice recognition dictionary registration window 301, "VK_RETURN:2" is input information indicating that the Enter key on the keyboard 131 has been pressed twice. Here, the control code of the Enter key is the control code for moving to the next input area. "VK_DOWN:1" is input information indicating that the down arrow key "↓" on the keyboard 131 has been pressed once.
[0046] The voice recognition dictionary 137 shown in the voice recognition dictionary registration window 301 is a dictionary set in the voice input device 101. In addition, the dictionary of the combination of "conversion result" and "hiragana reading of voice" at the beginning of each line of the voice recognition dictionary 137 shown in the voice recognition dictionary registration window 301 is provided to the voice recognition server 102.
[0047] If “110, touzayokin, VK_RETURN:1;” is registered in the voice recognition dictionary 137 of the voice input device 101, the voice input device 101 provides the voice recognition dictionary 137 of “110, touzayokin” to the voice recognition server 102.
[0048] When the voice information of "current account" is input, the voice recognition server 102 recognizes the voice and returns the input information of "110, current account" to the voice input device 101 using the voice recognition dictionary 137. Here, "110" indicates, for example, the account code of current account.
[0049] When the voice input device 101 receives the input information "110, Touza Yokin", it refers to the subject master 138, acquires the subject name corresponding to the subject code "110", and inputs the subject name into the input area. After that, the voice input device 101 inputs the control code VK_RETURN using the voice recognition dictionary 137 of "110, Touza Yokin, VK_RETURN:1;", and moves to the next input area.
[0050] In the speech recognition result display window 302, the CPU 121 displays the speech recognition (hiragana reading of the speech information) and the speech recognition result as the speech recognition result of the speech recognition server 102. The speech recognition result display window 302 can be used mainly for debugging.
[0051] The version information window 303 is displayed by operating the mouse 132 on the version information 205 in FIG.
[0052] 4A and 4B are diagrams showing an example of a setting screen 401 for the voice input software 136 that is displayed on the display 134 by the CPU 121 of the voice input device 101. In the setting area 411, the CPU 121 can perform overall settings for the voice input software 136 and settings related to the voice recognition dictionary 137 in response to user operations on the keyboard 131 and mouse 132.
[0053] 4A is a diagram showing a voice recognition screen in A mode 412, and FIG. 4B is a diagram showing a voice recognition screen in B mode 421. The CPU 121 selects the A mode 412 or the B mode 421 by operating the mouse 132.
[0054] First, the B mode 421 in Fig. 4(B) will be described. The start instruction unit 112 instructs the start of voice input based on an operation including a click of the mouse 132. The end instruction unit 113 instructs the end of voice input based on a movement operation of the pointer of the mouse 132. For example, the end of voice input can be instructed by moving the pointer of the mouse 132 downward, upward, rightward, or leftward. The CPU 121 can select operation modes 422 to 425 of the B mode 421. In the operation mode 422, the CPU 121 can select an operation mode when an instruction to end voice input is given by moving the pointer of the mouse 132 downward.
[0055] In the operation mode 423, the CPU 121 can select an operation mode to be used when an instruction to end voice input is given by moving the pointer of the mouse 132 rightward.
[0056] In the operation mode 424, the CPU 121 can select an operation mode when an instruction to end voice input is given by moving the pointer of the mouse 132 upward.
[0057] In the operation mode 425, the CPU 121 can select an operation mode to be used when an instruction to end voice input is given by moving the pointer of the mouse 132 leftward.
[0058] In the operation modes 422 to 425, external EXE mode, dictionary priority mode, no dictionary mode, cancel, etc. can be selected, respectively.
[0059] 5, the CPU 121 can input voice data into the input area of the sales input screen of the business software 135. In B mode 421, when the start instruction unit 112 instructs the start of voice input, the CPU 121 displays a voice recording message 512 and displays the pointer of the mouse 132 in the center of the voice recording message 512.
[0060] When the pointer of the mouse 132 is moved downward so as to exit the area of the voice recording in progress 512, the end instruction unit 113 instructs the voice input to end, and the operation mode selected in the operation mode 422 is performed.
[0061] When the pointer of the mouse 132 is moved rightward so as to exit the area of the voice recording in progress 512, the end instruction unit 113 instructs the voice input to end, and the operation mode selected in the operation mode 423 is performed.
[0062] When the pointer of the mouse 132 is moved upward so as to exit the area of the voice recording in progress 512, the end instruction unit 113 instructs the voice input to end, and the operation mode selected in the operation mode 424 is performed.
[0063] When the pointer of the mouse 132 is moved to the right so as to leave the area of the voice recording 512 to the left, the end instruction unit 113 instructs the voice input to end, and the operation mode selected in the operation mode 425 is performed.
[0064] For example, the external EXE mode is selected as the operation mode 423. In the external EXE mode, the input information acquisition unit 114 inputs voice via the microphone 133 between the instruction to start voice input and the instruction to end voice input, and acquires input information converted from the voice from the voice recognition server 102 using the voice recognition dictionary 137. The input unit 115 inputs the input information into multiple input areas of the sales input screen 502 based on the input information acquired by the input information acquisition unit 114. That is, the voice input device 101 inputs the input information into multiple input areas with a single voice input using the voice recognition dictionary 137. In the area to the right of the voice recording in progress 512, "EXE" is displayed, indicating the external EXE mode of the operation mode 423.
[0065] For example, the dictionary priority mode is selected as the operation mode 425. In the dictionary priority mode, the input information acquisition unit 114 inputs voice via the microphone 133 between the instruction to start voice input and the instruction to end voice input, and acquires input information converted from the voice from the voice recognition server 102 using the voice recognition dictionary 137. The input unit 115 then inputs input information corresponding to the input information acquired by the input information acquisition unit 114 into one input area on the sales input screen 502. That is, the voice input device 101 inputs input information into one input area with one voice input using the voice recognition dictionary 137. In the area to the left of the voice recording in progress 512, "dictionary priority" is displayed, indicating the dictionary priority mode of the operation mode 425.
[0066] For example, the no-dictionary mode is selected as the operation mode 422. In the no-dictionary mode, the input information acquisition unit 114 inputs voice via the microphone 133 between the instruction to start voice input and the instruction to end voice input, and acquires input information converted from the voice from the voice recognition server 102 without using the voice recognition dictionary 137. The input unit 115 inputs input information corresponding to the input information acquired by the input information acquisition unit 114 into one input area on the sales input screen 502. The voice recognition server 102 converts the voice information into input information (text information) without using the voice recognition dictionary 137. For example, the voice information of "Genkin" (good health) is converted into text information of "cash." The voice input device 101 inputs the input information into one input area with one voice input without using the voice recognition dictionary 137. "No dictionary" indicating the no-dictionary mode of the operation mode 422 is displayed in the area below the voice recording in progress 512.
[0067] For example, cancel is selected as the operation mode 424. In the case of cancel, the input information acquisition unit 114 does not acquire input information from the voice recognition server 102, and the input unit 115 does not input input information. In the area above recording voice 512, "cancel" is displayed, indicating that the operation mode 424 is canceled.
[0068] Next, mode A 412 in Fig. 4(A) will be described. Start instructing unit 112 instructs the start of voice input based on an operation including a click of mouse 132. End instructing unit 113 instructs the end of voice input based on a movement operation of the pointer of mouse 132. For example, any operation of moving the pointer of mouse 132 downward, upward, rightward, or leftward instructs the end of voice input, and the operation of the operation mode selected in operation mode 413 is performed.
[0069] In the operation mode 413, the CPU 121 can select an operation mode when an instruction to end voice input is given by moving the pointer of the mouse 132 downward, rightward, upward or leftward.
[0070] In the operation mode 413, similarly to the operation modes 422 to 425, it is possible to select an external EXE mode, a dictionary priority mode, a no-dictionary mode, a cancel mode, or the like.
[0071] 5, the CPU 121 can input voice data into the input area of the sales input screen of the business software 135. In the A mode 412, when the start instruction unit 112 instructs the start of voice input, the CPU 121 displays a voice recording message 511 and displays the pointer of the mouse 132 in the center of the voice recording message 511.
[0072] When the pointer of mouse 132 is moved downward, rightward, upward, or leftward so as to leave the area of voice recording 512, end instruction unit 113 instructs to end voice input, and the device operates in the operation mode selected in operation mode 413. In other words, no matter which direction the pointer is moved, the device operates in the operation mode selected in operation mode 413. The operation of each operation mode is the same as that of mode B 421.
[0073] 6(A) is a diagram showing an example of a journal entry screen displayed on the display 134 when the CPU 121 executes accounting software of the business software 135. The journal entry screen has multiple input areas such as "Debit Account Title," "Debit Amount," "Consumption Tax Amount," "Credit Amount," "Consumption Tax Amount," "Summary," "Debit Tax Category," and "Credit Tax Category."
[0074] When the user operates the mouse 132 to left-click on the input area 601 for "Debit Account Title," the CPU 121 displays today's date in the "Date" field and the voucher serial number in the "Voucher No." field.
[0075] Next, when the user operates the keyboard 131 and mouse 132 to right-click on the input area 601 while pressing the Ctrl key, the start instruction unit 112 issues an instruction to start voice input.
[0076] 6B, the CPU 121 displays a message 512 indicating that voice input has been started. The user can input voice into the input area 601 using the microphone 133.
[0077] For example, the user speaks, "Meeting fee 11,000 yen in cash, summary: Meeting with two people from XX Corporation, Yakitori Min-chan." After that, the user moves the pointer of the mouse 132 to the right so as to exit the area of the voice recording 512 to the right.
[0078] Then, the end instruction unit 113 gives an instruction to end the voice input. The CPU 121 operates in the external EXE mode because the operation mode 423 corresponding to the rightward movement operation has been selected as the external EXE mode.
[0079] In the external EXE mode, the input information acquisition unit 114 requests the speech recognition server 102 to recognize the speech information "Meeting fee 11,000 yen, cash, summary, meeting with two people from XX Corporation, Yakitori Min-chan" based on the speech recognition dictionary 137, and acquires input information corresponding to the speech information from the speech recognition server 102.
[0080] For example, the input information acquisition unit 114 acquires the subject code or character information that is the conversion result of the voice recognition dictionary 137 from the voice recognition server 102. For example, "meeting expenses" is converted into the subject code of meeting expenses, and "cash" is converted into the subject code of cash.
[0081] In external EXE mode, as shown in voucher No. 1888 in Figure 6 (C), the input unit 115 enters "Conference room" in the input area for the debit account item, "11,000" in the input areas for the debit amount and credit amount, "Cash" in the input area for the credit account item, and "Meeting for two people from XX Trading, Yakitori Minchan" in the input area for the summary.
[0082] The correspondence between the subject code and the subject name is registered in the subject master 138. The CPU 121 refers to the subject master 138 and displays "Conference Expenses" based on the subject code of the conference expenses, and displays "Cash" based on the subject code of the cash.
[0083] Here, if "754, kaigihi, VK_RETURN:1;" is registered in the speech recognition dictionary 137, the input information acquisition unit 114 acquires "754, kaigihi" from the speech recognition server 102. "754" is the subject code. The input unit 115 inputs the input information for "meeting expenses" corresponding to the subject code "754" into the input area 601, and then inputs the VK_RETURN control code for moving to the next input area, and moves to the next input area. For other items, inputting the VK_RETURN control code also moves to the next input area.
[0084] Even if "VK_RETURN:1;" is not registered in the voice recognition dictionary 137, the input unit 115 may input the input information for "meeting fee" into the input area 601, and then input the control code VK_RETURN to move to the next input area, and move to the next input area.
[0085] Furthermore, the input unit 115 inputs the text information "Meeting for two people from XX Trading, Yakitori Min-chan" into the input area for the summary based on the voice information of "Summary: Meeting for two people from XX Trading, Yakitori Min-chan." Since the voice information of "Summary" is an item name, the input unit 115 does not input "Summary" into the input area.
[0086] Furthermore, based on the voice information of "11,000 yen," if the input information acquired by the input information acquisition unit 114 is a numerical amount, the input unit 115 corrects the number to "11,000" in a predetermined format and inputs it into the input area.
[0087] Next, the case of slip No. 1889 in Figure 6(C) will be explained. This case is also similar to the case of slip No. 1888 above.
[0088] When the user operates the keyboard 131 and mouse 132 to right-click on the input area while pressing the Ctrl key, the start instruction unit 112 instructs the user to start voice input.
[0089] The user speaks, "Accounts receivable Chuo Sangyo 2.2 million yen Sales summary Customer management system construction costs Chuo Sangyo Fukuoka branch." After that, the user moves the pointer of the mouse 132 rightward so as to exit the area of the voice recording 512 to the right.
[0090] Then, the end instruction unit 113 gives an instruction to end the voice input. The CPU 121 operates in the external EXE mode because the operation mode 423 corresponding to the rightward movement operation has been selected as the external EXE mode.
[0091] In the external EXE mode, the input information acquisition unit 114 requests the speech recognition server 102 to recognize the speech information of "Accounts receivable Chuo Sangyo 2.2 million yen Sales amount Summary Customer management system construction cost Chuo Sangyo Fukuoka branch" based on the speech recognition dictionary 137, and acquires input information corresponding to the speech information from the speech recognition server 102.
[0092] For example, the input information acquisition unit 114 acquires an account code or character information that is a conversion result of the voice recognition dictionary 137 from the voice recognition server 102. For example, "accounts receivable" is converted into an account code for accounts receivable, and "sales amount" is converted into an account code for sales amount.
[0093] In the external EXE mode, as shown in voucher No. 1889 in Figure 6 (C), the input unit 115 enters "accounts receivable" in the input area for the debit account item, "Chuo Sangyo" in the input area for the debit sub-item, "2,200,000" in the input areas for the debit amount and credit amount, "sales" in the input area for the credit account item, and "customer management system construction costs, Chuo Sangyo Fukuoka Branch" in the input area for the summary.
[0094] The correspondence between the account code and the account name is registered in the account master 138. The CPU 121 refers to the account master 138 and displays "accounts receivable" based on the account code of the accounts receivable, and displays "sales amount" based on the account code of the sales amount.
[0095] Here, if "142, UriKakekin, VK_RETURN:1;" is registered in the speech recognition dictionary 137, the input information acquisition unit 114 acquires "142, UriKakekin" from the speech recognition server 102. "142" is an account code. The input unit 115 inputs "accounts receivable" corresponding to the account code "142" into the input area 601, then inputs the VK_RETURN control code, and moves to the next input area. For other items, input of the VK_RETURN control code also moves to the next input area. Note that even if "VK_RETURN:1;" is not registered in the speech recognition dictionary 137, the input unit 115 may input "accounts receivable" into the input area 601, then input the VK_RETURN control code, and move to the next input area.
[0096] Also, the input unit 115 differs from the case of slip No. 1888 in that it inputs "Chuo Sangyo" into the input area for the debit sub-item.
[0097] In Figures 6(A) to (C), the voice input device 101, by combining the voice input software 136 and the accounting software of the business software 135, performs coding of subjects and sub-subjects, forwarding control, correction of numerical items, etc., and makes corrections based on the characteristics of the accounting software, allowing input using only voice.
[0098] As described above, the input information acquisition unit 114 sequentially inputs a plurality of voices between the instruction to start voice input and the instruction to end voice input, and acquires a plurality of pieces of input information into which the plurality of voices are respectively converted from the voice recognition server 102. The input unit 115 inputs the plurality of pieces of input information acquired by the input information acquisition unit 114 into a plurality of input areas, respectively.
[0099] 7(A) is a diagram showing an example of a journal entry input screen displayed on the display 134 when the CPU 121 executes the blue return software of the business software 135. The journal entry input screen has multiple input areas for "transaction date," "subject," "transaction method," "summary," "client," and "amount."
[0100] When the user operates the keyboard 131 and mouse 132 to right-click on the transaction date input area 701 while pressing the Ctrl key, the start instruction unit 112 issues an instruction to start voice input.
[0101] 6B, the CPU 121 displays "Voice recording in progress" 512 to indicate that an instruction to start voice input has been issued. The user can input voice into the input area 701 using the microphone 133.
[0102] For example, the user speaks, "Transaction date: January 12, 2024, Item: Entertainment expenses, Transaction method: Cash, Summary: Izakaya Shinchan, Amount: 12,500 yen." Then, the user moves the pointer of the mouse 132 to the right so as to exit the area of the voice recording 512 to the right.
[0103] Then, the end instruction unit 113 gives an instruction to end the voice input. The CPU 121 operates in the external EXE mode because the operation mode 423 corresponding to the rightward movement operation has been selected as the external EXE mode.
[0104] In the external EXE mode, the input information acquisition unit 114 requests the speech recognition server 102 to recognize the speech information of "Transaction date: January 12, 2024, Item: Entertainment expenses, Transaction method: Cash, Summary: Izakaya Shinchan, Amount: 12,500 yen" based on the speech recognition dictionary 137, and acquires input information corresponding to the speech information from the speech recognition server 102.
[0105] For example, the input information acquisition unit 114 acquires the item code or character information that is the conversion result of the voice recognition dictionary 137 from the voice recognition server 102. For example, "entertainment expenses" is converted into the item code for entertainment expenses, and "cash" is converted into the item code for cash.
[0106] In external EXE mode, as shown in Figure 7(B), the input unit 115 inputs "2024 / 01 / 12" into the transaction date input area, "entertainment expenses" into the subject input area, "cash" into the transaction method input area, "Izakaya Shinchan" into the summary input area, and "12,500" into the amount input area.
[0107] The correspondence between the account code and the account name is registered in the account master 138. The CPU 121 refers to the account master 138 and displays "Entertainment expenses" based on the account code of the entertainment expenses, and displays "Cash" based on the account code of the cash.
[0108] The input information acquisition unit 114 sequentially inputs a plurality of voices between an instruction to start voice input and an instruction to end voice input, and acquires a plurality of pieces of input information into which the plurality of voices have been converted from the voice recognition server 102. The input unit 115 inputs the plurality of pieces of input information acquired by the input information acquisition unit 114 into a plurality of input areas, respectively.
[0109] Here, after inputting "Entertainment Expenses" into the input area, the input unit 115 inputs the control code VK_RETURN and moves to the next input area. For other items, input of the control code VK_RETURN also moves to the next input area.
[0110] Furthermore, based on the voice information of "Transaction Date January 12, 2024," the input unit 115 inputs the text information of "2024 / 01 / 12" into the input area for transaction date. Because the voice information of "Transaction Date" is an item name, the input unit 115 does not input "Transaction Date" into the input area. Similarly, because "Item," "Transaction Method," "Summary," and "Amount" are item names, the input unit 115 does not input them into the input area.
[0111] 7(A) and (B) has an input area for multiple item names. When the input information acquired by the input information acquisition unit 114 is an item name, the input unit 115 does not input the item name into the input area, but instead inputs input information corresponding to the voice following the item name into the input area for the item name.
[0112] Furthermore, if the input information acquired by the input information acquisition unit 114 is a date based on the voice information of "January 12, 2024," the input unit 115 corrects the date to "2024 / 01 / 12" in a predetermined format and inputs it into the input area.
[0113] Furthermore, if the input information acquired by the input information acquisition unit 114 is a numerical amount based on the voice information of "12,500 yen," the input unit 115 corrects the number to "12,500" in a predetermined format and inputs it into the input area.
[0114] In Figures 7(A) and (B), the voice input device 101, by combining the voice input software 136 and the blue return software of the business software 135, performs coding of subjects, forwarding control, correction of numerical items, etc., and performs corrections based on the characteristics of the blue return software, allowing input by voice only.
[0115] 8(A) is a diagram showing an example of a sales input screen displayed on the display 134 when the CPU 121 executes the sales software of the business software 135. The sales input screen has multiple input areas for customer information in the upper row and multiple input areas for product information in the lower row.
[0116] When the user operates the keyboard 131 and mouse 132 to right-click on the customer code input area 801 while pressing the Ctrl key, the start instruction unit 112 instructs the start of voice input of customer information.
[0117] 6B, the CPU 121 displays "Voice recording in progress" 512 to indicate that an instruction to start voice input has been issued. The user can input voice into the input area 701 using the microphone 133.
[0118] For example, the user speaks "Chuo Sangyo." After that, the user moves the pointer of the mouse 132 to the right so as to exit the area of the voice recording 512 to the right.
[0119] Then, the end instruction unit 113 gives an instruction to end the voice input. The CPU 121 operates in the external EXE mode because the operation mode 423 corresponding to the rightward movement operation has been selected as the external EXE mode.
[0120] In the external EXE mode, the input information acquisition unit 114 requests the voice recognition server 102 to recognize the voice information of "Chuo Sangyo" based on the voice recognition dictionary 137, and acquires input information corresponding to that voice information from the voice recognition server 102.
[0121] For example, if "C001, customer code" is registered in the voice recognition dictionary 137, the input information acquisition unit 114 acquires "C001, customer code" from the voice recognition server 102. "C001" is a customer code.
[0122] In the external EXE mode, as shown in FIG. 8(B), the input unit 115 inputs "C001" into the customer code input area, inputs "Chuo Sangyo Co., Ltd." into the customer name and billing address input area, and also inputs information about Chuo Sangyo Co., Ltd. into other input areas.
[0123] The customer master 139 stores the correspondence between the customer code and the customer information such as the customer name in Fig. 8(B). The CPU 121 refers to the customer master 139 and inputs the customer code of "C001" and the customer information such as the customer name of "Chuo Sangyo Co., Ltd." based on the customer code of "C001".
[0124] Thereafter, the input unit 115 displays a cursor in the product code input area 802 in FIG. 8(A) and sets the input area 802 to an input standby state.
[0125] The input information acquisition unit 114 inputs a single voice of "Chuo Sangyo" between the instruction to start voice input and the instruction to end voice input, and acquires input information into which the voice has been converted from the voice recognition server 102. The input unit 115 inputs a plurality of pieces of input information corresponding to the input information acquired by the input information acquisition unit 114 into a plurality of input areas, respectively.
[0126] Based on the voice of "Chuo Sangyo," the voice input device 101 inputs the corresponding customer code "C001" and the customer name "Chuo Sangyo Co., Ltd.", and automatically advances the cursor to the product code input area 802.
[0127] Next, when the user operates the keyboard 131 and mouse 132 to right-click on the product code input area 802 in FIG. 8(A) while pressing the Ctrl key, the start instruction unit 112 instructs the start of voice input of product information.
[0128] 6B, the CPU 121 displays "Voice recording in progress" 512 to indicate that an instruction to start voice input has been issued. The user can input voice into the input area 701 using the microphone 133.
[0129] For example, the user may say, "Pure silver plated wine cooler, quantity 1." Then, the user moves the pointer of the mouse 132 rightward so as to exit the area of the voice recording 512 to the right.
[0130] Then, the end instruction unit 113 gives an instruction to end the voice input. The CPU 121 operates in the external EXE mode because the operation mode 423 corresponding to the rightward movement operation has been selected as the external EXE mode.
[0131] In the external EXE mode, the input information acquisition unit 114 requests the speech recognition server 102 to recognize the speech information of "Pure silver plated wine cooler, quantity 1" based on the speech recognition dictionary 137, and acquires input information corresponding to the speech information from the speech recognition server 102.
[0132] For example, if "SET-SIL-0003, Jungin plating wine cooler" is registered in the speech recognition dictionary 137, the input information acquisition unit 114 acquires "SET-SIL-0003, Jungin plating wine cooler" from the speech recognition server 102. "SET-SIL-0003" is a product code.
[0133] In the external EXE mode, as shown in FIG. 8(B), the input unit 115 inputs "SET-SIL-0003" into the product code input area, inputs "Pure Silver Plated Wine Cooler" into the product name input area, inputs "1" into the quantity input area, and also inputs information about the Pure Silver Plated Wine Cooler into the other input areas.
[0134] The product master 140 stores the correspondence between the product code and product information such as the product name in Fig. 8(B). The CPU 121 references the product master 140 and inputs the product code "SET-SIL-0003" and product information such as the product name of "Pure Silver Plated Wine Cooler" based on the product code "SET-SIL-0003".
[0135] Thereafter, the input unit 115 displays the cursor in the input area for the next product code and puts that input area into an input standby state.
[0136] Here, the input unit 115 inputs the character information "1" into the input area for the quantity based on the voice information of "quantity 1." Since the voice information of "quantity" is an item name, the input unit 115 does not input "quantity" into the input area.
[0137] Based on the voice of "pure silver plated wine cooler," the voice input device 101 inputs the corresponding product code "SET-SIL-0003" and the product name of "pure silver plated wine cooler."
[0138] Next, we will explain how to input keywords for product names. If the product master 140 contains more than 10,000 products that can be entered on the sales input screen of Figure 8(A), it can be difficult to find the desired product from among them. Therefore, the voice input device 101 makes it possible to present product candidates by combining multiple keywords.
[0139] First, when the user operates the keyboard 131 and mouse 132 to right-click on the product code input area 802 in FIG. 8(A) while pressing the Ctrl key, the start instruction unit 112 instructs the start of voice input of product information.
[0140] 6B, the CPU 121 displays "Voice recording in progress" 512 to indicate that an instruction to start voice input has been issued. The user can input voice into the input area 701 using the microphone 133.
[0141] For example, the user speaks "Chinese lunch box." After that, the user moves the pointer of the mouse 132 rightward so as to exit the voice recording area 512 to the right.
[0142] Then, the end instruction unit 113 gives an instruction to end the voice input. The CPU 121 operates in the external EXE mode because the operation mode 423 corresponding to the rightward movement operation has been selected as the external EXE mode.
[0143] In the external EXE mode, the input information acquisition unit 114 requests the speech recognition server 102 to recognize the speech information of "Chinese lunch box" based on the speech recognition dictionary 137, and acquires the text information of "Chinese, Chinese" and "Lunch box, bento" corresponding to the speech information from the speech recognition server 102.
[0144] Here, the product code corresponding to Chinese food and the product code corresponding to bento lunch are not registered in the voice recognition dictionary 137. In this case, when the voice recognition server 102 receives voice information of "Chinese bento lunch," it returns multiple pieces of character information of "Chinese food, Chinese" and "bento lunch, bento" to the voice input device 101.
[0145] Since the character information acquired from the voice recognition server 102 is not a product code, the input unit 115 searches the product master 140 for product names that include the acquired multiple character information of "Chinese" and "bento". If there are multiple product names in the product master 140 that include the multiple character information of "Chinese" and "bento", the input unit 115 displays a selection screen 902 for candidates for those multiple product names. The selection screen 902 displays "C Sports Chinese Bento", "Chinese Colorful Mix Bento", and "Packed Sports C Chinese Bento" as product name candidates that include "Chinese" and "bento". The user selects the desired product name from the three product names on the selection screen 902. For example, "C Sports Chinese Bento" is selected.
[0146] Then, the input unit 115 refers to the product master 140 and acquires a product code based on the selected product name of "C Sports Chinese Lunch Box." Then, the input unit 115 refers to the product master 140 and inputs the product code, product name, warehouse code, warehouse name, unit price, etc. into the input area based on the acquired product code, as shown in No. 1 on the sales input screen 901 in Fig. 9.
[0147] Thereafter, the input unit 115 displays the cursor in the input area for the next No. 2 product code, and puts the input area into an input standby state.
[0148] Next, when the user operates the keyboard 131 and mouse 132 to right-click on the input area for the No. 2 product code on the sales input screen 901 while pressing the Ctrl key, the start instruction unit 112 instructs the user to start voice input of product information.
[0149] 6B, the CPU 121 displays "Voice recording in progress" 512 to indicate that an instruction to start voice input has been issued. The user can input voice into the input area 701 using the microphone 133.
[0150] For example, the user speaks, "Chinese bento colorful." After that, the user moves the pointer of the mouse 132 rightward so as to exit the voice recording in progress area 512 to the right.
[0151] Then, the end instruction unit 113 gives an instruction to end the voice input. The CPU 121 operates in the external EXE mode because the operation mode 423 corresponding to the rightward movement operation has been selected as the external EXE mode.
[0152] In the external EXE mode, the input information acquisition unit 114 requests the speech recognition server 102 to recognize the speech information of "Chinese lunch box colorful" based on the speech recognition dictionary 137, and acquires the text information of "Chinese, Chinese," "bento, bento," and "colorful, irodori" corresponding to the speech information from the speech recognition server 102.
[0153] Here, the product code corresponding to Chinese, the product code corresponding to bento, and the product code corresponding to color are not registered in the voice recognition dictionary 137. In this case, when the voice recognition server 102 receives voice information of "Chinese bento color," it returns multiple pieces of character information, "Chinese, Chinese," "bento, bento," and "color, irodori," to the voice input device 101.
[0154] Because the character information acquired from the voice recognition server 102 is not a product code, the input unit 115 searches for a product name that includes the acquired multiple pieces of character information, "Chinese," "bento," and "colorful," from the product master 140. If the product master 140 contains only one product name, "Chinese Colorful Mixed Bento," that includes the multiple pieces of character information, "Chinese," "bento," and "colorful," the input unit 115 does not display the selection screen 902.
[0155] The input unit 115 refers to the product master 140 and acquires a product code corresponding to a "Chinese Colorful Mixed Lunch Box" that includes "Chinese," "Lunch Box," and "Colorful" from the product master 140. Then, the input unit 115 refers to the product master 140 and inputs the product code, product name, warehouse code, warehouse name, unit price, etc. into the input area based on the acquired product code, as shown in No. 2 on the sales input screen 901 in Fig. 9 .
[0156] Thereafter, the input unit 115 displays the cursor in the input area for the next No. 3 product code, and puts the input area into an input standby state.
[0157] In Figures 8 and 9, the voice input device 101, by combining the voice input software 136 and the sales software of the business software 135, performs coding of customers and products, feed control, correction of numerical items, etc., and performs corrections based on the characteristics of the sales software, allowing input by voice only.
[0158] Fig. 10 is a flowchart showing the processing method of the voice input device 101. The flowchart in Fig. 10 is realized by the CPU 121 executing a program.
[0159] In step S1001, the CPU 121 sets the voice recognition dictionary 137 for converting voice into input information in accordance with the type of business software 135 used, using the dictionary setting unit 111 and voice recognition dictionary switching 201 in Fig. 2, and registers the voice recognition dictionary 137 in the voice recognition server 102. The CPU 121 also sets the pop-up menu 200 in Fig. 2, registers the voice recognition dictionary 137 in the voice recognition dictionary registration window 301 in Fig. 3, sets the setting screen 401 in Figs. 4(A) and (B), etc.
[0160] In step S1002, the CPU 121 waits until the user right-clicks while pressing the Ctrl key in the input area via the keyboard 131 and mouse 132. If the user right-clicks while pressing the Ctrl key in the input area, the process proceeds to step S1003. This input area becomes the input target.
[0161] The operation of right-clicking while pressing the Ctrl key is not limited to this, and any operation including clicking the mouse 132 may be used. The mouse 132 is not limited to this, and may be any pointing device.
[0162] 6B, and issues an instruction to start voice input. Then, the CPU 121 inputs voice via the microphone 133.
[0163] In step S1004, CPU 121 waits until the pointer of mouse 132 is moved upward, rightward, downward, or leftward so as to leave the area of voice recording in progress 512 in Fig. 6(B). When the pointer is moved upward, rightward, downward, or leftward, CPU 121 instructs end instruction unit 113 to end voice input.
[0164] 4B, when an upward movement operation is performed, the CPU 121 performs a cancellation process in accordance with the setting of the operation mode 424. The CPU 121 does not acquire input information through the input information acquisition unit 114, and does not input input information through the input unit 115. Thereafter, the process of the flowchart in FIG. 10 ends.
[0165] Also, in the case of the setting of B mode 421 in FIG. 4B, when a rightward movement operation is performed, CPU 121 sets external EXE mode in accordance with the setting of operation mode 423, and the process proceeds to step S1011.
[0166] Furthermore, in the case of the setting of B mode 421 in FIG. 4B, when a leftward movement operation is performed, CPU 121 sets dictionary priority mode in accordance with the setting of operation mode 425, and the process proceeds to step S1021.
[0167] Furthermore, in the case of the setting of B mode 421 in FIG. 4B, when a downward movement operation is performed, CPU 121 sets the no-dictionary mode in accordance with the setting of operation mode 422, and the process proceeds to step S1022.
[0168] Also, when the setting is A mode 412 in Figure 4(A), when a movement operation is performed in the upward, rightward, downward, or leftward direction, the CPU 121 proceeds to processing of cancel, external EXE mode, dictionary priority mode, or no dictionary mode, in accordance with the setting of operation mode 413, as in the above-mentioned B mode 421.
[0169] In step S1011, the CPU 121 uses the input information acquisition unit 114 to request the speech recognition server 102 to recognize the speech input in step S1003 using the speech recognition dictionary 137, and acquires the input information into which the speech has been converted from the speech recognition server 102.
[0170] In step S1012, if the acquired input information is a plurality of pieces of input information, CPU 121 sets the first input information as the processing target in order to process the input information in order from the first input information.
[0171] In step S1013, the CPU 121 determines whether the input information to be processed is an item name. The item name is, for example, "Summary" in Figures 6(A) to 6(C), "Transaction Date", "Item", "Transaction Means", "Summary", and "Amount" in Figures 7(A) and 7(B), or "Quantity" in Figures 8(A) and 8(B).
[0172] If the input information is an item name, the process proceeds to step S1014. If the input information is not an item name, the process proceeds to step S1015.
[0173] In step S1014, CPU 121 does not input the input information for the above item name into the input area using input unit 115, but instead inputs as many VK_RETURN control codes as necessary to move the cursor to the input area for the above item name. This makes the input area for the above item name the next input target. CPU 121 then returns to step S1012, processes the next input information, and repeats the above processing.
[0174] In step S1015, if the input information to be processed by the input unit 115 is a date or a number such as an amount of money, the CPU 121 corrects the input information.
[0175] When the input information is a date, such as the "transaction date" in Fig. 7(B), the CPU 121 corrects the date to a predetermined format. For example, the CPU 121 corrects the input information of the date "January 12, 2024" to the input information of the date "2024 / 01 / 12" in a predetermined format.
[0176] 7B, the CPU 121 corrects the input information of the amount of money or other numerical values into a predetermined format. For example, the CPU 121 corrects the input information of the amount of money "12,500 yen" into the input information of the amount of money "12,500" in a predetermined format.
[0177] In step S1016, CPU 121 inputs input information corresponding to the above input information into the input area. Details of step S1016 will be described with reference to FIG.
[0178] Fig. 11 is a flowchart showing the details of step S1016 in Fig. 10. In step S1101, CPU 121 determines whether the input area to be input is the customer code input area, as shown in Figs. 8(A) and (B). If it is the customer code input area, the process proceeds to step S1102. If it is not the customer code input area, the process proceeds to step S1103.
[0179] In step S1102, the CPU 121, as shown in FIG. 8A, refers to the customer master 139 through the input unit 115, and inputs customer information such as the customer code "C001" and the customer name of "Chuo Sangyo Co., Ltd." based on the customer code "C001" acquired from the voice recognition server 102. When inputting data into multiple input areas, the CPU 121 inputs a VK_RETURN control code for each input area, and moves to the next input area. Finally, the cursor is displayed in the product code input area 802 in FIG. 8A, and the input area 802 enters an input standby state. Thereafter, the process proceeds to step S1017 in FIG. 10.
[0180] In step S1103, CPU 121 determines whether the input area to be entered is an input area for a product code, as shown in Figures 8(A) and 8(B). If it is an input area for a product code, the process proceeds to step S1104. If it is not an input area for a product code, the process proceeds to step S1107.
[0181] In step S1104, the CPU 121 determines whether the input information acquired from the speech recognition server 102 is a product code. If it is a product code, the process proceeds to step S1106. If it is not a product code, the process proceeds to step S1105.
[0182] In step S1105, the CPU 121 searches the product master 140 for product names that include one or more pieces of input information (keywords) acquired from the voice recognition server 102. The multiple keywords are, for example, "Chinese" and "bento."
[0183] If there is one product name including one or more pieces of input information, the CPU 121 refers to the product master 140 and acquires the product code corresponding to that product name.
[0184] If there are multiple product names that include one or more pieces of input information, the CPU 121 displays a selection screen 902 of product name candidates, as shown in Fig. 9. The user selects the desired product name from the three product names on the selection screen 902. The CPU 121 references the product master 140 and acquires a product code based on the selected product name.
[0185] In step S1106, the CPU 121 uses the input unit 115 to refer to the product master 140 as shown in FIG. 8(B) or 9, and inputs the product code, product name, warehouse code, warehouse name, unit price, etc. into the input areas based on the acquired product code. When inputting data into multiple input areas, the CPU 121 inputs a VK_RETURN control code for each input area and moves to the next input area. Finally, the cursor is displayed in the input area for the next product code, and that input area is put into an input standby state. Thereafter, the process proceeds to step S1017 in FIG. 10.
[0186] In step S1107, the CPU 121 inputs, via the input unit 115, input information corresponding to the input information acquired from the speech recognition server 102 into the input area, as shown in Fig. 6(B) or 7(B). Thereafter, the CPU 121 inputs a control code VK_RETURN and moves to the next input area.
[0187] In Fig. 6(C), if input information of an amount is input after input of "Meeting Expenses", CPU 121 inputs that amount in the input area for the debit amount. Also, if input information of "Central Industry" that is not an amount is input after input of "Accounts Receivable", CPU 121 inputs "Central Industry" in the input area for the debit sub-item. Then, the process proceeds to step S1017 in Fig. 10.
[0188] In step S1017, when the CPU 121 receives a plurality of pieces of input information from the speech recognition server 102, the CPU 121 determines whether or not processing of all of the pieces of input information has been completed. If processing of all of the input information has not been completed, the CPU 121 returns to step S1012, sets the next piece of input information as the processing target, and repeats the above processing. If processing of all of the input information has been completed, the processing of the flowchart in FIG. 10 ends.
[0189] In step S1021, CPU 121 uses input information acquisition unit 114 to request speech recognition server 102 to recognize the speech input in step S1003 using speech recognition dictionary 137, and acquires input information into which the speech has been converted from speech recognition server 102. Thereafter, the process proceeds to step S1023.
[0190] In step S1022, CPU 121 causes input information acquisition unit 114 to request speech recognition server 102 to recognize the speech input in step S1003 without using speech recognition dictionary 137, and acquires input information into which the speech has been converted from speech recognition server 102. Thereafter, the process proceeds to step S1023.
[0191] In step S1023, the CPU 121 inputs the input information acquired from the voice recognition server 102 into the input area, and then ends the processing of the flowchart in FIG.
[0192] As described above, according to this embodiment, the user can instruct the start of voice input by pressing the Ctrl key and right-clicking, and can instruct the end of voice input by moving the pointer of the mouse 132, so that voice input can be performed with simple operations using the mouse 132.
[0193] Furthermore, the voice input device 101 inputs input information into multiple input areas based on input information acquired from the voice recognition server 102, and after inputting the input information into an input area, inputs a control code for moving to the next input area. This allows the user to input information into multiple input areas with a single voice input.
[0194] This embodiment can be realized by a computer executing a program. A computer-readable recording medium on which the above program is recorded and a computer program product such as the above program can also be applied as an embodiment of the present invention. Examples of recording media that can be used include flexible disks, hard disks, optical disks, magneto-optical disks, CD-ROMs, magnetic tapes, non-volatile memory cards, and ROMs.
[0195] Furthermore, the above-described embodiments are merely examples of specific embodiments for carrying out the present invention, and the technical scope of the present invention should not be construed as being limited by these embodiments. In other words, the present invention can be carried out in various forms without departing from its technical concept or main features. [Explanation of symbols]
[0196] 111 Dictionary Settings 112 Start instruction section 113 End instruction section 114 Input information acquisition unit 115 Input section
Claims
1. a start instruction means for instructing the start of voice input into one of the plurality of input areas displayed; an end instruction means for instructing the end of voice input; an input information acquisition means for inputting voice via a microphone during the period from the instruction to start voice input to the instruction to end voice input, and acquiring input information obtained by converting the voice; an input means for inputting input information into the plurality of input areas based on the input information acquired by the input information acquisition means, and inputting a control code for moving to the next input area after inputting the input information into the input area; A voice input device comprising:
2. the input information acquiring means inputs one voice during a period from the instruction to start voice input to the instruction to end voice input, and acquires input information into which the voice is converted; 2. The voice input device according to claim 1, wherein the input means inputs a plurality of pieces of input information corresponding to the input information acquired by the input information acquisition means into the plurality of input areas, respectively.
3. the input information acquiring means sequentially inputs a plurality of voices between the instruction to start the voice input and the instruction to end the voice input, and acquires a plurality of pieces of input information into which the plurality of voices are respectively converted; 2. The voice input device according to claim 1, wherein the input means inputs a plurality of pieces of input information corresponding to the plurality of pieces of input information acquired by the input information acquisition means into the plurality of input areas, respectively.
4. the plurality of input areas are input areas for a plurality of item names, 4. The voice input device according to claim 3, wherein when the input information acquired by the input information acquisition means is an item name, the input means does not input the item name into the input area, but instead inputs input information corresponding to the voice following the item name into the input area for the item name.
5. 5. The voice input device according to claim 4, wherein, when the input information acquired by the input information acquisition means is a date, the input means corrects the date to a predetermined format and inputs it into the input area.
6. 6. The voice input device according to claim 5, wherein, when the input information acquired by the input information acquisition means is a number, the input means corrects the number to a predetermined format and inputs the number into the input area.
7. The voice input device according to claim 2, characterized in that, when the input area instructed to start voice input is an input area for a product code, and the input information acquired by the input information acquisition means is character information rather than a product code, and there are multiple product names including the character information, the input means displays a selection screen for the multiple product names and inputs input information based on the selected product name into the input area.
8. The voice input device according to claim 7, characterized in that, when the input information acquired by the input information acquisition means is not a product code but a plurality of character information, the input means inputs input information based on a product name including the plurality of character information into the input area.
9. The voice input device according to claim 8, characterized in that, when there are multiple product names including the multiple pieces of character information, the input means displays a selection screen for the multiple product names and inputs input information based on the selected product name into the input area.
10. the start instruction means instructs the start of the voice input based on an operation including a click of a pointing device; 7. The voice input device according to claim 6, wherein the end instruction means instructs the end of the voice input based on a movement of a pointer of a pointing device.
11. The device further includes a dictionary setting means for setting a dictionary for converting speech into input information, the end instruction means instructs the end of the voice input based on a movement operation of a pointer of a pointing device in first to fourth directions, In the case of a movement operation of a pointer of a pointing device in a first direction, the input information acquisition means inputs voice via a microphone between the instruction to start voice input and the instruction to end voice input, and acquires input information converted from the voice using the dictionary, and the input means inputs the input information into the plurality of input areas based on the input information acquired by the input information acquisition means; In the case of a movement operation of a pointer of a pointing device in a second direction, the input information acquiring means inputs a voice via a microphone between the instruction to start the voice input and the instruction to end the voice input, and acquires input information converted from the voice using the dictionary, and the input means inputs input information corresponding to the input information acquired by the input information acquiring means into the one input area; In the case of a movement operation of a pointer of a pointing device in a third direction, the input information acquiring means inputs a voice via a microphone during a period from the instruction to start the voice input to the instruction to end the voice input, and acquires input information converted from the voice without using the dictionary, and the input means inputs input information corresponding to the input information acquired by the input information acquiring means into the one input area; 11. The voice input device according to claim 10, wherein, in the case of a movement operation of a pointer of a pointing device in a fourth direction, the input information acquisition means does not acquire input information, and the input means does not input input information.
12. 12. The voice input device according to claim 11, wherein the input area is an input area for sales software, an input area for accounting software, or an input area for blue tax return software.
13. start instruction means for instructing the start of voice input into a displayed input area based on an operation including a click of a pointing device; an end instruction means for instructing the end of voice input based on a movement operation of a pointer of a pointing device; an input information acquisition means for inputting voice via a microphone during the period from the instruction to start voice input to the instruction to end voice input, and acquiring input information obtained by converting the voice; an input means for inputting input information into the input area based on the input information acquired by the input information acquisition means; A voice input device comprising:
14. a start instruction means for instructing the start of voice input into the input area of the displayed product code; an end instruction means for instructing the end of voice input; an input information acquisition means for inputting voice via a microphone during the period from the instruction to start voice input to the instruction to end voice input, and acquiring input information obtained by converting the voice; an input means for inputting input information based on a product name including the plurality of character information into the input area when the input information acquired by the input information acquisition means is not a product code but a plurality of character information; A voice input device comprising:
15. A program for causing a computer to function as the voice input device according to any one of claims 1 to 14.
Citation Information
Patent Citations
Transcription system
JP2008107624A