Information processing program, information processing device, and information processing method
The information processing program and device automatically determine input formats for voice input fields, addressing the challenge of expert dependence in existing voice input methods, enhancing data entry efficiency and accuracy.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- KK TOSHIBA
- Filing Date
- 2023-06-13
- Publication Date
- 2026-05-08
AI Technical Summary
Existing voice input methods for data entry in manufacturing and maintenance sites require speech recognition experts to determine appropriate input formats, which is challenging in their absence.
An information processing program and device that includes an information acquisition unit and a format estimation unit to automatically determine input formats for voice input fields based on data from recording sheets, using input value examples and estimation algorithms.
Enables accurate and efficient voice-based data entry without the need for speech recognition experts by estimating appropriate input formats, improving data entry speed and reducing errors.
Smart Images

Figure 0007855550000001 
Figure 0007855550000002 
Figure 0007855550000003
Abstract
Description
[Technical Field]
[0001] Embodiments of the present invention relate to an information processing program, an information processing device, and an information processing method. [Background technology]
[0002] In manufacturing and maintenance sites, measurement values from measuring instruments and visual inspection results are sometimes entered into record data sheets such as forms and tables, and these data sheets are then shared among workers or with customers. In this case, the contents to be entered into the record data sheet are predetermined, and workers perform their tasks according to the work procedures and enter the obtained data into the designated locations on the record data sheet.
[0003] While typical form software requires data to be entered as text, entering text during work is time-consuming, creating a need for voice-based data entry. For example, by configuring input fields and their content in a separate application from the form software, a voice input method can be used to input values into selected fields when spoken. Furthermore, by specifying the next field to be entered during setup, values can be entered into fields consecutively. With such a voice input method, a speech recognition expert can pre-define the input format of the values to be entered, making it possible to perform voice input based on that format.
[0004] While the above-described voice input method is usually fine, according to the inventor's research, for example, In the absence of speech recognition experts, there is room for improvement in determining an appropriate input format, which can be challenging. [Prior art documents] [Patent Documents]
[0005] [Patent Document 1] Japanese Patent Publication No. 2008-52676 [Overview of the Initiative] [Problems that the invention aims to solve]
[0006] The problem that this invention aims to solve is, for example, The objective is to provide an information processing program, information processing device, and information processing method that can determine an appropriate input format even in the absence of a speech recognition expert. [Means for solving the problem]
[0007] The information processing program according to this embodiment implements an information acquisition unit and a format estimation unit on a computer. The information acquisition unit acquires, for each item, one or more items and information regarding the values of the input fields for the one or more items from a recording data sheet that includes voice input fields. The format estimation unit estimates the input format of the input fields based on the information. [Brief explanation of the drawing]
[0008] [Figure 1] A block diagram showing the configuration of an information processing device according to the first embodiment. [Figure 2] A diagram showing an example of report data according to the first embodiment. [Figure 3] A diagram showing an example of the initial settings of the input procedure storage unit according to the first embodiment. [Figure 4] A figure showing an example of a word list according to the first embodiment. [Figure 5] A diagram showing an example of a completed form according to the first embodiment. [Figure 6] A diagram showing the automatically configured input procedure storage unit according to the first embodiment. [Figure 7] A diagram showing an example of an input format correspondence table for guidance according to the first embodiment. [Figure 8] A flowchart illustrating the input procedure creation process according to the first embodiment. [Figure 9] A diagram showing an example of range selection according to the first embodiment. [Figure 10] A diagram showing an example of an alert according to the first embodiment. [Figure 11] A diagram showing another example of an alert according to the first embodiment. [Figure 12] A diagram showing an example of the input format display for a form according to the first embodiment. [Figure 13] A figure showing an example of how input values are displayed in a form according to the first embodiment. [Figure 14] A diagram showing an example of the input order display for a form according to the first embodiment. [Figure 15] A diagram showing another example of the input order display for a form according to the first embodiment. [Figure 16] A diagram showing the formal estimation process (without concatenation) according to the first embodiment. [Figure 17] A diagram showing an example of a format sorting rule according to the first embodiment. [Figure 18] A diagram showing formal estimation (without linking) according to the first embodiment. [Figure 19] A diagram showing the format check process according to the first embodiment. [Figure 20] A diagram showing the voice input processing according to the first embodiment. [Figure 21] A block diagram showing the configuration of the information processing device according to the second embodiment. [Figure 22] A figure showing an example of automatic range selection according to the second embodiment. [Figure 23] A flowchart illustrating the input procedure creation process according to the second embodiment. [Figure 24] A block diagram showing the configuration of the information processing device according to the third embodiment. [Figure 25] A flowchart illustrating the input procedure creation process according to the third embodiment. [Figure 26] A flowchart showing the formal estimation process (without linking) according to the third embodiment. [Figure 27] A flowchart illustrating the format check process according to the third embodiment. [Figure 28] A block diagram showing the configuration of the information processing device according to the fourth embodiment. [Figure 29] A flowchart showing the operation example display process according to the fourth embodiment. [Figure 30] A diagram showing an example of the hardware configuration of an information processing device according to the fifth embodiment. [Modes for carrying out the invention]
[0009] The embodiments will be described in detail below with reference to the drawings. In the following description, we will use as an example the case in which an information processing device such as a tablet terminal with voice input functionality performs input procedure creation processing and voice input processing. Note that the information processing device is not limited to a tablet terminal; any computer such as a PC (Personal Computer) or smartphone equipped with an external microphone and external speakers can be used as appropriate. Furthermore, the information processing device does not necessarily need to perform both input procedure creation processing and voice input processing; it is sufficient if it is a device that performs at least input procedure creation processing. That is, the program for input procedure creation processing and the program for voice input processing may be installed as a single information processing program, or as separate information processing programs. In either case, the information processing device realizes the respective functions of input procedure creation processing and voice input processing by executing the information processing program. Furthermore, if the information processing device performs only input procedure creation processing, the microphone may be omitted. In the following description, the user related to input procedure creation processing will be called the input procedure creator, and the user related to voice input processing will be called the voice inputter. The voice inputter and the person who creates the input procedure may be the same person or different people. Furthermore, multiple voice inputters and multiple input procedure creators may collaborate.
[0010] <First Embodiment> Figure 1 is a block diagram showing an example of the configuration of an information processing device according to the first embodiment. This information processing device 1 comprises a data storage unit 2, an input procedure storage unit 3, an input value example acquisition unit 4, a range selection unit 5, a format estimation unit 6, a format checking unit 7, a dictionary generation unit 8, a speech recognition unit 9, a speech synthesis unit 10, a recognition control unit 11, and a display unit 12. The information processing device 1 is a computer that realizes the functions of each unit by executing an information processing program installed in memory.
[0011] Here, as shown in Figure 2, the data storage unit 2 stores form data 2a, which represents data with voice input fields for each item, such as forms and tables. Form data 2a is an example of a recording data sheet that includes voice input fields for each item. For example, in form data 2a, values are entered into the input fields associated with each row indicated by row numbers 1 to 5 and each item in column numbers B to D, and values are entered into the input fields associated with each row indicated by row numbers 8 to 9 and each item in column numbers B to D. Note that form data 2a may also include a word sheet, which will be described later, on a separate sheet. The data storage unit 2 is an example of computer memory.
[0012] As shown in Figure 3, the input procedure storage unit 3 stores input procedures 3a for entering values into multiple input fields contained in the form data 2a, which is a record data sheet. The input procedure 3a includes, for example, a procedure number, input target item, input format, guidance, and reference value as item names. However, the procedure number, guidance, and reference value are optional additional items and may be omitted. Such an input procedure 3a also functions as a list of the order in which to enter data into multiple items. The order of the input procedures 3a may be determined from top to bottom in the list, or a separate number representing the order may be included as an item in the input procedure. The input procedure storage unit 3 is an example of computer memory.
[0013] Here, the item name for the procedure number represents a number indicating the order in which values are entered into the multiple input fields included in the report data 2a, as the value of the procedure number.
[0014] The item name of the input target item represents an identifier that identifies the input field of the input target item included in the report data 2a, as the value of the input target item. The identifier can be written using column and row numbers, for example, D2, if the report data is in tabular format as shown in Figure 2. In Figure 3, input procedure 3a shows an example where the first measurement results are entered into B1, B2, ..., B5, then into B8, C8, D8, and then into C1, C2, ..., C5 for the second measurement results.
[0015] The item name of the input format represents the input format estimated by the format estimation unit 6 as the value of the input format. Input procedure 3a includes an input format field where the estimated input format can be entered, and the input format estimated by the format estimation unit 6 is entered into this input format field. To clarify, the input format specifies the format of the value to be entered into the input field of the input target item. The recognition control unit 11, described later, creates a speech recognition dictionary based on the input format entered in the input format field. Since directly describing the speech recognition dictionary in the input procedure is difficult, this embodiment uses a simplified input format.
[0016] The main categories of input formats that can be used as appropriate include, for example, 'numbers', 'alphanumeric characters', 'date and time', 'date', 'time', and 'words'.
[0017] In the case of a 'numerical value', it may also be possible to specify the number base (decimal, hexadecimal, etc.), the number of digits in the integer part (if decimal, it may be possible to specify a range), the number of digits in the decimal part, etc.
[0018] For example, if the input is a decimal number with an integer part range of Mmin to Mmax and a fractional part range of Nmin to Mmin, the input format will be written as "decimal_Mmin~Mmax_Nmin~Nmax". Also, if there is no upper limit to the number of digits in the "number", the input format will be written as "decimal_Mmin~_Nmin~", and if there is no lower limit to the number, it will be written as "decimal_~Mmax_~Nmax". Furthermore, if there is no range for the number of digits in the "number" (integer part digits M, fractional part N), the input format will be written as "decimal_M_N". Additionally, if a positive or negative sign "+" (plus) or "-" (minus) is spoken before the number in the "number", the input format will be written as "[+-]decimal_M_N".
[0019] In the case of "alphanumeric," you can further specify patterns of letters, numbers, and symbols. For example, as an alphanumeric pattern, you could write "A" for letters, "D" for numbers, and "S" for symbols. Alternatively, you can use the character "$" for letters or numbers, and "!" for letters, numbers, or symbols. For example, a pattern of 3 letters, 2 numbers, 1 symbol, and 1 letter could be written as "alphanumeric_AAADDSA".
[0020] Furthermore, in the case of "alphanumeric characters," optional characters (characters that may or may not be entered) should be preceded by a '?'. For example, the pattern 'alphanumeric_AAA?DD' represents a pattern of two or three alphanumeric characters and two numbers. In addition to single-character patterns, patterns that allow for multiple characters may also be described.
[0021] For "Date and Time," "Date," and "Time," no additional specifications are required. However, for "Date," it may be possible to specify additional details, such as specifying which part of the date to enter, such as "Year Month Day," "Month Day," "Year Month," or "Day."
[0022] In the case of "words," as shown in Figure 4 as an example, a word list 2b is written on a separate sheet of the form data 2a, and the group name of the word list 2b is specified. The word list 2b includes the group name, the spelling of the word, the corresponding reading (the reading uttered by the user), and the read-aloud (the content read aloud when repeating in guidance, etc.). In this embodiment, since the word list 2b is generated automatically, a number is automatically assigned to the group name. Since only the spelling of the word can be obtained from the input value example, the reading and read-aloud content are generated from the spelling. The reading and read-aloud items are provided as separate items so that the input procedure creator can change them later, but at the generation stage, the same content should be generated and set. Any existing method can be used to generate the reading and read-aloud. For example, morphological analysis technology and a dictionary for morphological analysis can be used.
[0023] Note that the input format is not limited to the major categories mentioned above; any input format that can be converted to the speech recognition dictionary is acceptable, and combinations of individual input formats may also be used. For example, a concatenated format such as "date_month_day, word_day_of_week, (time_hour_minute)" can be used as the input format. In this case, after uttering the date, the word group "day of the week" (Sunday, Monday, Tuesday, ...) can be uttered, followed by the time, or not uttered at all. The parentheses within the concatenation indicate that the utterance is optional. In this case, the recognized results will be space-separated. For example, with "date_month_day, word_day_of_week, (time_hour_minute)", the recognized results may include "3 / 10(space)Friday(space)9:10" or "3 / 11(space)Saturday".
[0024] Furthermore, the input format may also use an OR of multiple input formats, such as "time_hours_minutes|time_hours_minutes_seconds" or "decimal_2_1|word_abnormal_present_abnormal". In this case, either "time_hours_minutes" or "time_hours_minutes_seconds" can be spoken, and either "decimal_2_1" or "word_abnormal_abnormal_abnormal" can be spoken. For example, with "time_hours_minutes|time_hours_minutes_seconds", "9:10" and "13:20:10" can be recognized, and with "decimal_2_1|word_abnormal_abnormal_abnormal", "12.3" and "abnormal" can be recognized.
[0025] Furthermore, the input format may also be a combination of the above-mentioned concatenated formats or OR operations. The method of describing the input format is also not limited to those mentioned above; any method that can be converted to a speech recognition dictionary is acceptable. Note that specifying the input format as narrowly as possible (narrow range of digits for numbers, fewer words for words) improves the speech recognition rate, but this limits the types of speech that can be input, making it less user-friendly for the voice inputter if the specification is too detailed.
[0026] The guidance item names are used as guidance values, and are played at the start of each step to provide work instructions when actually performing the input procedure and using voice input. The guidance is not limited to the item names of the input fields; any guidance related to voice input into the input field is acceptable. For example, in the case of an inspection input procedure, the guidance would include descriptions that instruct (guide) the content of the inspection procedure. Of all the guidance, only those that match a specific pattern, such as "XX time," will be used in subsequent input format checks, etc.
[0027] The item name for the reference value indicates the range of values that will serve as the standard for the values entered in the input field. The reference value is a value that indicates the range of values in the input field and may include an upper limit, a lower limit, or both upper and lower limits. The reference value is used to check whether the value entered via voice input is abnormal when actually performing the input procedure and using voice input.
[0028] The input procedure 3a described above may not only be stored in the input procedure storage unit 3 as separate data from the report data 2a, but may also be stored on the report data 2a. For example, in a typical spreadsheet application, multiple spreadsheet sheets can be created on a single file, so one of these may be used for the report data and another for storing the input procedure. In this case, the recognition control unit 11, which will be described later, will obtain the input procedure from the sheet used for storing the procedure. Also, when the input procedure 3a is stored separately from the report data 2a, it may be stored in storage such as a DB (Database), or it may be stored in memory. In this embodiment, it is assumed that there is a sheet for storing the input procedure on a separate sheet in the report data 2a file. In this case, by copying the sheet for storing the input procedure to another report data file, voice input can also be performed on the destination report data.
[0029] Note that input procedure 3a may also be created by manually entering information in areas other than the input format. For example, if a sheet for retaining procedures is created in the report data 2a file, input procedure 3a other than the input format may be created by manually entering a table similar to Figure 3 into the sheet for retaining procedures. Alternatively, input procedure 3a may also be created by clicking on the input target item to display a dialog box and manually entering information other than the input format in the dialog box.
[0030] The input value example acquisition unit 4 acquires, for each item, one or more items and information regarding the values of the input fields for one or more items from the form data 2a, which includes voice input fields. As information regarding the values of the input fields, for example, input value examples, which are examples of the values of the input fields, and formatting used to display the values of the input fields can be used as appropriate. In the first embodiment, input value examples are used as information regarding the values of the input fields. In this case, the input value example acquisition unit 4 can acquire the values as input value examples from at least one of the form data 2a in which the values of the input fields have been pre-filled, and the user interface that accepts input of values into the input fields. As for the form data 2a in which the values of the input fields have been pre-filled, if at least one of the input fields of the form data 2a has been filled in, an input value example can be acquired from that filled-in input field. The input value example acquisition unit 4 may also acquire information regarding the values of the input fields corresponding to the range of input fields selected by the range selection unit 5. Note that "information regarding the values of the input fields" may be appropriately read as "input hints for the input fields," "input hint information for the input fields," etc. The input value example acquisition unit 4 is an example of an information acquisition unit.
[0031] Figure 5 shows an example of pre-entered report data 2a. In this report data 2a, only the measurement results for the first and second measurements, and for noise measurement, only positions 1 and 2 of #1 are entered. The input value example acquisition unit 4 obtains input value examples from the pre-entered report data 2a by acquiring the value of the item pointed to by the identifier of the input target item for each input procedure. For example, for the report data 2a in Figure 5, the input value example for input procedure number 1 in Figure 3 only needs to acquire the value of the input target item with identifier B1, which is "10:30". Such pre-entered report data 2a may be one that was manually entered before setting up the voice input procedure, or new input value examples may be manually entered for the purpose of setting up. Also, there may be more than one pre-entered report data 2a. If there are multiple report data 2a, multiple input value examples will be obtained for a single input procedure. Furthermore, if the procedure is to be set from a user interface such as a dialog box, a text box for accepting example input values may be added to the dialog box, and the example input value acquisition unit 4 may acquire the example input values by acquiring the text entered in the text box.
[0032] The range selection unit 5 selects a range of multiple input fields from the report data 2a. To elaborate, the range selection unit 5 selects a range of multiple input fields as the range for estimating the input format. For example, the range selection unit 5 provides a user interface for selecting the range for which to set the input format. The range selection unit 5 may use the functions of an existing application that displays the report, or it may use other methods.
[0033] The format estimation unit 6 estimates the input format of an input field based on the input value examples acquired by the input value example acquisition unit 4. The format estimation unit 6 may also estimate a common input format for input fields in a range based on the input value examples acquired corresponding to the input fields in the range selected by the range selection unit 5. The common input format is estimated from multiple input value examples. The format estimation unit 6 may also estimate the input format for each input field in the range and estimate one of the estimated input formats as the common input format. Alternatively, the format estimation unit 6 may estimate the input format for each input field in the range and estimate an input format encompassing multiple of the estimated input formats as the common input format. The format estimation unit 6 may also adjust the estimated input format based on the number of input value examples acquired corresponding to the input fields in the selected range. Furthermore, the format estimation unit 6 may adjust the estimated input format based on reference values (e.g., upper and lower limits) included in the input procedure 3a. In any case, the format estimation unit 6 estimates the input format from the input value examples and updates the input format field in input procedure 3a with the estimated input format. Figure 6 shows an example of the updated input procedure 3a. In this case, the input format has already been set for input procedure 3a for all items except B4, C4, and D4.
[0034] The format checking unit 7 checks the validity of the estimated input format. For example, the format checking unit 7 checks the estimated input format based on the guidance included in the input procedure 3a.
[0035] For example, the format checking unit 7 generates example input values from the estimated input format, synthesizes the generated example input values into speech, and performs speech recognition on the speech data (speech waveform data) obtained from the speech synthesis using a speech recognition dictionary. The text of the speech recognition result represents an example input value, a command, or something else. The format checking unit 7 may also check the estimated input format by determining whether the obtained speech recognition result text matches the generated example input value. The format checking unit 7 may also classify the situation when there is no match with the example input value by determining whether a command was detected in the speech recognition result text. However, the determination regarding command detection is performed not only when the determination result regarding the match with the example input value is negative. The format checking unit 7 may also check the input format by calculating the recognition rate or misrecognition rate of speech recognition based on the number of determinations and the determination results. The format checking unit 7 may also notify the results of the input format check in an identifiable manner. The results of the input format check include, for example, judgment results regarding matches with example input values, judgment results regarding false command detection, judgment results regarding other misrecognitions, and statistical information such as the misrecognition rate, which can be used as appropriate. Furthermore, if a command is detected, the format check unit 7 may delete or modify the voice input command that matches the detected command, depending on the user's operation. In other words, if a conflict occurs with a voice input command, the format check unit 7 may resolve the issue by deleting the voice input command or changing the wording of the voice input command.
[0036] In any case, the format checking unit 7 performs at least one of the following: a check based on guidance and a check using speech synthesis.
[0037] In this case, when performing a format check based on guidance, it is assumed that guidance is set in input procedure 3a. The format check unit 7 checks whether there is any discrepancy between the wording of the guidance and the estimated input format. For example, if the wording of the guidance is "○○ time" and the input format is not a time format, it will be determined that there is a discrepancy.
[0038] Figure 7 shows an example of a rule table for guidance text conditions and corresponding input format rules. The "Guidance" column contains regular expressions that match parts of the guidance text as guidance conditions. The "Input Format" column contains the content that the input format should satisfy when the guidance conditions are met.
[0039] The format check unit 7 checks whether the set guidance text matches the guidance in the rule table for each input procedure corresponding to the selected range. If there is a matching guidance row, it checks whether the input format of that row matches the estimated input format. If they do not match, it notifies the input procedure creator accordingly.
[0040] In the case of format checking using speech synthesis, the format checking unit 7 checks whether or not input value examples that are difficult to recognize are generated for the estimated input format by performing speech recognition on the speech waveform data of the input value examples generated from the estimated input format.
[0041] To elaborate, the format check unit 7 converts the text of each input value example into audio waveform data using speech synthesis for each input procedure. The format check unit 7 can perform speech synthesis at any time, such as when a predetermined number of input value examples have been generated or when a predetermined operation by the user has been received. Furthermore, the format check unit 7 performs speech recognition on the obtained audio waveform data using the speech recognition dictionary generated by the dictionary generation unit 8 for the corresponding input format. At this time, the format check unit 7 also operates speech recognition for command recognition, which will be used in the recognition control unit 11 described later. If the speech recognition for command recognition is different from the speech recognition for input values, the format check unit 7 recognizes the input value examples generated using both speech recognition processes. When the speech recognition for input values is used as the speech recognition for command recognition, the generated speech recognition dictionary also includes the dictionary for command recognition, and speech recognition is performed using that dictionary. This provides the speech recognition results for the generated input value examples. Furthermore, the format check unit 7 compares the speech recognition result with the generated input value example, and if they all match, it determines that the estimated input format is valid.
[0042] The dictionary generation unit 8 generates a speech recognition dictionary based on the estimated input format. For example, the dictionary generation unit 8 generates a speech recognition dictionary capable of recognizing speech utterances of the input format set in the input format field of input procedure 3a. Here, the speech recognition dictionary is assumed to contain grammar and be uniquely generated from the input format. The dictionary generation unit 8 also generates a dictionary for command recognition in parallel.
[0043] The speech recognition unit 9 is controlled by the format checking unit 7 and the recognition control unit 11, and converts speech waveform data into text using a speech recognition dictionary. For example, the speech recognition unit 9 is assumed to use speech recognition that uses a speech recognition dictionary containing a predetermined grammar to recognize utterances for speech input. In this type of speech recognition, speech waveform data of utterances that conform to the grammar is recognized, and speech waveform data of other utterances is rejected. The format of the speech recognition dictionary depends on the speech recognition technology used, and may be a grammar written in text, a binary dictionary, or an object held in memory. Furthermore, the speech waveform data may be obtained from the user's utterance, or it may be obtained by text-to-speech synthesis.
[0044] In this embodiment, in addition to inputting items by voice, the system can also execute commands other than item input, such as undoing input, proceeding to the next input step, or ending input, by speaking them aloud. At this time, the format check unit 7 and the recognition control unit 11 also perform voice recognition for command recognition.
[0045] Command recognition can use the same speech recognition method as described above, or a different speech recognition method. For example, instead of grammatical speech recognition, a speech recognition method that detects utterances of keywords included in a keyword list can be used. Speech recognition is not limited to these methods; general speech recognition can also be used. In this case, the speech recognition dictionary should consist of a word dictionary containing the spelling and pronunciation of the words to be recognized, and a grammar dictionary for determining whether or not to reject the recognition result.
[0046] Furthermore, if the speech recognition for command recognition is different from the speech recognition used for inputting items, the speech recognition unit 9 recognizes the input value examples generated using both speech recognitions. Also, if the same speech recognition used for inputting items is used for command recognition, the speech recognition unit 9 includes the dictionary for command recognition in its speech recognition dictionary and performs speech recognition using that dictionary.
[0047] The speech synthesis unit 10 is controlled by the format checking unit 7 and the recognition control unit 11, and converts text into speech waveform data. The speech synthesis unit 10 can play the text as speech from the speaker by sending the speech waveform data to the speaker. The speech synthesis unit 10 is used when checking the validity of the input format estimated by the format checking unit 7, when the recognition control unit 11 plays guidance at the start of each input procedure, and when repeating the input content.
[0048] The recognition control unit 11 performs speech recognition using a speech recognition dictionary in response to the user's utterance and inputs the resulting speech recognition text into the input field according to the identifier that identifies the input field of the input item. For example, the recognition control unit 11 performs speech recognition on the user's utterance according to the input procedure and inputs the recognition result into the input field of the form data 2a. Alternatively, the recognition control unit 11 may perform speech recognition using a speech recognition dictionary generated from the input format of the input procedure and input the recognition result into the input field of the input item of the input procedure.
[0049] For example, the recognition control unit 11 displays the form data on the display unit 12 and retrieves one input procedure from the input procedure storage unit 3 according to the order of the input procedure. The recognition control unit 11 uses the speech synthesis unit 10 to synthesize the guidance content set in the input procedure. The recognition control unit 11 uses the speech recognition dictionary (and command recognition dictionary) generated from the input format set in the input procedure to perform speech recognition using the speech recognition unit 9. If a command is recognized, the recognition control unit 11 executes the content of the corresponding command. If something other than a command is recognized, the recognition control unit 11 synthesizes the speech recognition result and repeats it, then compares the recognition result with the reference value set in the input procedure, and if it is within the reference value, it inputs the recognition result into the input target item. If it is outside the reference value, the recognition control unit 11 notifies the user of this fact via speech synthesis and prompts them to re-enter the information, or forces them to enter it.
[0050] The display unit 12 displays data corresponding to the processing of the information processing device 1. For example, the display unit 12 displays the report data 2a and displays the estimated input format in the input fields included in the report data 2a. Alternatively, the display unit 12 may generate one or more example input values from the estimated input format and display the generated example input values in the input fields included in the report data 2a. Alternatively, the display unit 12 may generate a range of values that can be entered into the input fields from the estimated input format and display the generated range in the input fields included in the report data 2a. Alternatively, the display unit 12 may display the report data 2a and, based on the input procedure 3a, superimpose the order in which values are entered into multiple input fields included in the report data 2a.
[0051] Specifically, for example, the display unit 12 includes a microphone for recording the user's voice, a speaker for playing back the voice, a display for displaying stored data and various information, and a control circuit for controlling the content displayed on the display. The display unit 12 displays the report data 2a on the display or the like. The display unit 12 may also superimpose the data from the input procedure 3a onto the report data 2a as appropriate. The display unit 12 may be called an output unit because it includes a speaker and a display, or it may be called an input / output unit because it further includes a microphone.
[0052] Next, the operation of the input procedure creation process and the voice input process by the information processing device configured as described above will be explained using Figures 8 to 20. In the following explanation, the input procedure creator creates an input procedure list other than the input format on partially completed form data, and then the information processing device 1 executes the input procedure creation process to complete the input procedure list. Furthermore, the information processing device 1 copies the completed input procedure list to the form data that has not yet been entered, thereby enabling voice input processing for the form data that has not yet been entered.
[0053] (Input procedure creation process: Figure 8) In step ST1, the information processing device 1 creates an input procedure list as input procedure 3a, with all settings except the input format already configured, in response to the input procedure creator's actions. In other words, in step ST1, input procedure 3a has no input format set.
[0054] In step ST2, the range selection unit 5 determines whether there are any remaining input procedures 3a for which an input format has not been set. If not, the input procedure creation process ends. On the other hand, if, as a result of the determination in step ST2, there are still unset input formats remaining, the range selection unit 5 proceeds to step ST3.
[0055] In step ST3, as shown in Figure 9, the range selection unit 5 selects a range 5a on the report data 2a that includes multiple input fields corresponding to one or more input target items, in response to the operation of the cursor c by the input procedure creator, and presses the displayed input format estimation button 5b. Note that the range selection unit 5 may also select a range that includes only one input field.
[0056] In step ST4, the input value example acquisition unit 4 has the identifiers of the input fields within the selected range as input target items (O1, O2, ..., O n ) are extracted from the input procedure list. However, n represents the number of procedure O within the selected range. In the case of Figure 9, the extracted procedure(O1, O2, ..., O n ) are the three rows (n=3) corresponding to the input field identifiers B1, C1, and D1 indicated by the input target items in input procedure 3a.
[0057] In step ST5, the input value example acquisition unit 4 obtains input value examples (E1, E2, ..., E) for items within the selected range. m The following input values are obtained from the form data 2a. However, m represents the number of input value examples within the selected range. In the case of Figure 9, the extracted input value examples (E1, E2, ..., E m The two possible times are "10:30" and "12:00" (m=2).
[0058] In step ST6, the form estimation unit 6 calculates the acquired procedure (O1, O2, ..., O n) and an example of its input value (E1, E2, ..., E m Based on this, a form estimation process is performed to estimate the input format. The input format of input step 3a is updated (set) based on the result of the form estimation process. Details of the form estimation process in step ST6 will be described later.
[0059] In step ST7, the format check unit 7 performs a format check on the estimated and updated input format and obtains a list of misrecognitions to be notified, determining whether notification is required or not, and if notification is required. Details of the format check process in step ST7 will be described later.
[0060] In step ST8, the format check unit 7 determines whether the format check result requires notification or not. If not, the unit proceeds to step ST11. On the other hand, if the result of the determination in step ST8 indicates that notification is required, the unit proceeds to step ST9.
[0061] In step ST9, the format check unit 7 notifies the display unit 12 of the misrecognition list as the format check result if notification is required. The display unit 12 displays the format check result.
[0062] Here, we will explain an example of a display that includes an alert in the format check results using Figures 10 and 11.
[0063] If a command is misdetected as a result of speech recognition of an example input value, speaking that example input value will result in the command being incorrectly recognized. In this case, for example, the system may have obtained a recognition result that recognizes a command from the example input value, or it may have used a command recognition dictionary simultaneously with the example input value to recognize the command, and the command was recognized as the result. In either case, the display unit 12 displays on the screen a message indicating that "there is a possibility of misdetection of the command," linked to the input item where the misdetection occurred, the input format (including words in the word list in the case of words), and the example input value. This is particularly useful because, in the case of "words," it may be possible to prevent misdetection of the command by changing the word being entered.
[0064] Figure 10 shows an example of how a dialog box 7a notifies the user of the format check results indicating a false positive for a command. In this example, words such as "No abnormality" and "Abnormality detected" conflict with the command "Above," and the dialog box notifies the user that the command has been falsely detected. The input field in which the false positive occurred and the estimated input format are also notified. The method of displaying the notification content is not limited to this; for example, the input field in which the false positive occurred could be highlighted.
[0065] At the bottom of dialog 7a in Figure 10, buttons bt1 to bt3 are displayed: "Change Input Format," "Change Command," and "Keep as is." By operating any of these buttons, you can choose how to deal with the possibility of false command detection.
[0066] When the input procedure creator operates the button bt1 indicating "Change input format", the display unit 12 prompts the user to change the input format by moving the display to the corresponding input procedure section of the procedure retention sheet, displaying a dialog box for changing the input format, etc.
[0067] Furthermore, if the input procedure creator operates button bt2 indicating "Change Command", the display unit 12 automatically changes the wording of the misdetected command. For example, multiple candidates can be provided for a single command, and if a misdetection occurs, the command can be changed to one selected from the other candidates, or it can be changed to wording with a similar meaning by referring to word2vec or a thesaurus. If a change is made, the format check unit 7 performs a format check again to check whether the same problem occurs with the changed command. If the input procedure creator selects button bt3 indicating "Leave as is", the display unit 12 does nothing.
[0068] On the other hand, even if the command misdetections described above do not occur, misrecognition is still possible. For example, both the number "9" and the letter "Q" are pronounced "kyuu," and the system cannot distinguish between them. In addition, some words may be difficult to recognize.
[0069] Therefore, the format check unit 7 may compare the speech recognition result with the generated input value examples and calculate the speech recognition rate. Since such input value examples are not always generated, the input value generation may be configured to always include examples that are prone to problems, such as "9" and "Q". If the speech recognition rate is lower than a predetermined threshold, the input procedure creator will be notified accordingly, as in the case of command misdetection. In addition to the fact that the recognition rate has decreased, the specific locations where misrecognition occurred can also be obtained from comparing the speech recognition result with the generated input value examples, so the details of the problem may be notified at the same time.
[0070] Figure 11 shows an example of a display that notifies the user of a format check result indicating a decrease in recognition rate via dialog 7b. In this example, instead of button bt2 indicating "Change command," button bt4 indicating "Re-select range" is displayed. Button bt4 is displayed because, in the case of a decrease in recognition rate, changing the selection range changes the input format and may mitigate the decrease in recognition rate. In this way, the format check unit 7 eliminates the need for the voice input procedure creator to actually speak and verify the operation. Furthermore, by displaying an alert, the problem can be communicated to the input procedure creator, prompting improvements.
[0071] When performing a format check all at once after estimating all input formats, such alerts can be displayed by highlighting the problematic selection and showing dialogs 7a and 7b below it, as shown in Figures 10 and 11.
[0072] Alternatively, the voice input procedure creator may perform a format check each time they select a range or estimate the input format. In this case, if there are any problems after the format check, an alert should be displayed in the same way. The method of notifying the alert is not limited to displaying dialogs 7a and 7b; any method can be used, such as displaying it in the status bar or notifying by voice.
[0073] Subsequently, in step ST10, the display unit 12 takes action on the format check results in accordance with the input procedure creator's actions. Alternatively, the display unit 12 may leave the format check results as they are, depending on the input procedure creator's actions.
[0074] In step ST11, the display unit 12 overlays the input format and example input values onto the report data according to the input procedure creator's operation. During the overlay display, the input procedure creator checks whether there are any problems with the set input format and whether there are any input procedures for which the input format has not been set.
[0075] For example, the display unit 12 can display a menu or the like, and then perform the overlay display by selecting a function from the menu.
[0076] Figure 12 shows an example of a display where the input format during the input procedure is superimposed on the form data 2a. The display unit 12 displays menu 12a at the top of the form data 2a. Input items may be displayed, for example, by color.
[0077] The input format may display the contents of input procedure 3a in the input procedure storage unit 3 as is, or it may be converted and displayed in a way that is easy for the input procedure creator to understand and view. For example, in the example in Figure 12, the input format for time and numbers is displayed as is, but for word input, instead of group names, a word list is displayed with words connected by "|".
[0078] Figure 13 shows an example of generating example input values from the input format and overlaying them onto the report data 2a. Overlaying the example input values allows for a more intuitive understanding of what values can be entered than simply displaying the input format itself. In both cases, input items for input procedures where the input format is not set are indicated accordingly. In the example in Figure 13, the item's color is changed and an icon is displayed.
[0079] In addition to the input format and input value examples, guidance, the order of input, etc. may also be superimposed and displayed so that it is possible to check whether the entire input procedure is correctly set. As the display of the order of input, as shown in FIGS. 14 and 15, the order number may be displayed, or the order may be displayed with an arrow.
[0080] After step ST11, the information processing apparatus 1 returns to step ST2 and repeatedly executes steps ST2 to ST11 until all the input formats are set.
[0081] Next, the details of the format estimation process in step ST6 will be described.
[0082] (Details of the format estimation process in step ST6) The format estimation process varies depending on whether concatenation is allowed in the input format.
[0083] (When concatenation is not allowed: FIG. 16) FIG. 16 shows a flowchart of the format estimation process when there is no concatenation.
[0084] In steps ST601 to ST603, the format estimation unit 6 estimates the major classification G of the input format that matches each of the input value examples (E1, E2,..., E m ) within the selection range (i ≤ m). i For example, when the input value example E i consists only of a numerical value such as "-12.34", a decimal point, and a sign, it can be seen that this is a numerical value. FIG. 17 shows an example of such a rule for classifying major categories. In this example of the rule, a regular expression is described in the conditions. In descending order of the priority of such a table, when the input value example E
[0085] [[ID=3�]] i matches the regular expression of the condition, it can be regarded as the corresponding major classification G i i i i .
[0086] If it does not match any of them, the input value example E i represents the major classification G of the lowest "word".i It will be distributed to the appropriate category. For example, input value example E i If the value is "2023 / 03 / 07", it matches the regular expression "\d\d\d\d / \d?\d / \d?\d", so it is determined to be a "date". This is not the only way to categorize, and the rules are not limited to this condition. Alternatively, instead of having such rules, you could write a program that performs a similar process.
[0087] In step ST604, if the format estimation unit 6 determines that the OR input format is not acceptable, it proceeds to step ST605. Whether or not the OR input format is acceptable is predetermined.
[0088] In step ST605, the format estimation unit 6 determines the major classification of the input format (G1, G2, ..., G m The formal estimation unit 6 selects one major category G from among the selected major category G. The formal estimation unit 6 stores examples of input values that match the selected major category G in the memory of the data storage unit 2 or other memory.
[0089] To clarify, the selected range contains multiple input values (e.g., E1, E2, ..., E m If there are multiple input value examples corresponding to multiple items, or if there are multiple input value examples corresponding to a single input item. In either case, the major categories of input formats that match each input value example (G1, G2, ..., G m ) need to be integrated and one major category G selected.
[0090] The integration of input formats may encompass all input formats. Alternatively, the narrowest input format may be adopted, and input value examples that do not match the adopted input format may be treated as invalid examples.
[0091] When determining the input format to encompass all, if the major categories are the same, the details of the input format are integrated. If the major categories are different, the format is aligned to the most frequent major category, and for input value examples that do not fit that major category, the input format is re-estimated to match the selected major category.
[0092] For example, the input value "10:30" looks like a "time" on its own, but it can also be considered an "alphanumeric" string. Therefore, if the most frequent major category is "alphanumeric," then it should be treated as an "alphanumeric" string, and the subsequent detailed format should be estimated accordingly. If the major category cannot be determined and re-estimation is impossible, one major category G may be selected, and input value examples that do not match it may be treated as invalid. Alternatively, if OR is also used as an input format, the result of concatenating the OR values may be used as the estimated major category.
[0093] In step ST606, the formal estimation unit 6 determines the input value examples (E1, E2, ..., E) that match the selected major category G. m The details of the input format are estimated and integrated for each of the following.
[0094] To elaborate, once the main category G is determined, the details of the input format are estimated using example input values corresponding to each main category G. For example, if the main category G is "Numerical (decimal)", details such as whether there are signs like + or -, the number of characters before the decimal point (i.e., the number of digits in the integer part), and the number of characters after the decimal point (i.e., the number of digits in the decimal part) are determined.
[0095] Furthermore, if the main category G is "alphanumeric," an alphanumeric pattern can be obtained by replacing each character with 'A' if it's a letter, 'D' if it's a number, and 'S' if it's any other symbol.
[0096] If the main category G is "Word," the example input value will be used as the word representation. For example, add a new group to word list 2b shown in Figure 4, set a new group name, and set the example input value as the word representation. The reading and speech content will be automatically generated and set using the method described above. The detailed input format should be "word_(group name)".
[0097] In step ST607, the form estimation unit 6 uses the following example input values (E1, E2, ..., E m The details of the input format are adjusted according to the number of ). This results in a single input format.
[0098] To elaborate, the details estimated for each input value example are integrated for each major category. When determining the input format to encompass everything, the integration of input format details is done, for example, as follows.
[0099] In the case of numerical values, for example, if the main category is "Numerical Value (Decimal)" and the estimated individual input formats are "Decimal_3_2 (3 integer digits, 2 decimal digits)" and "Decimal_3_1 (3 integer digits, 1 decimal digit)", the decimal part can be either one or two digits. Therefore, the combined input format becomes "Decimal_3_1~2 (3 integer digits, 1~2 decimal digits)".
[0100] In the case of alphanumeric characters, for example, if the main category is "alphanumeric" and the estimated individual input formats are "alphanumeric_AAADDSA" and "alphanumeric_AADDDS", the third character can be either a letter or a number, and the last letter is optional. Therefore, the combined input format becomes "alphanumeric_AA$DDSA?".
[0101] In the case of words, for example, if the main category is "Words" and the estimated individual input formats are "Word_1" and "Word_2", a new group is created by integrating the groups in the word list, and this group is used as the input format. That is, a new group "3" is created by integrating the word list of group "1" and the word list of group "2", and all the words from "1" and "2" are added to word group "3". The integrated format used is then "Word_3". In this case, if there are no longer any input formats that refer to word groups "1" or "2", the words from groups "1" and "2" are removed from word list 2b.
[0102] By integrating in this way, we can obtain the details of the final input format. Note that this is not the only integration method; we could also select the single input format with the most votes.
[0103] On the other hand, in step ST604, if the format estimation unit 6 determines that an OR of the input format is acceptable, it proceeds to step ST608.
[0104] In step ST608, the form estimation unit 6 takes the OR of each major category. For example, the form estimation unit 6 takes the major categories of the input format (G1, G2, ..., G m By taking the unique elements (removing duplicates) and connecting the resulting major categories with OR, we can obtain the OR major categories G1|G2|…|G k We obtain (k≦m). Also, the formal estimation unit 6 determines the major classification of OR G1|G2|…|G k Each of the major categories G j The corresponding example input values are stored in the memory of the data storage unit 2, etc. (j ≤ k).
[0105] After step ST608, the formal estimation unit 6 determines the major category G which will be an element of OR. j Each time, the loop process is executed between steps ST609 and ST613. This loop process includes steps ST610 to ST611 and step ST612, as described above.
[0106] In step ST610, the formal estimation unit 6 determines the major category G j The details of the input format are determined from the corresponding input value examples. Step ST610 is executed in the same way as step ST606.
[0107] In step ST611, the format estimation unit 6 adjusts the details of the input format according to the number of input value examples used. Step ST611 is executed in the same way as step ST607.
[0108] In step ST612, the formal estimation unit 6 determines the major category G corresponding to the input value example. j Replace this with the details of the input format. This will change the main category G j One corresponding input format is obtained. The loop processing of the above steps ST609~ST613 is major category G j It is executed each time.
[0109] After step ST607 or ST613, in step ST614, the form estimation unit 6 performs each step of input procedure 3a (O1, O2, ..., O nUpdate (set) the input format to the obtained input format.
[0110] Furthermore, if the user selects only a small range in the range selection unit 5 and there are few input value examples, determining the input format based solely on those input value examples may result in an overly detailed input format. As mentioned earlier, an overly detailed input format may impair the convenience of the user performing voice input. For example, if only one input value example, "12.34," is provided, the input format will be "decimal_2_2."
[0111] However, depending on the voice inputter, the "0" in "12.30" might be omitted, resulting in a number like "12.3". It would be ideal if the example input values included "12.30" and "12.3", but this is not always the case. Conversely, there may be cases where only "12.3" is given as an example input value, but "12.30" is actually also required.
[0112] To accommodate such examples, if there are few input value examples obtained within the selected range, the input format can be relaxed. That is, for decimal numbers, include variations with one less digit or one more digit (e.g., "decimal_M_N" → "decimal_(M-1)~(M+1)_(N-1)~(N+1)"), and for alphanumeric numbers, replace the letter-only pattern part with a letter or number pattern (e.g., "alphanumeric_AAA" → "alphanumeric_$$$"), and so on.
[0113] Once the input format has been estimated as described above, update the input format of the input procedure currently under consideration (the input procedure that uses the range-selected items as input items).
[0114] This completes the formal estimation process in step ST6 when concatenation is not permitted.
[0115] (If connection is permitted: Figure 18) Figure 18 shows a flowchart of the formal estimation process when connections are present.
[0116] In steps ST621 to ST624, the form estimation unit 6 calculates the input values for the selected range (E1, E2, ..., E m For each of the above, a loop process is executed. This loop process includes steps ST622 to ST623.
[0117] In step ST622, the formal estimation unit 6 determines the input value example E i Split by a space, input value example E i parts e1, e2, ..., e il The subscript "il" is the input value example E. i This represents the number of parts e included in the expression. For example, input value example E i If we divide it into three parts, then il = 3.
[0118] In step ST623, the form estimation unit 6 determines the input value example E i parts e1, e2, ..., e il For each of these, the corresponding major category G1 i G2 i ,…,G il i Select this option.
[0119] To clarify, if the input format allows not only simple 'numbers' or 'words' but also combinations thereof (concatenation or OR), then first, consider the example input value E. i Separate the parts with spaces, e1, e2, ..., e il The major category G1 that matches i G2 i ,…,G il i Search for it.
[0120] For example, input value example E i If it is "2023 / 03 / 07 Tue 13:00", then the part e1 "2023 / 03 / 07" is in the major category G1 of "Date". i And the "fire" part e2 is in the main category G2 of the 'words'. i And the part "13:00" e il This is the major category G of 'Time'. il iThis matches. If there is only one example input value, you can select the one with the highest priority among the multiple major categories, and then, as described above, estimate the detailed format for each major category.
[0121] In step ST624, the formal estimation unit 6 determines the selected major category G1 i G2 i ,…,G il i The concatenated data is added to the list of major category concatenations. For example, 'Date, Word, Time', which is a concatenation of 'Date', 'Word', and 'Time', is added to the list of major category concatenations. Major Category Concatenation G1 i G2 i ,…,G il i Example input value E for the delimiter i As for the part, e1 is "2023 / 03 / 07" from "Date", e2 is "Fire" from "Word", and e is "13:00" from "Time". il This is obtained. Note that "major category linking" does not necessarily require linking, so it can also be called "combination of major categories."
[0122] The formal estimation unit 6 processes the above steps using input value example E i The loop process from steps ST621 to ST624 is executed, repeating the process as many times as necessary.
[0123] Subsequently, in steps ST625 to ST629, the formal estimation unit 6 executes a loop process if there are two or more connections in the list of major category connections. This loop process includes steps ST626 to ST628.
[0124] In step ST626, the formal estimation unit 6 extracts two major category links from the list of major category links. The formal estimation unit 6 also extracts the two major category links (G1 i G2 i ,…,G il i ),(G1 j G2 j ,…,G jl jA difference algorithm is applied to ).
[0125] For example, if there are multiple input value examples, the content of the major category linkage may differ among the input value examples. For instance, if the input value examples are "A12C Black 12.3", "A12M White 45.6", and "B23F Red", there will be three candidates for major category linkage: "Hexadecimal, Word, Decimal", "Alphanumeric, Word, Decimal", and "Hexadecimal, Word". From these, the major category is selected to maximize the number of matches.
[0126] Two of the candidates are selected from the list of linked major categories, and the problem is treated as a difference calculation problem involving OR in the subsequences. The order of the major categories can then be determined using existing difference algorithms. For example, algorithms used in diff, which find differences between sequences by calculating edit distance, LCS (Longest Common Subsequence), and SES (Shortest Edit Script), can be used as appropriate.
[0127] The result of the diff algorithm is one of the following for each character: "identical," "added," "deleted," or "replaced." Using the diff algorithm on each major category (or their OR), select the one with the smallest distance.
[0128] The substitution cost between major categories should be 0 if one is contained within the other (i.e., they are considered the same, such as 'decimal|alphanumeric' and 'alphanumeric', or 'hexadecimal' and 'alphanumeric'), and 1 if they are not contained within each other.
[0129] Furthermore, applying a difference algorithm to the two major linked categories, "hexadecimal, word, decimal" and "alphanumeric, word, decimal," yields "hexadecimal and alphanumeric substitution, word, decimal" as difference information.
[0130] Furthermore, in step ST627, the form estimation unit 6 determines the two major classifications that have been extracted (G1 i G2 i ,…,Gil i ),(G1 j G2 j ,…,G jl j Remove ) from the list of linked major categories.
[0131] In step ST628, the formal estimation unit 6 absorbs the difference using the difference information obtained as a result of applying the difference algorithm, and concatenates one major category (G1 i,j G2 i,j ,…,G i,jl i,j Create ).
[0132] For example, the format estimation unit 6 absorbs the differences in the aforementioned difference information "hexadecimal and alphanumeric substitution, words, decimal numbers" and converts the "substitution," "addition," and "deletion" parts into combinations of broad categories. For example, in "hexadecimal and alphanumeric substitution, words, decimal numbers," if the substitution part is changed to a broader inclusion relationship, it becomes "alphanumeric numbers, words, decimal numbers." In this way, "substitution of ○○ and ×× (○○ is included in ××)" with an inclusion relationship should be written as "××."
[0133] If you encounter a substitution between XX and YY without an inclusion relationship, you can use OR, i.e., XX|YY.
[0134] Additionally, "adding ○○" and "deleting ××" can be treated as optional major categories, i.e., "(○○)" and "(××)".
[0135] At this point, it's important to record which part of which input value example matches which major category.
[0136] Subsequently, the formal estimation unit 6 performs the linking of the created major classifications (G1 i,j G2 i,j ,…,G i,jl i,j Add ) to the list of major category links.
[0137] The form estimation unit 6 executes the loop process of steps ST625 to ST629 so as to repeat the above process until only one concatenation of major classifications remains.
[0138] After that, in steps ST630 to ST633, the form estimation unit 6 executes a loop process to obtain the details of the input format for each major classification G i , G2 i , …, G il i ) in the remaining concatenation of major classifications. The loop process includes steps ST631 to ST632. i i
[0139] In step ST631, the form estimation unit 6 determines the details of the input format from the subset (e i ) of the input value examples Ei corresponding to the major classification G i1 , e i2 , …, e im ).
[0140] i In the case of the above example, the remaining concatenation of major classifications G i is 'alphanumeric, word, (decimal number)'. Here, the input value examples E i that match the major classification G i of 'alphanumeric' are "A12C", "A12M", "B23F", the input value examples E i that match the major classification G i of 'word' are "black", "white", and the input value examples E i that match the major classification G i of 'decimal number' are "12.3", "45.6". After that, for each major classification G i , the detailed format may be estimated from the input value examples E
[0141] In step ST632, the form estimation unit 6 determines the subset (e i ) of the input value example E i1 , e i2 , …, e im) number (Example input value E used) i The details are adjusted by the number of ( ). After that, the formal estimation unit 6 determines the major category G based on the obtained details. i Replace it.
[0142] The formal estimation unit 6 processes the above steps for each major category G i The loop process from steps ST630 to ST633 is executed, repeating each time.
[0143] In step ST634, the format estimation unit 6 sets the input format obtained as described above as the input format for the corresponding multiple input procedures. In step ST634, if there are few input value examples, the input format may be relaxed, as in step ST614. In any case, this completes the format estimation process in step ST6 when concatenation is permitted.
[0144] (Format checking process) The following format checking process describes format checking using speech synthesis. As explained in format checking section 7, the format checking process generates example input values from the input format, generates speech data using speech synthesis, performs speech recognition using the speech recognition dictionary generated from the input format, and checks the format by seeing if misrecognition occurs.
[0145] Figure 19 shows a flowchart of the format check process.
[0146] In step ST701, the format checking unit 7 generates a speech recognition dictionary from the input format it initially focuses on, using the dictionary generation unit 8.
[0147] In step ST702, the format check unit 7 generates multiple input value examples from the estimated input format. For example, for parts of the input value examples that can be generated from the input format, where multiple candidates can be generated, the candidates can be selected using random numbers. For example, in the case of "numerical value", if the input format is "decimal_2~4_1", the number of digits in the integer part of the input value example can be randomly selected from 2 to 4, and the digits of each digit can be randomly determined from 0 to 9 (however, the first digit can be from 1 to 9).
[0148] In the case of "alphanumeric," if the input format is "alphanumeric_AA$D?", the first and second characters of the example input value are selected from "A-Za-z", the third character is selected from "A-Za-z0-9", and the fourth character is determined by a random number to see whether or not to include it. If you want to include a random number for the fourth character of the example input value, you should select one character from "0-9". Similarly, multiple example input values can be generated using random numbers for other major categories.
[0149] In step ST703, the format check unit 7 creates an empty misrecognition list. The empty misrecognition list is a list that includes the type of misrecognition, an example input value, and detailed information. For example, the types of misrecognition can be command misdetection and misrecognition as appropriate. For detailed information, for example, if the type is command misdetection, the command in question can be used. Also, for example, if the type is misrecognition, the detailed information can be recognition result and error location as appropriate.
[0150] Subsequently, in steps ST704 to ST711, the format checking unit 7 executes a loop process to perform a format check by speech recognition for each of the generated input value examples. This loop process includes steps ST705 to ST710.
[0151] In step ST705, the format check unit 7 performs speech synthesis for each generated input value example and obtains speech data as a result of the speech synthesis.
[0152] In step ST706, the format check unit 7 performs speech recognition on the speech data of the speech synthesis result using the generated speech recognition dictionary. At this time, the format check unit 7 uses the speech recognition dictionary to perform speech recognition of the input value example and speech recognition of the command. A first dictionary for the input value example and a second dictionary for the command may be used as the speech recognition dictionary. When using the first and second dictionaries, the format check unit 7 performs speech recognition of the input value example and speech recognition of the command separately. Alternatively, a dictionary for both the input value example and the command may be used as the speech recognition dictionary. When using a dictionary for both, the format check unit 7 performs speech recognition of the input value example and speech recognition of the command simultaneously. There are three possible recognition results in step ST706: (1) when the input value example is recognized, (2) when the input value example is not recognized and the command is incorrectly detected, and (3) when there is a misrecognition other than incorrect command detection. In case (1) above, speech recognition is successful. In cases (2) or (3) above, speech recognition has failed. Furthermore, if the command in (2) above is misdetected, that is, if the command is obtained as the recognition result, speaking the example input value will result in the command being incorrectly recognized.
[0153] In step ST707, the format check unit 7 determines whether the text of the speech recognition result matches the input value example. If they match, speech recognition is successful, and the process proceeds to step ST711. If the result is negative, speech recognition is unsuccessful, and the process proceeds to step ST708.
[0154] In step ST708, the format check unit 7 determines whether the reason for the failure of speech recognition was a false command detection. If it was a false command detection, it proceeds to step ST709; otherwise, it proceeds to step ST710. According to step ST708, when speech recognition fails in step ST707, the situation can be classified, such as whether the failure was due to a false command detection.
[0155] In step ST709, the format check unit 7, having determined that a command was recognized as a result of speech recognition, adds information regarding the command misdetection to the misrecognition list and proceeds to step ST711.
[0156] In step ST710, the format check unit 7 adds information about the misrecognition to the misrecognition list because the recognition result of the speech recognition is a misrecognition other than a command misrecognition, and proceeds to step ST711.
[0157] In step ST711, the format check unit 7 performs the processing in steps ST705 to ST710 for all generated input value examples, and then terminates the loop processing in steps ST704 to ST711.
[0158] In step ST712, the format check unit 7 determines whether or not there is information about a command misdetection in the misrecognition list. If there is information about a command misdetection, the unit proceeds to step ST714; otherwise, it proceeds to step ST713.
[0159] In step ST713, the format check unit 7 calculates the misrecognition rate by dividing the number of misrecognitions in the misrecognition list by the number of input value examples, and determines whether the misrecognition rate is above a threshold. If the misrecognition rate is above the threshold, the format check unit 7 proceeds to step ST714; otherwise, it proceeds to step ST715.
[0160] In step ST714, the format check unit 7 determines that there is information about a command misrecognition or that the misrecognition rate is high and therefore notification of the format check result is necessary, and notifies the display unit 12 of the misrecognition list corresponding to the format check result and terminates the process.
[0161] In step ST715, the format check unit 7 terminates processing because it determines that notification of the format check result is unnecessary if the misrecognition list is empty or the misrecognition rate is low.
[0162] (Voice input processing) Based on the input procedure 3a obtained through the input procedure creation process, the information processing device 1, together with another voice inputter, executes the voice input process as shown in Figure 20.
[0163] In step ST21, the recognition control unit 11 of the information processing device 1 adds a flag (=execution flag) indicating whether or not it has been performed for each step number of the input procedure 3a, and sets all execution flags to OFF.
[0164] In step ST22, the recognition control unit 11 determines whether there is an input procedure 3a with the execution flag of the procedure number set to OFF. If not, it terminates the voice input processing. On the other hand, if, as a result of the determination in step ST22, there are still procedure numbers with the execution flag set to OFF in input procedure 3a, the recognition control unit 11 proceeds to step ST23.
[0165] In step ST23, the recognition control unit 11 retrieves the procedure corresponding to the earliest procedure number among the procedure numbers for which the execution flag is OFF from input procedure 3a.
[0166] In step ST24, the dictionary generation unit 8 generates a speech recognition dictionary from the input format set in the retrieved procedure, under the control of the recognition control unit 11.
[0167] In step ST25, the speech synthesis unit 10, under the control of the recognition control unit 11, synthesizes the guidance included in the extracted procedure into speech and plays the resulting audio data through the speaker.
[0168] In step ST26, the speech recognition unit 9 waits for the voice inputter to speak, under the control of the recognition control unit 11. At this time, the voice inputter speaks as appropriate.
[0169] In step ST27, the speech recognition unit 9 performs speech recognition using the speech recognition dictionary in response to the speech inputter's utterance, under the control of the recognition control unit 11.
[0170] In step ST28, the recognition control unit 11 determines whether or not a command was recognized as a result of speech recognition. If a command is recognized, the unit proceeds to step ST29; otherwise, it proceeds to step ST30.
[0171] In step ST29, the recognition control unit 11 executes the recognized command and returns to step ST22.
[0172] In step ST30, the recognition control unit 11 synthesizes the recognition result text, which does not contain any commands, into speech and plays the resulting speech data through the speaker. As a result, the content spoken by the voice inputter is repeated through the speaker.
[0173] In step ST31, the recognition control unit 11 inputs the value represented by the recognition result text into the input field of the report data 2a, based on the identifier of the input target item included in the extracted procedure.
[0174] In step ST32, the recognition control unit 11 turns on the flag for executing the procedure and returns to step ST22.
[0175] The recognition control unit 11 will then execute all the steps included in the list of input steps 3a until all the execution flags are turned ON.
[0176] The above voice input processing is just one example, and a voice recognition dictionary is generated each time a step in the procedure number is performed, but this is not the only way. For example, the voice recognition dictionary could be generated all at once at the beginning, and then the necessary dictionary could be extracted and used during voice recognition. Also, repetition is not required. Alternatively, after the recognition result is repeated, the system could ask the voice inputter for confirmation such as "Is this OK?", and only execute the input if OK; if NG, the procedure could be restarted from the utterance in step ST26. In this case, it is possible to prevent the input of incorrect values due to misrecognition.
[0177] As described above, according to the first embodiment, the input value example acquisition unit 4 acquires, for each item, one or more items and input value examples, which are information about the values of the input fields for one or more items, from the form data 2a, which includes the voice input fields. The format estimation unit 6 estimates the input format of the input fields based on the input value examples. In this way, by estimating the input format of the input fields based on the input value examples of the input fields, an appropriate input format can be determined even when there is no voice recognition expert present.
[0178] To elaborate, conventional methods pre-define the format of input values and perform speech recognition based on that format. Generally, performing speech recognition based on the input format allows for rejecting recognition results for utterances that do not conform to the set input format, and rejecting recognition results that are mistakenly identified as noise, thereby improving the accuracy of speech recognition. Therefore, setting an appropriate input format is important.
[0179] However, under conventional methods, it is difficult for workers other than speech recognition specialists to determine the appropriate input format, and setting the appropriate input format for all input fields is extremely time-consuming.
[0180] In contrast, according to the first embodiment, as mentioned above, an appropriate input format can be determined even when a speech recognition expert is not available. Furthermore, the information regarding the values in the input fields is not limited to the example input values; it may also be formatted as described later. Even when implemented in this way, the same effects as described above can be obtained.
[0181] Furthermore, according to the first embodiment, an input procedure 3a for inputting values into multiple input fields included in the report data 2a is stored in the input procedure storage unit 3. The input procedure 3a includes an identifier that identifies the input field of the input target item included in the report data 2a, and an estimated input format. Therefore, in addition to the effects described above, the input procedure 3a can be managed separately from the report data 2a. Note that the input procedure 3a is not limited to the memory of the information processing device 1, but may also be stored in a database device that can communicate with the information processing device 1, or in a separate sheet of the report data 2a. In addition to the identifier and input format, the input procedure 3a may also include a procedure number, guidance, reference value, procedure group name, etc., as appropriate. Even with these modifications, the input procedure 3a can be managed separately from the report data 2a, as described above.
[0182] Furthermore, according to the first embodiment, the range selection unit 5 selects a range of multiple input fields from the form data 2a. The input value example acquisition unit 4 acquires input value examples corresponding to the input fields in the selected range. The format estimation unit 6 estimates the input format common to the input fields in the range based on the correspondingly acquired input value examples. Therefore, in addition to the effects described above, the configuration of selecting a range of multiple input fields reduces the effort required to acquire input value examples compared to selecting individual input fields.
[0183] Furthermore, according to the first embodiment, the dictionary generation unit 8 generates a speech recognition dictionary based on the estimated input format. The recognition control unit 11 performs speech recognition on the user's utterance using the speech recognition dictionary and inputs the obtained speech recognition result text into the input field according to the identifier in the input field. Therefore, in addition to the effects described above, it is possible to perform voice input in response to the user's utterance.
[0184] Furthermore, according to the first embodiment, the information regarding the value of the input field is an example of an input value, which is an example of the value of the input field. The input value example acquisition unit 4 acquires the value as an input value from at least one of the following: the form data 2a in which the values of the input fields have been pre-entered, and the user interface that accepts the input of values into the input fields. Therefore, in addition to the effects described above, input value examples can be acquired from various sources.
[0185] Furthermore, according to the first embodiment, the format estimation unit 6 may estimate the input format for each input field of the range and estimate one of the estimated input formats as the common input format. In this case, in addition to the effects described above, for example, if one of the multiple input formats encompasses another, the input formats can be integrated by adopting the narrowest input format.
[0186] Furthermore, according to the first embodiment, the format estimation unit 6 may estimate the input format for each input field of the range, and estimate an input format that encompasses multiple of the estimated input formats as a common input format. In this case, in addition to the effects described above, the input formats can be integrated, for example, by determining an input format that encompasses all of the multiple input formats.
[0187] Furthermore, according to the first embodiment, the format estimation unit 6 adjusts the estimated input format based on the number of corresponding input value examples obtained. Therefore, in addition to the effects described above, when the selected range (number of cells) is small, the user's convenience can be maintained by slightly relaxing the input format, such as increasing the number of digits in the input format (allowing +1 digit).
[0188] Furthermore, according to the first embodiment, the input procedure 3a may further include upper and lower limits (reference values) indicating the range of values in the input field. The format estimation unit 6 may adjust the estimated input format based on the upper and lower limits included in the input procedure 3a. In this case, adjustments such as slightly loosening the input format to satisfy the conditions of the upper and lower limits can be made, thereby estimating a more appropriate input format.
[0189] Furthermore, according to the first embodiment, the input procedure 3a further includes guidance regarding voice input to the input field. The format check unit 7 checks the estimated input format based on the guidance. Therefore, in addition to the effects described above, the configuration that checks the estimated input format with guidance can improve the accuracy of the estimated input format. For example, it can check for invalid input formats, such as when the guidance wording is "○○ time" but the input format is not in time format. Also, for example, since the input format of input fields for multiple items is estimated from example input values and the validity of the input format is checked, even non-speech recognition experts can set an appropriate input format, and the time required to create the input procedure can be reduced compared to manually considering the input format.
[0190] Furthermore, the format checking unit 7 may also check the input format by using rules for format checking. Alternatively, the format checking unit 7 may be implemented as a trained model obtained by training a machine learning model using existing data. Here, as existing data, for example, a dataset can be used that includes input data containing example input values and the input format estimated from the example input values, and output data containing the results of checking the estimated input format. Even with this transformation, the estimated input format can be checked, thus improving the accuracy of the estimated input format.
[0191] Furthermore, according to the first embodiment, the format checking unit 7 generates example input values from the estimated input format, synthesizes the generated example input values into speech, and performs speech recognition on the speech data obtained from the speech synthesis using a speech recognition dictionary. The format checking unit 7 checks the estimated input format by determining whether the text of the obtained speech recognition result matches the generated example input values. Therefore, in addition to the effects described above, the accuracy of the estimated input format can be improved by checking the estimated input format using speech recognition.
[0192] Furthermore, according to the first embodiment, the format check unit 7 may classify the situation in cases where there is no match by determining whether or not a command was detected from the text of the speech recognition result. In this case, the validity of the input format can be checked by confirming that there are no false detections of commands.
[0193] Furthermore, according to the first embodiment, the format check unit 7 checks the input format by calculating the recognition rate or misrecognition rate of speech recognition based on the number of times it has made a determination and the result of the determination. For example, the misrecognition rate may be calculated by dividing the result determined to be a misrecognition by the number of times it has made a determination, which is the number of input value examples. Alternatively, the recognition rate may be calculated by dividing the result determined to be a successful recognition by the number of times it has made a determination, which is the number of input value examples. In this case, in addition to the effects described above, the validity of the input format can be statistically checked.
[0194] Furthermore, according to the first embodiment, the format check unit 7 notifies the user of the results of the input format check in an identifiable manner. Therefore, in addition to the effects described above, the configuration that notifies whether the failure of speech recognition was due to a command misdetection allows the user to take appropriate action depending on the situation in which speech recognition failed. However, the format check unit 7 may also notify the user of concerns about a decrease in recognition rate, and may notify the user of the factors causing the decrease in recognition rate (command conflicts, procedures with low recognition rates, frequently mistaken characters, etc.). For example, it may notify the user of homophones such as "9" and "Q". In this way as well, the user can be prompted to take appropriate action depending on the situation in which speech recognition failed.
[0195] Furthermore, according to the first embodiment, the format check unit 7 may change the voice input command that matches the detected command. In this case, the conflict with the command can be resolved by changing the wording of the command, thus avoiding unforeseen situations caused by false command detection.
[0196] Furthermore, according to the first embodiment, the display unit 12 may display the report data 2a and display the estimated input format in the input fields included in the report data 2a. Similarly, the display unit 12 may generate one or more example input values from the estimated input format and display the generated example input values in the input fields included in the report data 2a. Also similarly, the display unit 12 may generate a range of values that can be entered into the input fields from the estimated input format and display the generated range in the input fields included in the report data 2a. Therefore, in addition to the effects described above, since the estimated input format and the example input values generated from that input format are displayed on the report data 2a, the input position, input format, and example input values on the report data 2a can be checked at a glance. Accordingly, the estimated input format can be displayed with high readability along with the report, and the input format can be checked automatically.
[0197] To elaborate, traditionally, even after determining the input format, the display of the input format and the input item were separated, resulting in poor overview. Furthermore, it was difficult to determine whether the set input format was truly appropriate just by looking at the input format. Therefore, in the conventional method, checking whether the input format was appropriate required actually trying out the input, which was even more time-consuming.
[0198] In contrast, according to the first embodiment, as shown in Figure 12 for example, the display of the input format and the input target items are close together, resulting in high readability, and it is easy to determine whether the set input format is appropriate by looking at the input format.
[0199] Furthermore, according to the first embodiment, the input procedure 3a includes an identifier that identifies the input field of the input target item included in the report data 2a, an estimated input format, and the order (procedure number) in which values are entered into the multiple input fields included in the report data 2a. The display unit 12 displays the report data 2a and, based on the input procedure 3a, superimposes the order in which values are entered into the multiple input fields included in the report data 2a. Therefore, in addition to the effects described above, the user can visually confirm whether the input order has been set correctly. However, the display unit 12 may also superimpose information (guidance, reference values) within the input procedure 3a onto the report data 2a. This allows the user to visually confirm whether the information (guidance, reference values) within the input procedure 3a has been set correctly, as described above.
[0200] <Second Embodiment> In the first embodiment, the range selection unit 5 selected a range on the form data 2a for which to estimate the input format, in response to the input procedure creator's operation. In contrast, in the second embodiment, the range selection unit 5 automatically selects a range, and then automatically performs all operations of estimating the input format of the selected range and setting it in the input procedure 3a. That is, unlike the first embodiment, the range selection unit 5 automatically estimates multiple ranges to which the same input procedure should be set.
[0201] Accordingly, as shown in Figure 21, the range selection unit 5 performs at least one of the following in addition to the functions described above: range selection using guidance in input procedure 3a and range selection using the position of the input target item in input procedure 3a.
[0202] The range selection unit 5 selects a range of input fields based on the guidance, for example, by selecting a range of input fields corresponding to guidance with the same wording. Specifically, the range selection unit 5 may consider the range of input fields corresponding to the same guidance included in input procedure 3a as one range.
[0203] In the case of range selection using item position, the range selection unit 5 selects a range of input fields based on the situation in which the identifier of the input field of the input target item indicates a nearby location. Specifically, if the identifier of the input target item included in input procedure 3a indicates an adjacent input field, the range of input fields identified by that identifier may be considered as one range. However, it is preferable to obtain an example input value from the input field of the input target item, estimate the input format for one of the example input values using the format estimation unit 6, and include only if there are N or more adjacent input fields with the same input format in the range of input fields (where N is a predetermined threshold).
[0204] Figure 22 illustrates the list of form data 2a, input value examples, and input procedures, as well as the ranges selected based on them. For example, B1-D1, B2-D2, ..., B5-D5 have common guidance ("Inspection Time", "Starting Current", ..., "Appearance", respectively), so they are set to one range rg1, rg2, ..., rg5, respectively. For B8-D9, B8-B9, C8-C9, and D8-D9 have the same guidance ("Noise, Position 1", ..., "Noise, Position 3"), so they are set to one range rg81, ..., rg83, as shown by the dashed line. Furthermore, since the input format estimated from the already entered input value examples is decimal_2_1 for items adjacent to B8-D8, when all are combined, B8-D9 becomes one range rg89, as shown by the dashed line.
[0205] The other configurations are the same as in the first embodiment.
[0206] Next, the operation of the information processing device configured as described above will be explained using the flowchart in Figure 23.
[0207] Now, as described above, step ST1 is executed to create an input procedure other than the input format.
[0208] After step ST1, in step ST3a, the range selection unit 5 selects a range of input fields, for example, based on guidance. The range selection unit 5 also selects a range of input fields, for example, based on the identifier of the input target item included in input procedure 3a. Specifically, for example, if the identifier of the input target item indicates an adjacent input field, the range selection unit 5 selects the range of that adjacent input field. In this case, the input target item corresponding to an input field that does not belong to any range is treated as a range consisting of a single item.
[0209] After step ST3a, a loop process is executed in the selected range between steps ST3b and ST3d. This loop process includes steps ST4 to ST7 as described above, as well as step ST3c.
[0210] Step ST4 is the process of extracting the steps within the selected range from input step 3a.
[0211] Step ST5 is the process of obtaining example input values for the extracted procedure from the report data 2a.
[0212] Step ST6 is a process that estimates the input format from the acquired input value examples.
[0213] Step ST7 is a process that checks the estimated input format.
[0214] After step ST7, in step ST3c, the format estimation unit 6 does not notify immediately even if the result of the format check requires notification, but instead adds up the results of the format check processing. That is, the format estimation unit 6 adds up the number of command false detections and false recognitions in the format check processing, and merges the false recognition list.
[0215] After loop processing for each range, the format check unit 7 executes steps ST8 and ST9 as described above. As a result, the format check unit 7 refers to the summation result of the format check process and notifies the input procedure creator from the display unit 12 whether or not there is a command misdetection and whether the misrecognition rate is above a threshold (step ST9).
[0216] As described above, the processes from step ST10 onward are executed. This completes the process of creating the input procedure.
[0217] As described above, according to the second embodiment, the input procedure 3a further includes guidance including the item name of the input target item. The range selection unit 5 selects the range of the input field based on the guidance or identifier. Therefore, in addition to the effects described above, the time required to set up the input procedure 3a can be further reduced because the range to which the same input format should be set is automatically selected.
[0218] <Third Embodiment> In the first embodiment, the input format was estimated based on an example of an input value. In contrast, in the third embodiment, the input format is estimated using the display formatting already set in the report data 2a.
[0219] Accordingly, as shown in Figure 24, the information processing device 1 is equipped with a format setting acquisition unit 13 instead of the input value example acquisition unit 4.
[0220] The formatting acquisition unit 13 acquires the formatting of the selected range from the report data 2a. Formatting is the setting used to display the value in the input field. For example, the formatting acquisition unit 13 can use a function provided in an existing report editing application. For example, in Excel®, when displaying a value as a cell (input field), you can choose whether to display it as a number or as a date and time. The formatting acquisition unit 13 is an example of an information acquisition unit that acquires the formatting contained in the record data sheet as information about the value in the input field.
[0221] On the other hand, the format estimation unit 6 estimates the input format based on the acquired formatting, instead of using the input value example. For example, the format estimation unit 6 can mechanically convert one formatting to one input format. That is, if the format is "number (number of decimal places N)", the input format is "decimal_1~_N", if the format is "date (2012 / 3 / 14)", the input format is "date_yearmonthday", and so on. Since there is only one formatting for each item, one input format is obtained for each item. Subsequently, the integration of multiple input formats and adjustments according to the selection range are performed in the same manner as in the first embodiment.
[0222] The other configurations are the same as in the first embodiment.
[0223] With the above configuration, as shown in Figures 25, 26, and 27, an example of input values (E1, E2, ..., E m Instead of ), formatting (F1,F2,...,F m By using this method, the input procedure creation process, format estimation process, and format checking process are executed in the same manner as described above.
[0224] As described above, according to the third embodiment, in the report data 2a, the information regarding the value of the voice input field is the formatting used to display the value of the input field. The formatting acquisition unit 13 acquires the formatting included in the report data 2a as information regarding the value of the voice input field. Therefore, in addition to the effects described above, there is no need to separately prepare an example of an input value.
[0225] In the third embodiment, instead of using an example input value, the formatting used to display the value in the input field is obtained from the report data 2a, but the embodiment is not limited to this. For example, in the third embodiment, the formatting acquisition unit 13 may acquire information including the example input value and the formatting from the report data 2a. In this case, the formatting acquisition unit 13 is an example of an information acquisition unit that acquires information including the example input value and the formatting (information regarding the value in the input field) from a record data sheet. With such modifications, the effects of the first and third embodiments can be obtained simultaneously.
[0226] <Embodiment 4> In addition to the first embodiment, the fourth embodiment presents an actual operation example to the input procedure creator in order to confirm whether the set input procedure operates correctly.
[0227] Along with this, as shown in FIG. 28, the information processing apparatus 1 further includes an operation example display unit 14 in addition to the configuration shown in FIG. 1.
[0228] Here, instead of the voice inputter speaking to perform voice input, the operation example display unit 14 displays an operation example by automatically generating an input value and automatically generating voice data using voice synthesis. To supplement, the operation example display unit 14 generates an example of an input value from the estimated input format, synthesizes the generated example of an input value, and provides the voice data obtained by voice synthesis to the recognition control unit 11 in place of the user's speech, thereby displaying an operation example of voice input by the recognition control unit 11. The operation example may be saved as a video and displayed by video playback, or may be displayed superimposed on the form data 2a. After the operation example display unit 14 superimposes and displays the operation example on the form data 2a, the values input to the form data 2a as the operation example are collectively deleted according to the operation of the input procedure creator.
[0229] Other configurations are the same as those in the first embodiment.
[0230] Next, the operation of the information processing apparatus configured as described above will be described with reference to FIG. 29. This operation is different in that in the voice input process shown in FIG. 20, instead of the user's speech and voice recognition (steps ST26, ST27), one example of an input value is automatically generated from the input format, synthesized into voice, and the synthesized voice data is played back and voice recognized. The following will be described in order.
[0231] Steps ST21 to ST25 are executed as described above.
[0232] After step ST25, in step ST26a, the operation example display unit 14 generates an example of an input value from the estimated input format.
[0233] In step ST26b, the operation example display unit 14 obtains voice data by voice-synthesizing the generated input value example.
[0234] In step STS26c, the operation example display unit 14 provides the obtained voice data to the recognition control unit 11. The recognition control unit 11 performs voice recognition using a voice recognition dictionary, and inputs the text of the obtained voice recognition result into the input field according to the identifier that identifies the input field of the input target item. By providing the voice data to the recognition control unit 11, the operation example display unit 14 displays the operation example of voice input by such a recognition control unit 11.
[0235] Hereinafter, in the same manner as described above, the processing after step ST28 is executed.
[0236] As described above according to the fourth embodiment, the operation example display unit 14 generates an input value example from the estimated input format, voice-synthesizes the generated input value example, and provides the voice data obtained by the voice synthesis to the recognition control unit 11 in place of the user's utterance, thereby displaying the operation example of voice input by the recognition control unit 11. Therefore, in addition to the effects of the first embodiment, since how the set input procedure operates is displayed and reproduced as an actual operation example, it is possible to easily confirm whether the input procedure is appropriate.
[0237] <The Fifth Embodiment> The fifth embodiment is a specific example of each of the above embodiments and each modification example, and is a form in which the above-described information processing apparatus 1 is realized by a computer. The information processing apparatus 1 may be realized by either a general-purpose computer such as a personal computer or a dedicated computer such as an inspection apparatus (embedded system).
[0238] Figure 30 is a block diagram illustrating the hardware configuration of an information processing device 1 according to the fifth embodiment. This information processing device 1 includes a CPU (Central Processing Unit) 31, RAM (Random Access Memory) 32, ROM (Read Only Memory) 33, storage 34, display device 35, input device 36, and communication device 37, all of which are connected by buses.
[0239] The CPU 31 is a processor that performs arithmetic and control processing according to a program. The CPU 31 uses a predetermined area of the RAM 362 as a working area and, in cooperation with programs stored in the ROM 33 and storage 34, performs processing in each part of the information processing device 1 described above. The CPU 31 and each of the processors may also be called processing circuits.
[0240] RAM32 is a type of memory such as SDRAM (Synchronous Dynamic Random Access Memory). RAM32 functions as a workspace for the CPU31.
[0241] ROM33 is a memory that stores programs and various information in a way that prevents rewriting.
[0242] Storage 34 is a device that writes and reads data from magnetic recording media such as HDDs (Hard Disk Drives), semiconductor storage media such as flash memory, or magnetically recordable storage media such as HDDs, or optically recordable storage media. Storage 34 writes and reads data from the storage media in response to control from the CPU 31. Storage 34 is an example of computer memory.
[0243] The display device 35 is a display such as an LCD (Liquid Crystal Display). The display device 35 displays various information based on display signals from the CPU 361.
[0244] The input device 36 is an input device such as a mouse and a keyboard. The input device 36 receives information input by the user as an instruction signal and outputs the instruction signal to the CPU 31.
[0245] The communication device 37 communicates with external devices via a network in response to control from the CPU 31. The communication device 37 may also be called a communication circuit.
[0246] The instructions shown in the processing procedure described in the above-described embodiment can be executed based on a software program. A general-purpose computer system can store this program in advance and, by reading this program, can obtain effects similar to those of the control operation of the information processing device 1 described above. The instructions described in the above-described embodiment are recorded as a program that can be executed by a computer on a magnetic disk (flexible disk, hard disk, etc.), optical disk (CD-ROM, CD-R, CD-RW, DVD-ROM, DVD±R, DVD±RW, Blu-ray® Disc, etc.), semiconductor memory, or similar non-transitory recording medium. Any storage format is acceptable as long as it is a recording medium that can be read by a computer or embedded system. The storage medium may also be called a non-transitory computer-readable storage medium. A computer can read the program from this recording medium and, based on this program, have the CPU execute the instructions described in the program, thereby realizing an information processing method with operations similar to the control operation of the information processing device 1 in the above-described embodiment. For example, the information processing method includes, for each item, obtaining one or more items and information about the values of the input fields for one or more items from the form data 2a, which includes voice input fields, and estimating the input format of the input fields based on that information. Of course, when a computer obtains or reads a program, it may do so via a network.
[0247] In addition, an OS (Operating System) running on a computer, middleware such as database management software, a network, etc., based on instructions of a program installed from a recording medium into a computer or an embedded system, may execute a part of each process for realizing this embodiment.
[0248] Furthermore, the recording medium in this embodiment is not limited to a medium independent of a computer or an embedded system, and also includes a recording medium that has downloaded and stored or temporarily stored a program transmitted via a LAN, the Internet, or the like.
[0249] Also, the recording medium is not limited to one, and even when the processes in this embodiment are executed from a plurality of media, it is included in the recording medium in this embodiment, and the configuration of the medium may be any configuration.
[0250] Note that the computer or embedded system in this embodiment is for executing each process in this embodiment based on a program stored in a recording medium, and may have any configuration such as a device consisting of one of a personal computer, a microcomputer, etc., or a system in which a plurality of devices are network-connected.
[0251] Also, the computer in this embodiment is not limited to a personal computer, and includes an arithmetic processing unit, a microcomputer, etc. included in an information processing device, and generically refers to devices and apparatuses capable of realizing the functions in this embodiment by a program.
[0252] (Modifications of each embodiment) Note that each embodiment and each modification may be expressed as an information processing method or an information processing program including each step of the information processing apparatus 1 described above.
[0253] According to at least one of the embodiments described above, even when an expert in speech recognition is absent, an appropriate input format can be determined. The same applies to at least one of the modifications described above.
[0254] While several embodiments of the present invention have been described, these embodiments are presented as examples only and are not intended to limit the scope of the invention. These embodiments can be carried out in a variety of other forms, and various omissions, substitutions, and modifications can be made without departing from the spirit of the invention. These embodiments and their variations are included in the scope and spirit of the invention, as well as in the claims and their equivalents. [Explanation of symbols]
[0255] 1... Information processing device, 2... Data storage unit, 2a... Form data, 3... Input procedure storage unit, 3a... Input procedure, 4... Input value example acquisition unit, 5... Range selection unit, 6... Format estimation unit, 7... Format check unit, 8... Dictionary generation unit, 9... Speech recognition unit, 10... Speech synthesis unit, 11... Recognition control unit, 12... Display unit, 13... Format setting acquisition unit, 14... Operation example display unit.
Claims
1. An information acquisition unit acquires, for each item, one or more items and information regarding the values of the input fields for those one or more items from a recording data sheet that includes an input field for voice input. A format estimation unit estimates the input format of the input field, which is the format of the value to be entered into the input field, based on the aforementioned information. An information processing program that enables computers to realize this.
2. The input procedure for entering values into multiple input fields included in the aforementioned recording data sheet is stored in the computer's memory. The input procedure includes an identifier that identifies the input field for the input item included in the recording data sheet, and the estimated input format. The information processing program according to claim 1.
3. The computer further implements a range selection unit that selects a range of multiple input fields from the aforementioned recording data sheet. The information acquisition unit acquires the information corresponding to the input fields within the selected range. The format estimation unit estimates the input format common to the input fields in the range based on the correspondingly acquired information. The information processing program according to claim 2.
4. The input procedure further includes guidance that includes the item name of the input target item, wherein the item name is used as content to be played back by speech synthesis at the start of each input procedure. The range selection unit selects the range of the input field based on the guidance or the identifier. The information processing program according to claim 3.
5. A dictionary generation unit generates a speech recognition dictionary based on the estimated input format, A recognition control unit that performs speech recognition using the speech recognition dictionary in response to the user's utterance and inputs the resulting speech recognition text into the input field according to the identifier, The information processing program according to claim 2, further enabling the computer to perform the above.
6. The aforementioned information is an example of an input value, which is an example of a value in the input field. The information acquisition unit acquires the values as input values from at least one of the following: the recording data sheet in which the values in the input fields are pre-entered, and the user interface that accepts the input of values into the input fields. The information processing program according to claim 1.
7. The aforementioned information is the formatting used to display the value in the input field, The information acquisition unit acquires the formatting settings included in the recording data sheet as the information. The information processing program according to claim 1.
8. The aforementioned information includes an example input value, which is an example of the value in the input field, and the formatting used to display the value in the input field. The information acquisition unit acquires the information, including the input value example and the formatting, from the recording data sheet. The information processing program according to claim 1.
9. The format estimation unit estimates the input format for each input field within the range, and estimates one of the estimated input formats as the common input format. The information processing program according to claim 3.
10. The format estimation unit estimates the input format for each input field within the range, and estimates an input format that encompasses multiple of the estimated input formats as the common input format. The information processing program according to claim 3.
11. The format estimation unit adjusts the estimated input format based on the number of corresponding pieces of information obtained. The information processing program according to claim 3.
12. The input procedure further includes upper and lower limits indicating the range of values in the input field. The format estimation unit adjusts the estimated input format based on the upper and lower limits included in the input procedure. The information processing program according to claim 3.
13. The computer further implements a format checking unit. The input procedure further includes guidance regarding voice input to the input field, which includes text to be played back by speech synthesis at the start of the voice input. The format checking unit checks the estimated input format based on the guidance. The information processing program according to claim 2.
14. A format checking unit checks the estimated input format by generating example input values from the estimated input format, performing speech synthesis on the generated example input values, performing speech recognition on the speech data obtained from the speech synthesis using the speech recognition dictionary, and determining whether the text of the obtained speech recognition result matches the generated example input values. The information processing program according to claim 5, further enabling the computer to perform the above.
15. The format check unit, based on the results of the determination, classifies the situation in the case of a mismatch by determining whether or not a command was detected from the text of the speech recognition result. The information processing program according to claim 14.
16. The format checking unit checks the input format by calculating the recognition rate or misrecognition rate of the speech recognition based on the number of times the judgment was made and the result of the judgment. The information processing program according to claim 14.
17. The format checking unit notifies the results of the input format check in an identifiable manner. The information processing program according to claim 14.
18. The format check unit, when it detects the command, modifies the voice input command to match the command in response to the user's operation. The information processing program according to claim 15.
19. A display unit that displays the aforementioned recording data sheet and displays the estimated input format in the input fields included in the recording data sheet. The information processing program according to claim 1, further enabling the computer to perform the above.
20. A display unit that generates one or more example input values from the estimated input format and displays the generated example input values in the input fields included in the recording data sheet. The information processing program according to claim 1, further enabling the computer to perform the above.
21. A display unit that generates a range of values that can be entered into the input field from the estimated input format and displays the generated range in the input field included in the recording data sheet. The information processing program according to claim 1, further enabling the computer to perform the above.
22. The display unit is further implemented in the computer. An input procedure including an identifier that identifies the input field of the input target item included in the recording data sheet, the estimated input format, and the order in which values are entered into the multiple input fields included in the recording data sheet is stored in the computer's memory. The display unit displays the recording data sheet and, based on the input procedure, superimposes the order in which the values are entered into the multiple input fields included in the recording data sheet. The information processing program according to claim 1.
23. An operation example display unit generates an example of voice input operation by the recognition control unit by generating an example of input value from the estimated input format, synthesizing the generated example of input value into speech, and providing the speech data obtained from the speech synthesis to the recognition control unit in place of the user's utterance. The information processing program according to claim 5, further enabling the computer to perform the above.
24. An information acquisition unit acquires, for each item, one or more items and information regarding the values of the input fields for those one or more items from a recording data sheet that includes an input field for voice input. A format estimation unit estimates the input format of the input field, which is the format of the value to be entered into the input field, based on the aforementioned information. An information processing device equipped with the following features.
25. For each item, obtain one or more items and information regarding the values of the input fields for those one or more items from the recording data sheet, which includes the voice input fields. Based on the above information, estimate the input format which is the input format of the input field and the format of the value to be entered into the input field, An information processing method comprising the following:
Citation Information
Patent Citations
Table calculation software cooperation processor and its method
JP1996329158A
Computer-executable program and method, and processor
JP2008052676A
Authoring device, authoring method, and program
JP2017102939A
Summary generation device, summary generation method, and summary generation program
JP2017167433A
Smart Selection Engine
US20140372854A1