Information processing apparatus, method, and computer program product

By generating templates related to the input order and using voice recognition technology, the problem of inconvenient voice data input in existing technologies has been solved, achieving efficient and convenient data input, especially in manufacturing and maintenance inspection sites.

CN115620724BActive Publication Date: 2026-05-15KK TOSHIBA
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
KK TOSHIBA
Filing Date
2022-02-28
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

In manufacturing and maintenance sites, existing technologies require pre-setting multiple range combinations when using voice input data, resulting in inconvenient and inefficient data input.

Method used

The generation unit generates templates related to the input order, the voice recognition unit identifies the user's speech, the range of input objects is determined, and the decision unit and control unit provide emphasis display and data input, thereby improving the efficiency and convenience of data input.

Benefits of technology

It enables efficient and convenient voice data input across multiple projects, shortening data input time and reducing the complexity of user operations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115620724B_ABST
    Figure CN115620724B_ABST
Patent Text Reader

Abstract

An information processing device, method, and program are related. Efficiency and convenience of data input based on sound can be improved. An information processing device of one embodiment includes a generation unit, a sound recognition unit, and a determination unit. The generation unit generates a template related to one or more items that are likely to be designated, with reference to an input order related to an input target item selected from a plurality of items, with respect to a data table for recording including the plurality of items. The sound recognition unit performs sound recognition on a user's utterance to generate a sound recognition result. The determination unit determines an input target range related to one or more items designated by the user's utterance from among the plurality of items, based on the template and the sound recognition result.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is based on Japanese Patent Application 2021-117888 (filed on July 16, 2021), and enjoys priority from that application. This application incorporates the entire contents of that application by reference. Technical Field

[0002] The embodiments of the present invention relate to information processing apparatus, methods, and programs. Background Technology

[0003] Sometimes, in the manufacturing or maintenance / inspection areas, the measured values ​​from measuring equipment and the results of visual inspections are entered into documents and forms for sharing among operators or with customers. The content of each data entry in the documents and forms is predetermined, and operators perform the work in the prescribed order, entering the results into the designated data entry locations.

[0004] In standard document processing software, data is entered via text. However, text input is time-consuming during the workflow, leading to a demand for voice input. One approach is to define input fields and their corresponding content in a separate application, allowing values ​​to be entered into the selected field during speech input. Furthermore, the next field to be entered can be specified during setup, enabling continuous input. However, when inputting values ​​across multiple ranges, it's impractical to pre-define all possible ranges for speech input. Summary of the Invention

[0005] The problem to be solved by the present invention is to provide an information processing apparatus, method and program that can improve the efficiency and convenience of voice-based data input.

[0006] One embodiment of the information processing apparatus includes a generation unit, a voice recognition unit, and a decision unit. The generation unit generates a template related to one or more items that may be specified, based on an input order related to input object items selected from the plurality of items, using a data table containing records of multiple items. The voice recognition unit performs voice recognition on a user's speech and generates a voice recognition result. The decision unit determines, based on the template and the voice recognition result, a range of input objects related to one or more items specified by the user's speech from the plurality of items.

[0007] The information processing device based on the above structure can improve the efficiency and convenience of voice-based data input. Attached Figure Description

[0008] Figure 1This is a block diagram illustrating the information processing apparatus of the first embodiment.

[0009] Figure 2 This is a diagram illustrating an example of a data table for recording in this embodiment.

[0010] Figure 3 This is a diagram showing an example of an input sequence list stored in the sequence storage unit according to the first embodiment.

[0011] Figure 4 This is a diagram showing an example of a speech template stored in the template storage section according to the first embodiment.

[0012] Figure 5 This is a flowchart illustrating the input processing of the information processing apparatus according to the first embodiment.

[0013] Figure 6 This is a flowchart showing the details of the input object range determination process in step S505.

[0014] Figure 7 This is a diagram illustrating an example of the determination dictionary of the first embodiment.

[0015] Figure 8 This is a diagram illustrating another example of the determination dictionary of the first embodiment.

[0016] Figure 9 This is a diagram illustrating a specific example of the input processing of the information processing apparatus according to the first embodiment.

[0017] Figure 10 This is a diagram illustrating a specific example of the input processing of the information processing apparatus according to the first embodiment.

[0018] Figure 11 This is a diagram illustrating a specific example of the input processing of the information processing apparatus according to the first embodiment.

[0019] Figure 12 This is another example of a diagram showing an emphasis display of the range of input objects.

[0020] Figure 13 This is a diagram illustrating an example of how the input sequence list of a variation of the first embodiment is updated.

[0021] Figure 14 This is a diagram illustrating an example of how the input sequence list of a variation of the first embodiment is updated.

[0022] Figure 15 This is a diagram illustrating an example of how adding a speech template can be prompted.

[0023] Figure 16This is a diagram showing the input sequence list stored in the sequence storage unit according to the second embodiment.

[0024] Figure 17 This is a diagram illustrating an example of a value dictionary in the second embodiment.

[0025] Figure 18 This is a diagram illustrating an example of the first-range dictionary of the second embodiment.

[0026] Figure 19 This is a diagram illustrating an example of the second scope dictionary of the second embodiment.

[0027] Figure 20 This is a diagram illustrating an example of a voice recognition dictionary according to the second embodiment.

[0028] Figure 21 This is a flowchart illustrating the input processing of the information processing apparatus according to the second embodiment.

[0029] Figure 22 This is a diagram illustrating an example of the first keyword recognition dictionary of the third embodiment.

[0030] Figure 23 This is a diagram illustrating an example of the second keyword recognition dictionary of the third embodiment.

[0031] Figure 24 This is a diagram illustrating an example of the syntax recognition dictionary of the third embodiment.

[0032] Figure 25 This is a flowchart illustrating the input processing of the information processing apparatus according to the third embodiment.

[0033] Figure 26 This is a diagram illustrating a specific example of the voice recognition processing according to the third embodiment.

[0034] Figure 27 This is a diagram illustrating a specific example of the voice recognition processing according to the third embodiment.

[0035] Figure 28 This is a block diagram illustrating the voice recognition processing of the fourth embodiment.

[0036] Figure 29 This is a diagram illustrating an example of the keyword recognition dictionary of the fourth embodiment.

[0037] Figure 30 This is a diagram illustrating an example of the first grammar recognition dictionary of the fourth embodiment.

[0038] Figure 31 This is a diagram illustrating an example of the second grammar recognition dictionary of the fourth embodiment.

[0039] Figure 32 This is a diagram illustrating an example of a voice recognition dictionary according to the fourth embodiment.

[0040] Figure 33 This is a flowchart illustrating the input processing of the information processing apparatus according to the fourth embodiment.

[0041] Figure 34 This is a diagram illustrating a specific example of the voice recognition processing according to the fourth embodiment.

[0042] Figure 35 This is a diagram illustrating an example of the operation of the information processing apparatus according to the fifth embodiment.

[0043] Figure 36 This is a diagram illustrating an example of the hardware structure of an information processing device.

[0044] (Symbol Explanation)

[0045] 10: Information processing device; 20: Document data; 21: Data item; 22: Input position; 30, 130, 160: Input sequence list; 70, 80: Judgment dictionary; 71, 81: Range specification template; 72, 73: Standard expression; 91: Bold outline; 101: Order storage unit; 102: Template storage unit; 103: Voice recognition unit; 104: Voice synthesis unit; 105: Generation unit; 106: Decision unit; 107: Judgment unit; 108: Control unit; 109: Input unit; 131: Group; 141: Speech template; 150: Message; 170: Value dictionary; 180, 190: Range dictionary; 200 320: Syntax recognition dictionary; 220: First keyword recognition dictionary; 230: Second keyword recognition dictionary; 240: Syntax recognition dictionary; 261~263, 271, 272, 341, 343: Interval; 281: Cache section; 290: Keyword recognition dictionary; 310: Syntax recognition dictionary; 342: Preset period; 351: Copy object range; 352: Input object range; 361: CPU; 362: RAM; 363: ROM; 364: Storage device; 365: Display device; 366: Input device; 367: Communication device; 1101: String; 1201: Input object range. Detailed Implementation

[0046] Hereinafter, the information processing apparatus, method, and program of this embodiment will be described in detail with reference to the accompanying drawings. Furthermore, in the following embodiments, portions marked with the same reference numerals are assumed to perform the same operations, and repeated descriptions are omitted where appropriate.

[0047] (First Embodiment)

[0048] In the first embodiment, it is envisioned that data input values ​​for documents are entered through voice input based on the user's free speech.

[0049] Reference Figure 1 The block diagram illustrates the information processing apparatus of the first embodiment.

[0050] The information processing apparatus 10 of the first embodiment includes an order storage unit 101, a template storage unit 102, a voice recognition unit 103, a voice synthesis unit 104, a generation unit 105, a decision unit 106, a judgment unit 107, a control unit 108, and an input unit 109.

[0051] The sequence storage unit 101 stores an input sequence list related to the input order of multiple items contained in the record data table. The record data table is, for example, a data table used to input values ​​for a certain item, such as a document data table, inspection data table, or experimental data table.

[0052] The template storage unit 102 stores templates and dictionaries used to detect user speech.

[0053] The voice recognition unit 103 performs voice recognition on user speech acquired from a microphone (not shown) or the like, and generates a voice recognition result. As a voice recognition process, the voice recognition unit 103 included in the information processing device 10 includes a voice recognition processing engine, which can perform voice recognition processing within the device or send voice data related to speech to the cloud or the like, and the voice recognition unit 103 obtains the result after voice recognition processing in the cloud.

[0054] The sound synthesis unit 104 generates synthesized sound to inform the user of content such as guidance. The generated synthesized sound can be output from a speaker (not shown) or the like.

[0055] The generation unit 105 refers to the input sequence list and generates templates related to one or more items that may be specified, based on the input order related to the input object items. The input object items are the items selected as processing objects from multiple items contained in the record data table.

[0056] Based on the template and the voice recognition results, the decision unit 106 determines the range of input objects related to one or more items specified by the user's speech among multiple items.

[0057] In addition to determining various conditions, the determination unit 107 also determines the range specification statement used to specify the range of input objects and the value statement representing the value input into the range of input objects.

[0058] In addition to performing various controls, the control unit 108 also highlights the range of input objects on a recording data table displayed on a monitor (not shown).

[0059] In addition to inputting various data, the input unit 109 also inputs values ​​related to value statements within the input object range.

[0060] Next, refer to Figure 2 This illustrates an example of document data envisioned in this embodiment.

[0061] In this embodiment, the envisioned data table for recording is, for example, a two-dimensional grid format where values ​​are input into a table, similar to that used in spreadsheet software. Hereinafter, document data 20 will be used as an example to illustrate this data table for recording. Figure 2 In the example, the horizontal axis represents column numbers (A~D), and the vertical axis represents row numbers (1~7). In document data 20, there are data items 21 related to test items such as inspection date, "no dirt," and "no defects," and input positions 22 for users to input values ​​for each data item 21. Data items 21 can also be categorized into groups such as "appearance check" and "action check."

[0062] Input position 22 can be specified based on column number and row number. Figure 2 In the example, input position 22 corresponding to data item 21 "Inspection Date" can be expressed as "D2", with a value of "2021 / 02 / 15". Input position 22 corresponding to data item 21 "Is there any dirt?" can be expressed as "D3", with a value of "No abnormalities".

[0063] Furthermore, not limited to document data from table calculation software based on 2D arrangement, even free descriptions such as randomly configured input positions 22 can be processed by the information processing device 10 of this embodiment, as long as the input position 22 corresponding to the data item 21 can be uniquely determined by the user's statement.

[0064] Next, refer to Figure 3 This illustrates an example of an input sequence list stored in the sequence storage unit 101.

[0065] Figure 3 The input sequence list 30 shown is a list representing the order of input positions of user input values. The input sequence list 30 is a table that maps sequence numbers, input positions, guides, and entered flags. The input sequence list 30 can be used in conjunction with... Figure 2The document data 20 shown can be kept on the same data file or kept as different data. Keeping it on the same data file means keeping both document data 20 and input sequence list 30 in one data file. Furthermore, any standard table management software can manage multiple data tables in a consolidated manner, so it can be kept on the same data table as document data 20 or on a different data table.

[0066] The sequence number indicates the order in which multiple data items 21 are entered for document data 20. The input location is uniquely specified. Figure 2 The identifier for input position 22 is shown. For example, when input position 22 is specified using column number and row number, an identifier such as "D2" is used. The guide is the content played by the synthesized sound generated by the sound synthesis process performed by the sound synthesis unit 104, such as the name of the data item 21. The input flag is a flag indicating whether a value has been entered at input position 22. For example, if a value has been entered, it is assigned "1", and if no value has been entered, it is assigned "0 (zero)".

[0067] exist Figure 3 In the example, the sequence number "1", the input position "D2", the guide "check date" and the input flag "1" are respectively stored as one entry in the input sequence list 30.

[0068] Next, refer to Figure 4 This describes an example of a speech template saved in the template storage section 102.

[0069] exist Figure 4 The table shown stores multiple possible speaking patterns, or speaking templates, when the user specifies a range.

[0070] As a template for speeches, examples could include "summarize ${number}", "summarize up to ${endpoint}", or "summarize from ${start} to ${endpoint}". "${}" means treating any speech as an object. Figure 4 In the example, expressed as "${number}", it represents the replacement object part when there are statements related to the number. For example, consider statements related to the numbers "3" and "5". "${start}" and "${end}" represent the start and end points of a specified range based on the input sequence list, respectively. For example, including Figure 3 The input sequence number 30 shown can be used, along with the input position identifier and guiding content. Specifically, in representing... Figure 3When the sequence number of the input sequence list 30 is in the range of "2" to "4", it can correspond to statements such as "summarize from the 2nd to the 4th", "summarize from D3 to D5", or "summarize from dirt to print". Furthermore, the start and end points can be specified without following the same category in the input sequence list 30. That is, the start point can be specified according to the guiding category, and the end point can be specified according to the category of the input position, as in "summarize from dirt to D5".

[0071] Next, refer to Figure 5 The flowchart illustrates the input processing of the information processing apparatus 10 in the first embodiment. Furthermore, for example, the control unit 108 sets the input sequence list's already-input flag to a non-input state before performing input processing for document data. Figure 3 In the example, the input flag of input sequence list 30 is set to zero.

[0072] In step S501, the determination unit 106 determines the data item 21 to be processed, i.e., the input item. For example, the input position with the smallest sequence number among the entries whose input flag is zero can be set as the input item. Furthermore, when the processing of the next input item, i.e., the processing of step S501, is the second time or later, the determination can be made as follows: If the input object range (described later) is one input item, or if the input object range includes multiple input positions but not input items, the determination unit 106 can set the input position next to the input position in the current processing as the input item. On the other hand, if the input object range includes multiple input positions and also includes input items, the determination unit 106 can set the input position in the input object range with a larger sequence number than the last one and which has not been input as the input item.

[0073] In step S502, the control unit 108 emphasizes the input location of the input object item. For example, it may display a box surrounding the input location with a thick line.

[0074] In step S503, the sound synthesis unit 104 synthesizes sound for guidance that prompts input related to the input position and plays the synthesized sound. For example, it can synthesize sound for guidance phrases corresponding to the input position in the input sequence list 30 and play them from a speaker or the like.

[0075] In step S504, the determination unit 107 determines whether the sound recognition unit 103, which acquired the user's speech, has generated a sound recognition result. If a sound recognition result has been generated, the process proceeds to step S505; otherwise, the process of step S504 is repeated. Furthermore, in the sound recognition processing of the sound recognition unit 103, padding, repetitions, etc., in the speech are removed using existing removal techniques as needed.

[0076] In step S505, the generation unit 105, the determination unit 106, and the judgment unit 107 perform an input object range determination process based on the voice recognition result. As a result of the determination process, an input object range and a value statement relating to the value input to the input position are generated. For details regarding the input object range determination process, please refer to... Figure 6 To be described later.

[0077] In step S506, the control unit 108 emphasizes the range of input objects determined in step S505. The method of emphasizing the range is the same as in step S502.

[0078] In step S507, the voice synthesis unit 104 plays a confirmation message prompting the user to confirm the range of input objects. The confirmation message can be a fixed phrase such as "Is it okay?" or it can be a message that prompts the user to repeat the specified range of input objects. The voice synthesis unit 104 synthesizes the message containing the range of input objects and plays the synthesized sound.

[0079] In step S508, the determination unit 107 determines whether the input content is confirmed. For example, it can determine that the input content is confirmed when a statement indicating agreement or affirmation, such as "Yes" or "OK," is detected from the user, or when the user presses a predetermined button. If the input content is confirmed, the process proceeds to step S509. If the input content is not confirmed, that is, when a statement denying the input content is made or when the voice input is restarted, the process returns to step S504 and repeats the same process.

[0080] In step S509, the input unit 109 inputs value-based data (e.g., numerical values ​​or strings) into the input positions included in the input object range.

[0081] In step S510, the input unit 109 sets the input position to "input" according to the input position where the value is input. Specifically, the input unit 109 may, for example, set the input flag of the corresponding input sequence list to "1".

[0082] In step S511, the determination unit 107 determines whether there are any unentered input positions in the input sequence list. If there are unentered input positions, the process returns to step S501 and repeats the same process. On the other hand, if there are no unentered input positions, that is, if values ​​are entered at all input positions, the input processing of the document data 20 by the information processing device 10 ends.

[0083] Next, refer to Figure 6 The flowchart illustrates the details of the input object range determination process in step S505.

[0084] In step S601, the generation unit 105 generates a decision dictionary based on the input object items and the input sequence list. The decision dictionary is a dictionary that stores multiple templates representing possible input positions or combinations of multiple input positions spoken by the user when a statement is made as a specified unentered input position. For details regarding the decision dictionary, please refer to [link to relevant documentation]. Figure 7 To be described later.

[0085] In step S602, the determination unit 107 compares the sound recognition result with the determination dictionary to determine whether the sound recognition result includes a range-specified statement that is intended to encompass a specified range. Specifically, if the sound recognition result includes a portion consistent with the range-specified template, the consistent portion in the sound recognition result is determined to be a range-specified statement. If the sound recognition result includes a range-specified statement, the process proceeds to step S603; if the sound recognition result does not include a range-specified statement, the process proceeds to step S605.

[0086] In step S603, the determination unit 107 determines the part of the string that is later than the specified range of speech in the string that is the sound recognition result as the value (e.g., a string) that the user wants to input at the input position, and determines the part that is later than the specified range of speech as the value speech.

[0087] In step S604, the decision unit 106 determines one or more input locations specified by the range specification statement as the input object range. Then, it proceeds to... Figure 5 The processing of step S506. Furthermore, the processing order of steps S603 and S604 is not restricted; either can be executed first.

[0088] In step S605, the determination unit 107 determines the overall sound recognition result as a speech.

[0089] In step S606, the decision unit 106 determines the current input object item as the input object range. Then, it proceeds to... Figure 5The processing of step S506 is as follows. Furthermore, the processing order of steps S605 and S606 is not restricted; either can be executed first.

[0090] exist Figure 5 as well as Figure 6 Although not shown in the flowchart, the voice recognition unit 103 can control the switching between recording the user's speech and starting and stopping the voice recognition processing. For example, after the output of the guidance sound in step S503 and the output of the confirmation message in step S507, the voice recognition unit 103 can start recording the user's speech and performing voice recognition processing at the designated time. Furthermore, after generating a voice recognition result, the voice recognition processing is stopped to prevent the synthesized sound output from the information processing device 10 from looping back to the microphone and being processed along with the user's speech.

[0091] Furthermore, in cases where signal processing is not applied, such as the reverberation of synthesized sound from the information processing device 10 to the microphone, the aforementioned switching control may not be performed, and the switching control may still be performed. Figure 5 The input processing shown performs voice recognition processing.

[0092] Next, refer to Figure 7 This is an example of the determination dictionary of the first embodiment.

[0093] Figure 7 The decision dictionary 70 shown includes multiple range specification templates 71. The range specification templates 71 include standard expressions. Standard expressions are generated by expanding the permutation object portion of the speech template for any input into a standard expression.

[0094] For example, the standard expression 72 with ID "1" can be generated by expanding "${number}" of the permutation object part contained in the speech template into "(?<value>\d+)". The standard expression 72 with ID "2" corresponds to the case where only the permutation object part "${endpoint}" contained in the speech template is included, and can be generated by expanding it into an identifier of the input order that is later than the current input order and a set of leading sums.

[0095] The standard expression 72 with ID "3" corresponds to the case where the substitution object part contained in the speech template has both "${start point}" and "${end point}". The start point can be expanded to the input position of the entry other than the last entry in the input sequence list and the leading sum set, and the end point can be expanded to the input position of the entry other than the first entry in the input sequence list and the leading sum set.

[0096] When a range is specified for a statement, the number of statements, start point, and end point can be determined from the portion consistent with the standard expression of the range-specified template. That is, when a number is specified, the start point is the sequence number in the input sequence list corresponding to the current input item, and the end point is the sequence number corresponding to the value obtained by subtracting 1 from the number of statements. Specifically, when in... Figure 3 The entry with sequence number "2" represents the input object item. In the case of stating "summarize 4 without exception," the range specified in the statement is "summarize 4." The sequence number corresponding to the endpoint is calculated as 2 + 4 - 1 = 5, continuing until the input position (D3~D6) corresponding to the entry with sequence number "5," which is the input object range. In other words, the input position of any entry in the input sequence list 30, the identifier of the input position, or any specified guidance, is within the input object range.

[0097] Similarly, when a start and end point are specified, for example, when a sound recognition result such as "summarizing from dirt to printed text" is obtained, if referring to... Figure 3 Given the input sequence list, the input position corresponding to the guide "dirt" at the starting point is "D3", and the input position corresponding to the guide "print" at the ending point is "D5". Therefore, the input object range can be set to "D3~D5".

[0098] Furthermore, the decision dictionary 70 is not limited to being generated in step S601, but can also be generated before the sound recognition processing in step S504.

[0099] In addition, not limited to Figure 7 For example, the standard expression can also be expanded based on other expressions that can specify the input positions contained in the input sequence list.

[0100] Next, refer to Figure 8 This illustrates another example of the determination dictionary in the first embodiment.

[0101] It can also be like Figure 8 As shown in the decision dictionary 80, multiple ranges are specified from one template to designate template 81. For example, Figure 7 The range specified by ID "2" in template 71 is expressed as "summarize up to (?<endpoint> (D4) | (D5) | (D6) | (D7) | ... | (defect) | (printing) | (lit status) | (working status) | ...)". Furthermore, "A | B" is a non-terminal symbol indicating identification of "A" or "B".

[0102] On the other hand, Figure 8In the range specification template 81, each state can also be split. For example, the standard expression 72 with ID "2-1" is "aggregate until (?<endpoint> (D4) | (defect))", the standard expression 72 with ID "2-2" is "aggregate until (?<endpoint> (D5) | (printing))", the standard expression 72 with ID "2-3" is "aggregate until (?<endpoint> (D6) | (lighting state))", and the standard expression 72 with ID "2-4" is "aggregate until (?<endpoint> (D7) | (operating state))".

[0103] In addition, regarding the Figure 7 whose ID is "3" in the range specification template 71 that can form a combination of start and end points, in the Figure 8 range specification template 81, for a group of sequence numbers i, j (i, j are natural numbers where i < j), it can be expanded in the input position and union of the entries included in the input sublist.

[0104] Next, referring to Figure 5 and Figure 6 the flowchart and Figures 9 to 12 illustrate a specific processing example of the input processing of the information processing apparatus 10 according to the first embodiment.

[0105] Figure 9 It is a diagram showing an example of displaying the document data 20 based on the control unit 108 when starting processing for the entry with sequence number "2" in the input sublist 30 shown in Figure 3 .

[0106] Through the processing of step S502, the entry with sequence number "2" has an input position of "D3", so the input position of "D3" in the document data 20 is highlighted. In the Figure 9 example, an example of highlighting is shown by surrounding the cell of D3 with a thick frame 91.

[0107] Here, through the processing of step S503, for example, the voice synthesis unit 104 generates a synthesized voice of "Is there 'dirt'" as a voice guidance using the guiding item "dirt" in the input sublist 30 and notifies the user. Suppose a user who hears this synthesized voice speaks "Aggregate from dirt to printing without abnormality" in this case.

[0108] Through the processing of step S504, the voice recognition unit 103 generates a voice recognition result of "Aggregate from dirt to printing without abnormality".

[0109] Next, in step S601, the generation unit 105 generates a result related to Figure 7The determination dictionary related to the entry with the sequence number "2". In step S602, the determination unit 107 determines whether the sound recognition result includes a range-specified statement. Here, a part of the sound recognition result, such as "summarize from dirt to printed text", is consistent with the standard expression 72 with ID "3" in the determination dictionary, which is a range-specified template 71 including a start and end point, so "summarize from dirt to printed text" is determined to be a range-specified statement. As a result, in step S603, the statement "no abnormality" after "summarize from dirt to printed text" is set as a value statement. In step S604, the determination unit 106 determines the range-specified statement "summarize from dirt to printed text" and Figure 3 The input sequence list is set to a range of input objects "D3~D5" starting from the input position "D3" corresponding to "dirt" applied to standard expression 72 and ending at the input position "D5" corresponding to "print".

[0110] As a result, through the emphasis display processing in step S506, such as Figure 10 As shown, the input range, encompassing three cells from "D3" to "D5," is highlighted with a thick box (91). This allows users to easily determine whether they can define the desired input range using their own input.

[0111] Consent is obtained from the user through steps S507 and S508. Figure 10 In the case of speaking, as shown in the example of setting the range of input objects, such as... Figure 11 As shown, through the processing in step S509, the string "1101" "No anomaly" is entered as the value in each cell of the input object range, namely D3, D4, and D5. Afterwards, although not illustrated, the process... Figure 3 The input sequence table shown has the input flag set to "1" for the entries corresponding to input positions "D3~D5".

[0112] Furthermore, in order to proceed with the next step, regarding the input object items in the case of returning to step S501, the input object range includes multiple input positions "D3~D5", and also includes the current input object item "D3". Therefore, the decision unit 106 selects the input position in the input object range that is greater than the last sequential number "4" (input position D5) and has not been entered. Figure 3 In the example, the position numbered "5" (input position D6) determines the next input item.

[0113] Additionally, although not illustrated, when in Figure 9If a user utters "There is garbled text" in the specified state, the voice recognition unit 103 generates a voice recognition result of "There is garbled text" through the processing in step S504. In this case, through step S602, it is determined that there is no specified template with a standard expression consistent with the voice recognition result "There is garbled text". Through step S605, the voice recognition result "There is garbled text" is set as a value statement. In step S606, the input position "D3" of the input object item is set as the input object range, and the string "There is garbled text" is entered at the input position "D3".

[0114] In addition, Figure 9 as well as Figure 10 In the example, the input position is surrounded by a thick line for emphasis, but it is not limited to this; it can also be emphasized by making the border blink or coloring.

[0115] Figure 12 This shows another example of an emphasis display for the range of input objects. Figure 12 The image shows an example where the input range 1201 is colored differently from other cells, thus emphasizing the input range 1201. This emphasis can be any display scheme as long as it differs from the display scheme of cells outside the input range 1201.

[0116] According to the first embodiment described above, based on the input sequence list and the input object items, a dictionary for summarizing input for multiple items is generated. The user's speech is processed through voice recognition, and one or more input positions as the input object range and the values ​​to be input at those positions are extracted. The values ​​to be input into the input object range are then input. Thus, by simply speaking while performing the task and specifying the desired input range regarding the task results and inspection results, the user can summarize and input values ​​into one or more input positions at once for record data tables such as document data.

[0117] That is, there is no need to switch between specified ranges of modes, and users do not have the hassle of additional settings, allowing for seamless individual and aggregated input. As a result, efficient voice data input is possible, improving the efficiency and convenience of voice-based data input. Consequently, it can also shorten the operation time for inputting data into record tables such as documents.

[0118] (A variation of the first embodiment)

[0119] In the information processing apparatus 10 of the first embodiment, it is envisioned that an input sequence list is generated by obtaining items that can be used as guides from document data before use, but it is also possible to update the template by adding terms for specifying the range of input objects to the input sequence list during use.

[0120] Reference Figure 13 as well as Figure 14 This example illustrates how to update the input sequence list.

[0121] exist Figure 13 In the input sequence list 130 shown, in order to... Figure 2 The terms “appearance check” and “action check” in the test project are used to specify the range of input objects, and the user adds the items of group 131 to the input sequence list.

[0122] exist Figure 14 In this context, a new statement template 141, such as "summarize ${group}", has been added. Therefore, for example, if a statement like "summarize appearance" is obtained as a result of voice recognition, the determination unit 107 can refer to... Figure 13 The input sequence list shown summarizes the "Dirty (D3)," "Flaw (D4)," and "Print (D5)" entries that correspond to "Appearance" as the range of input objects.

[0123] In addition to the user manually updating the input sequence list and speech templates, the information processing device 10 can also learn the user's preference for specifying the range of input objects and provide the user with suggestions for adding new speech templates or providing alternative speech templates. Alternatively, it can automatically add new speech templates.

[0124] Figure 15 This illustrates an example of an additional setting of a speech template related to the range of input objects, prompted by the information processing device 10.

[0125] Imagine a scenario where, during multiple data input processes for a document, the same input range, including multiple input locations, is specified more than a predetermined number of times. The determination unit 107 can also determine that the user is specifying this input range at a high frequency, prompting the user to add a new statement template so that the input range can be specified with a shorter statement. Specifically, if the user specifies the input range more than a predetermined number of times using multiple words including a start and end point, such as "summarize from D3 to D5," the determination unit 107 determines that an additional statement template should be added.

[0126] Control unit 108 displays, for example, a message 150 prompting the user to add a speech template related to the input object range, such as "Set in a way that allows for comprehensive verification?". If the user answers "Yes," a name is input via voice or text, such as one that the user can easily specify as the input object range. Figure 15 In the examples, terms like "appearance," "outer appearance," or "exterior" are acceptable. Therefore, for example... Figure 13 as well as Figure 14 As shown, the input sequence list and the statement template are updated. In subsequent executions, the user can state statements such as "summarize appearances," "summarize external features," or "summarize external features," thus specifying the range of input objects from D3 to D5.

[0127] Furthermore, when prompting the addition of speech templates, the decision is not limited to rules based on a predetermined number of attempts; learned models from machine learning can also be used. For example, existing supervised learning can be employed, using training data that takes input object items as input data and specifies the user's preference for a range of input objects as correct answer data. Alternatively, the addition settings for speech templates can be recommended using a learned model that generates the learning results.

[0128] Furthermore, according to a variation of the first embodiment described above, a speech template is added based on the user's preference for specifying the range of input objects. Therefore, as a speech template, in addition to allowing the user to easily specify a name for the range of input objects and efficiently input values ​​via voice, the efficiency and convenience of voice-based data input are improved.

[0129] (Second Implementation)

[0130] In the second embodiment, the processing differs from the first embodiment in that it performs sound recognition on specific types of speech. In the first embodiment, free speech is assumed, but in noisy environments such as work sites, accurate sound recognition processing may not be possible. Therefore, the information processing apparatus of the second embodiment performs input processing only on speech that follows a specific input format, thereby improving the accuracy of sound recognition processing and enabling sound-based input processing of document data even in noisy environments.

[0131] In the information processing apparatus of the second embodiment, the generation unit 105 generates a voice recognition dictionary during voice recognition processing, which is used to recognize only input forms of speech based on specific grammar. The template storage unit 102 stores the voice recognition dictionary. In the second embodiment, the voice recognition dictionary is also referred to as a grammar recognition dictionary. Apart from the generation unit 105 and the template storage unit 102, the same operations as in the first embodiment are performed, so the description is omitted here.

[0132] Reference Figure 16 The input sequence list stored in the sequence storage unit 101 in the second embodiment is explained.

[0133] Figure 16 The input sequence list 160 shown includes, in addition to, Figure 3 In addition to the input sequence list 30, it also includes items in input form 161.

[0134] The input format is the form of speech that is only subject to a specific grammatical structure in the speech recognition process, and is used to generate the speech recognition dictionary. For example, terms are specified such as "date", "word (no abnormality | needs to be replaced)", "word (normal action | abnormal action)", or the pattern to be recognized by the speech recognition unit 103 is specified such as "numerical value (3 digits of integer part)", "numerical value (2 digits of integer part, 1 digit of decimal part)", "5 characters of English numerals".

[0135] Specifically, in the inspection date data item of the document data, the input format is set to "Date" to identify only the date. In the dirt data item, the input format is set to "Word (No Abnormalities | Need to Replace)" to indicate whether the acceptance is "No Abnormalities | Need to Replace".

[0136] Next, refer to Figures 17 to 19 This illustrates an example of a voice recognition dictionary generated by the decision unit 107.

[0137] The voice recognition dictionary in the second embodiment includes three types of dictionaries: a value dictionary for recognizing the value of the input form of the input object item, a first range dictionary for recognizing range-specified speech starting from the input object item, and a second range dictionary for recognizing range-specified speech that can be input at any time.

[0138] first, Figure 17 This shows an example of a value dictionary.

[0139] Value dictionary 170 is a dictionary that maps sequence numbers to grammar templates. Using value dictionary 170, at the input position corresponding to the sequence number, only the sound recognition result consistent with the grammar template is input. The dictionary used for numerical input uses other defined syntax such as "$integer N digits" for simplicity. Sequence numbers "2~4" identify "no abnormality" or "needs replacement", and sequence numbers "5~6" identify "normal action" or "abnormal action".

[0140] Next, Figure 18 Here is an example of a dictionary with the first range.

[0141] The first range dictionary 180 is a dictionary that maps sequence numbers to standard expressions. Using the first range dictionary 180, only consecutive input sequences with the same input format can be input. For example, starting with an input sequence numbered "2", only input sequences with the same format up to sequence number "3" or sequence number "4" can be input. Specifically, ranges can be specified using "summarize 2", "summarize 3", and direct endpoints such as "summarize until defective" and "summarize until printed". Additionally, "("No abnormality" | "Needs replacement")" and... Figure 17 The value dictionary is the same. On the other hand, when the input order starts from the sequence number "3", the range of inputs with the same form is up to the sequence number "4", so it is not possible to input a statement like "summarize 3", and only statements like "summarize 2" are accepted.

[0142] In addition, Figure 18 In the example, the word of the "guide" item set in the input sequence list is used, but similarly to the first embodiment, other items such as sequence numbers can be used to specify the range of input objects.

[0143] Next, Figure 19 Here is an example of a dictionary in the second range.

[0144] The second-range dictionary 190 is a dictionary that maps sequence numbers to standard expressions. The second-range dictionary 190 is designed to process statements regardless of the sequence number of the input items, even if the input items have any sequence number. It is generated for each group of multiple consecutive input sequences with the same input format, in order to group input positions with the same input format together. Figure 19In the example, to represent a range of sequence numbers "2~4", one can use statements like "from dirt to printing" to specify the start and end points, and statements like "summarize the appearance" to specify a "group" of the input sequence list. Following standard expression 73, "("No abnormality" | "Needs replacement")" and Figure 18 The first range dictionary 180 shown is the same. Similarly, the second range dictionary 190 can also use other items such as ordinal numbers to specify the range of input objects.

[0145] In addition, Figures 17 to 19 The example shown is that the value dictionary 170, the first range dictionary 180 and the second range dictionary 190 are generated as other dictionaries, but it is not limited to this. It can also be generated as a voice recognition dictionary that maps the template of the value to the standard expression of the range of speech.

[0146] Next, Figure 20 Here is an example of a voice recognition dictionary.

[0147] By Figures 17 to 19 The value dictionary 170, the first range dictionary 180, and the second range dictionary 190 shown are combined to generate various sound recognition dictionaries 200 (also called syntax recognition dictionaries 200) for each sequentially numbered entry when the input object item is selected. Figure 20 In the example, the entry with the sequence number "2" is extracted from each of the value dictionary 170, the first range dictionary 180, and the second range dictionary 190 when it is the entry for the input object item.

[0148] Furthermore, the input format specified in the voice recognition dictionary 200, such as "summarizing up to (defect | printing)," can also be found from... Figure 7 The range shown specifies the template generation.

[0149] Next, refer to Figure 21 The flowchart illustrates the input processing of the information processing device 10 in the second embodiment.

[0150] Furthermore, in the second embodiment, it is envisioned that a value dictionary, a first range dictionary, and a second range dictionary are generated in advance based on an input sequence list before the input processing performed by the information processing device 10 is executed, but they can also be generated before the voice recognition processing in the input processing is executed.

[0151] In step S2101, when the entry corresponding to the order number in the input sequence list is set as the input object item in step S501, the generation unit 105 generates a sound recognition dictionary corresponding to the order number of the input object item based on the value dictionary, the first range dictionary, and the second range dictionary.

[0152] In step S504, the voice recognition unit 103 performs voice recognition processing based on a voice recognition dictionary. In this voice recognition processing, only speech in input formats that exist in the voice recognition dictionary is accepted; therefore, speech in other input formats is rejected, and no voice recognition result is generated. Thus, if voice recognition processing of the user's speech cannot be performed for a certain period, a synthesized voice message such as "Voice recognition failed. Please say it again" can be output to the user to prompt them to re-enter the speech.

[0153] Furthermore, in step S505, the determination unit 107 determines the input object range. However, if the voice recognition processing can obtain information about which string in the input form using the value dictionary, the first range dictionary, and the second range dictionary identifies the user's speech, it can also determine that the speech is a range-specified speech when the speech is identified using the first range dictionary and the second range dictionary. Furthermore, if the voice recognition processing cannot obtain information indicating which part of the string existing in the voice recognition dictionary was used, the input object range can be determined by referring to the input sequence list and the range specification template, similar to the first embodiment.

[0154] According to the second embodiment described above, a voice recognition dictionary (grammar recognition dictionary) is generated that specifies the form of speech including a specified input range. Voice recognition processing is performed using this voice recognition dictionary, so that only values ​​consistent with the input form are recognized by voice. Therefore, voice recognition accuracy can be improved, and similarly to the first embodiment, the efficiency and convenience of voice-based data input can be improved.

[0155] (Third implementation)

[0156] In the third embodiment, unlike the embodiments described above, in addition to performing voice recognition processing based on the input format of the second embodiment, keyword detection-type voice recognition processing that does not require voice interval detection is also performed. By using keyword detection-type voice recognition processing that does not require voice interval detection, the range of input objects can be determined even in the middle of a speech, and prompts can be provided to the user.

[0157] In the information processing apparatus 10 of the third embodiment, the generation unit 105 generates a keyword detection-type voice recognition dictionary (keyword recognition dictionary) and a voice recognition dictionary (grammar recognition dictionary) specifying the grammar as the input form. The voice recognition unit 103 uses the keyword recognition dictionary and the grammar recognition dictionary to perform two types of voice recognition processing. Other structures are the same as in the above embodiment, so descriptions are omitted.

[0158] Reference Figure 22 as well as Figure 23 Here is an example of a keyword recognition dictionary according to the third embodiment. The keyword recognition dictionary includes a first keyword recognition dictionary for detecting statements within a range starting from a corresponding sequence number, and a second keyword recognition dictionary for detecting statements that can be entered at any time.

[0159] Figure 22 The first keyword recognition dictionary 220 shown corresponds to the second embodiment. Figure 18 The dictionary 180 in the first range shown corresponds to the sequence number, keyword list, and syntax used for the values. The keyword list is the dictionary's keywords used to detect the corresponding statements. The syntax used for the values ​​is as described later. Figure 24 The grammar recognition dictionary corresponds to the grammar sequence number. The first keyword recognition dictionary 220 stores keywords indicating the endpoint of the input object range when the entries for each sequence number are input object items. That is, for example, when the sequence number is "2", as items that can input a common value, "dirt", "defect", and "printed characters" can be listed as three items. Therefore, the keyword "summarize up to defect" indicates the input positions (D3, D4) for "dirt" and "defect". In addition, the keyword "summarize up to print" indicates the input positions (D3~D5) for "dirt", "defect", and "printed characters".

[0160] Figure 23 The second keyword recognition dictionary 230 shown corresponds to the second embodiment. Figure 19 The second range dictionary 190 shown corresponds to the sequence number, keyword list, and syntax used for the value. The second keyword recognition dictionary 230 is used to detect statements that are unrelated to the sequence number, relate to the range of specified start and end points that can be entered at any time, and statements from specified groups.

[0161] That is, the first keyword recognition dictionary 220 and the second keyword recognition dictionary 230 are used to detect specified statements. The first keyword recognition dictionary 220 and the second keyword recognition dictionary 230 can also be derived from the second embodiment. Figure 18 The first range dictionary 180 shown and Figure 19 The syntax of the range-specified portion in the second range dictionary 190 shown is generated by expanding the non-terminal symbols. Furthermore, both the first keyword recognition dictionary 220 and the second keyword recognition dictionary 230 are equally capable of performing input processing when either one is generated.

[0162] Next, refer to Figure 24 This illustrates an example of the syntax recognition dictionary of the third embodiment.

[0163] Figure 24 The grammar recognition dictionary 240 shown is the same as that in the second embodiment. Figure 17 The value dictionary 170 shown is the same as the dictionary used to identify value statements. The syntax recognition dictionary 240 includes sequence numbers and syntax templates. The syntax templates are similar to... Figure 17 The syntax templates shown are the same. Figure 24 In the example, for the input format with sequence numbers "2~4", refer to... Figure 16 The input sequence list 160 shown is either "No abnormality" or "Needs to be replaced", so it is set to "No abnormality | Needs to be replaced | Skip".

[0164] "Skip" refers to skipping input within a range of input objects. The input unit 109 may also, in the case of a skipped input as a value statement, not perform any input within the range of input objects, or input a predetermined symbol such as "N / A". Furthermore, for ease of explanation, "skip" is described in the syntax recognition dictionary of the third embodiment, but the input unit 109 can perform the same processing even when it is determined to be a value statement in the first embodiment and in the value dictionary 170 included in the second embodiment.

[0165] In addition, Figure 24 In the example, the summary records the sequence numbers of groups that are common input forms, but grammar templates can also be set separately for each sequence number. Alternatively, the grammar recognition dictionary 240 can be generated for each input form, not for each sequence number.

[0166] Next, refer to Figure 25 The flowchart illustrates the input processing of the information processing device 10 in the third embodiment.

[0167] In step S2501, the control unit 108 emphasizes the input position of the input item and then displays the content that can be recognized by sound to the user based on the current input format of the input item. For example, if the input format is "word (no abnormality | needs to be replaced)", the text "no abnormality, needs to be replaced" can be displayed on the screen. The display location can be, for example, on the document data, or outside the status bar, or displayed in another window. Alternatively, it is not limited to displaying text; a synthesized sound such as "Please say whether there is no abnormality or needs to be replaced" can be generated and played to notify the user.

[0168] In step S2502, the voice recognition unit 103 begins voice recognition processing using a keyword recognition dictionary corresponding to the input sequence number, and also begins voice recognition processing using a syntax recognition dictionary corresponding to the current input sequence number. Furthermore, hereinafter, voice recognition processing using a keyword recognition dictionary will be referred to as keyword detection, and voice recognition processing using a syntax recognition dictionary will be referred to as syntax recognition. Specifically, the keyword recognition dictionary includes a first keyword recognition dictionary and a second keyword recognition dictionary. The syntax recognition dictionary is a value dictionary.

[0169] In step S2503, the determination unit 107 determines whether a keyword has been detected. That is, it determines whether the user speaks a keyword contained in the keyword recognition dictionary. If the user speaks a keyword contained in the keyword recognition dictionary, it is determined that a keyword has been detected. If a keyword is detected, the process proceeds to step S2504; if no keyword is detected, the process proceeds to step S2508.

[0170] In step S2504, the voice recognition unit 103 temporarily stops the grammar recognition, that is, the voice recognition processing using the grammar recognition dictionary.

[0171] In step S2505, the decision unit 106 determines the input object range based on the detected keywords. The keyword recognition dictionary contains a list of keywords indicating the range of statements, so the decision unit 106 can determine the input object range from the number of keywords, the start point, and the end point of the string. Specifically, for example, if the keyword "summarize until the defect" is detected, according to the keyword recognition dictionary, "defect" represents the end point, so the input object range is the range that starts at the current input object item's input position and ends at the input position corresponding to the item "defect".

[0172] In step S2506, the control unit 108 emphasizes the range of input objects. At this time, the "input format" in the input sequence list corresponding to the specified range of input objects may also be displayed as the content of the currently recognizable value.

[0173] In step S2507, in order to guard against speech from the user, the voice recognition unit 103, similar to step S2502, uses the first keyword recognition dictionary and the second keyword recognition dictionary corresponding to the current sequence number, and uses the grammar template that uses the keyword detection to apply the value corresponding to the keyword detected in step 2503, to start grammar recognition respectively.

[0174] In step S2508, the determination unit 107 determines whether a sound recognition result based on a grammar recognition dictionary has been obtained. If a sound recognition result is obtained, the process proceeds to step S2509; otherwise, it returns to step S2503 and repeats the same process.

[0175] In step S2509, the voice recognition unit 103 stops the voice recognition processing that uses the grammar recognition dictionary.

[0176] In step S2510, the decision unit 106 sets the voice recognition result output by the syntax recognition as a value speech.

[0177] Then, similarly to the first and second embodiments, a synthesized sound related to the confirmation message is played. If the input content is determined, the input unit 109 inputs a string related to the value statement at the input position of the document data. Afterwards, the entry in the input sequence list related to the input position where the value was input is set to "inputted," and the process is repeated for the next uninputted input position. Figure 25 The input processing is shown above. This concludes the input processing of the information processing apparatus 10 according to the third embodiment.

[0178] Next, refer to Figure 26 as well as Figure 27 This section describes a specific example of the voice recognition processing in the third embodiment. Figure 26 as well as Figure 27 These are diagrams illustrating the timing of the voice recognition processes for keyword detection and grammatical recognition along the time series.

[0179] Figure 26 In the case of an entry numbered "2" as an input item, imagine a scenario where the user specifies the range of statements.

[0180] In the first keyword recognition dictionary, for example, using and Figure 22 The list of keywords related to the sequence number "2" is used in the second keyword recognition dictionary. Figure 23 The entire keyword list is used to begin keyword detection. On the other hand, in the grammar recognition dictionary, [the following is used]... Figure 24 The grammar templates corresponding to the sequence numbers "2~4" are used to begin grammar pattern recognition. In addition, the speech is also recorded.

[0181] exist Figure 26In the example, it is assumed that in section 261, the user says "Summarize until printing". In this case, it matches the keyword list "Summarize until printing" in the first keyword recognition dictionary 220, so the speech "Summarize until printing" is detected as a keyword. As a result of detecting the keyword, while maintaining the state of continuous recording, the voice recognition processes of keyword detection and grammatical pattern recognition are each stopped. Additionally, in addition to recording, the voices of the subsequent speeches until the voice recognition process related to grammatical pattern recognition starts again can also be cached.

[0182] The detected keyword is a keyword indicating the end point. When referring to the input sequence list, the input position "D5" corresponding to "printing" included in the detected keyword is extracted. Therefore, the determination unit can determine the input object range "D3 - D5" with the input position "D3" of the input object item as the starting point and the input position "D5" as the end point.

[0183] In keyword detection, there is no need to determine the voice detection interval. So even in the middle of a speech, after the keyword detection is completed, the input object range corresponding to the speech can be immediately updated and emphasized on the document data. Additionally, a string in the input form that can be input within the updated input object range can also be displayed.

[0184] Next, in section 262, the voice recognition process related to keyword detection starts again. Additionally, the grammatical template with the sequence number corresponding to the detected keyword in the grammatical recognition dictionary 240 shown in Figure 24 (the grammatical template corresponding to the sequence numbers "2 - 4" of Figure 24 ) is used to start the voice recognition process related to grammatical pattern recognition again. Here, since the input object range has been determined, the restarted voice recognition process is not rejected, and the voice recognition process related to grammatical pattern recognition continues until a voice recognition result is generated. Here, it is assumed that the speech "No abnormality" is generated as the voice recognition result. In this case, in the grammatical recognition dictionary 240, it matches the grammatical template with the sequence numbers "2 - 4" of the input object range, so "No abnormality" is generated as the voice recognition result recognized by the grammatical pattern. When the speech "No abnormality" is detected, various determinations are made until the input ends in section 263, and the voice recognition processes related to keyword detection and grammatical pattern recognition are temporarily stopped.

[0185] The string "No abnormality" detected in the grammatical pattern recognition type is a value speech, so it is input to the input object range "D3 - D5".

[0186] Next, Figure 27In the case where the input object is an item numbered "2", imagine a scenario where the user only makes a value statement.

[0187] exist Figure 27 In the voice recognition processing involving keyword detection in interval 271, no keywords are detected, and no voice recognition result is generated. On the other hand, in the grammar recognition-type voice recognition processing, a "no abnormality" statement is detected, and a voice recognition result is generated. With a voice recognition result generated, in interval 272, until each judgment and input ends and then resumes, the voice recognition processing for both keyword detection and grammar recognition types is stopped. Since it is a value statement, the input object range remains the input position "D3" corresponding to sequence number 2, and the string "no abnormality" is input into the input object range "D3".

[0188] According to the third embodiment described above, the voice recognition processing involved in the grammar recognition shown in the second embodiment is used for the detection of value-based speech, and the voice recognition processing related to keyword detection, which does not require determining the voice detection interval, is used for the detection of range-specified speech. Therefore, when range-specified speech is recognized, the display switches to the corresponding input object range, allowing the user to immediately determine whether the spoken input object range is as desired; if not, the user repeats the speech. Furthermore, content based on the input format can be displayed when the input object range is updated, eliminating the situation where the user doesn't know what to say, and allowing the user to easily grasp the content to be said. Thus, similar to the first embodiment, the efficiency and convenience of voice-based data input can be improved.

[0189] (Fourth implementation)

[0190] In the third embodiment, it is envisioned that range-specified statements are detected by voice recognition processing related to keyword detection, but in the fourth embodiment, it is envisioned that in the voice recognition processing related to keyword detection, range-specified statements and value statements are detected by voice recognition processing related to syntax recognition to detect the end of the range-specified statement.

[0191] Reference Figure 28 A block diagram illustrating the information processing apparatus 10 of the fourth embodiment.

[0192] The information processing apparatus 10 of the fourth embodiment includes, in addition to the information processing apparatus of the third embodiment, a buffer unit 281.

[0193] The caching unit 281 caches audio data related to the user's speech in a manner that allows for at least a predetermined period of backtracking.

[0194] Next, Figure 29 An example of a keyword recognition dictionary for keyword detection according to the fourth embodiment is shown.

[0195] Figure 29 The keyword recognition dictionary shown is configured to detect the end of a specified range of statements. For example, keywords related to the end part such as "summary" or "summary up to the end" can be used. The keyword recognition dictionary can be generated, for example, by extracting the end part of a specified range of templates. Furthermore, keyword detection can be performed up to where the end part is extracted. For example, as long as detection accuracy can be ensured in application, the term "summary" can be set as a keyword, or the term of the end part can be set to be longer if the term "summary" is too short and the detection accuracy deteriorates.

[0196] Reference Figure 30 as well as Figure 31 This describes the syntax recognition dictionary for syntax recognition in the fourth embodiment.

[0197] As a syntax recognition dictionary, it is a range dictionary used to recognize by backtracking sound when a keyword is detected. It uses a first syntax recognition dictionary related to the input range starting from the input object item, a second syntax recognition dictionary related to the input range that can be detected at any time, and a value dictionary.

[0198] Figure 30 This is an example from the first grammar recognition dictionary 300. Unlike the first grammar recognition dictionary 300, which records grammar points instead of keywords, it also contains... Figure 22 The first keyword recognition dictionary 220 shown is the same.

[0199] Figure 31 This is an example of the second grammar recognition dictionary 310. The second grammar recognition dictionary 310 differs from other dictionaries in that it is not a keyword but rather a grammar point, but... Figure 23 The second keyword recognition dictionary 230 shown is the same. The value dictionary uses... Figure 24 The syntax recognition dictionary shown is sufficient.

[0200] Next, Figure 32 An example of a syntax recognition dictionary of the fourth embodiment is shown.

[0201] Figure 32 The grammar recognition dictionary 320 shown, for example, is obtained from... Figure 30 The first grammar recognition dictionary 300 shown extracts grammar templates starting with the sequence number "2", from... Figure 31 The second grammar recognition dictionary 310 shown is generated by extracting all grammar templates. Alternatively, it can be generated as follows: Figure 32In this way, instead of summarizing them into one dictionary, the keyword recognition dictionary 290, the first grammar recognition dictionary 300, the second grammar recognition dictionary 310, and the value dictionary are used separately as the sound recognition dictionary.

[0202] Next, refer to Figure 33 The flowchart illustrates the input processing of the information processing apparatus of the fourth embodiment.

[0203] Furthermore, in the information processing apparatus 10 of the fourth embodiment, the buffer unit 281 buffers a predetermined period T of sound. The buffer may also maintain at least the most up-to-date predetermined period T, discarding past sound that exceeds the predetermined period T. The length of the predetermined period T may be a predetermined time length such as 30 seconds, or it may be the length of "longest number of beats × 1 beat" in a speech pattern used to specify the range of input objects, such as guiding phrases in the input sequence list, or it may be a length calculated based on these values.

[0204] In step S3301, the voice recognition unit 103 uses a keyword recognition dictionary to begin grammatical recognition using a grammar recognition dictionary corresponding to the current order for keyword detection. Specifically, the grammar recognition dictionary includes a value dictionary.

[0205] In step S3302, when the voice recognition unit 103 detects a keyword in step S2503, it traces back from the current time point to the voice data corresponding to a predetermined period T. For the cached voice data within the predetermined period T, it uses the voice data from the beginning of the period T to the end of the keyword as a voice interval and performs grammatical recognition using a range dictionary (first grammar recognition dictionary, second grammar recognition dictionary) corresponding to the current order. Furthermore, it is considered that depending on the setting of the predetermined period T, there may sometimes be multiple voice intervals within the cached voice data. In this case, it is envisioned that multiple voice recognition results could be obtained, but the voice recognition result corresponding to the latest voice interval is sufficient.

[0206] In step S3303, the voice recognition unit 103 uses a keyword recognition dictionary and a value dictionary corresponding to the determined range of input objects to begin grammatical recognition for keyword detection. Subsequent processing is the same as the input processing in the third embodiment.

[0207] Next, refer to Figure 34 This section describes a specific example of the voice recognition processing in the fourth embodiment. Figure 34 This is a diagram showing the timing of the voice recognition processing for keyword detection and grammatical recognition along the time series.

[0208] Here, it was executed and used. Figure 29 The keyword detection of the keyword recognition dictionary 290 shown uses the same method as... Figure 24 The same value dictionary is used for voice recognition processing related to grammar recognition, and the buffer unit 281 buffers the voice for at least a predetermined period T. Let's say the interval 341 starts when the user says "summarize until printing". When referring to the keyword recognition dictionary 290, "summarize until printing" exists in the keyword list, so the statement "summarize until printing" is detected as a keyword. After detecting the statement "summarize until printing", the voice recognition processing related to keyword detection and the voice recognition processing related to grammar recognition are stopped.

[0209] The voice recognition unit 103 takes "summarize up to" as the end of the voice interval, and performs voice recognition processing related to grammatical type recognition using the range dictionary (first grammar recognition dictionary 300 and second grammar recognition dictionary 310) on the cached voice from the time of the predetermined backtracking period 342 up to that end. Figure 34 In this case, the result is set to "summarize until the text is printed". Therefore, compared with... Figure 30 The entries numbered "3" in the first grammar recognition dictionary 300 shown are consistent, so "summarize until printing" can be detected as a range-specified statement.

[0210] Subsequently, in interval 343, voice recognition processing begins again, involving keyword detection using keyword recognition dictionary 290 and grammatical recognition using a value dictionary corresponding to the range-specified speech within the predetermined period 342. Specifically, for the range-specified speech "summarize until printing," the grammar used for the corresponding entry in the first grammatical recognition dictionary is "2-4," so the grammar used is... Figure 24 The syntax template "No anomaly | Need to be replaced | Skip" in the value dictionary sequence number "2-4" is used to perform syntax identification. Here, there is a statement "No anomaly" in interval 343, which is consistent with the syntax template "2-4", so "No anomaly" can be detected as a value statement.

[0211] Furthermore, if no sound recognition result is generated through the sound recognition processing related to grammatical recognition within the predetermined period 342, it is set to not being a range-specified speech, and the sound recognition processing related to keyword detection and the sound recognition processing related to grammatical recognition using a value dictionary can be restarted.

[0212] Alternatively, if a keyword is detected in interval 343 through voice recognition processing related to keyword detection, the cached voice can be backtracked for a certain period based on the end of the speech related to that keyword, and grammatical recognition using a range dictionary can be performed.

[0213] According to the fourth embodiment described above, grammatical recognition is performed on audio cached for a predetermined period of time, based on the audio recognition result from keyword detection, and the range of input objects is determined from the range-specific speech. This reduces the amount of keyword list contained in the keyword recognition dictionary used in the audio recognition processing related to keyword detection. Therefore, in addition to the same effects as the third embodiment, compared to keyword detection, using a grammatical recognition dictionary capable of detecting various forms of speech increases the ability to detect patterns that specify a range of speech.

[0214] (Fifth implementation)

[0215] In the above embodiments, the range of input objects related to the input position used to input values ​​for document data is determined. However, in the case of creating daily reports of document data, it is also envisioned that a situation may arise where values ​​input to previously generated document data are to be copied. Therefore, in the fifth embodiment, the range of copyable values ​​to be input to the input position is determined.

[0216] Reference Figure 35 This explains the operation of the information processing device 10 in the fifth embodiment.

[0217] Figure 35 The upper layer is the document data in the current input with the inspection date of "2021 / 02 / 15", and the lower layer is the document data of the past input value with the inspection date of "2021 / 02 / 01".

[0218] The information processing device 10, for example, is triggered by a specific term and copies values ​​from past document data, thus performing processing to determine the scope of the copy target 351. For example, the information processing device 10 may also switch to a mode that sets the scope of the copy target 351 for the document data that is the copy source when the voice recognition unit 103 generates a voice recognition result such as "copy mode".

[0219] The range of objects to be copied, 351, can be determined in the same way as the range of input objects shown in the above embodiment. That is, if an identifier for the input position is used, it can be expressed as "D3 to D5", and if... Figure 35 The test numbers shown can be expressed as "the 1st to the 3rd".

[0220] After determining the copy target range 351, the input target range 352 "D3~D5" is determined through the input processing of the information processing device 10 of the first to fourth embodiments. When the input target range 352 is determined, the input unit 109 copies the value of the copy target range 351 to the input target range 352. Specifically, for the input target range 352, "no abnormality" is input at D3 and D4, and "slightly light" is input at D5.

[0221] Furthermore, the order in which the copy object range 351 and the input object range 352 are set is not limited; the copy object range 351 can be determined after the input object range 352 is determined. Additionally, the data used as the copy source for setting the copy object range 351 is not limited to past document data; it can also be copied from other input locations in the currently entered document data. Alternatively, different data formats such as text files can be used as the copy source.

[0222] According to the fifth embodiment described above, the information processing device sets an input object range, determines a copy object range related to the values ​​copied to the input object range, and copies the values ​​in the copy object range to the input object range. Thus, similar to the embodiments described above, the efficiency and convenience of voice-based data input can be improved.

[0223] Next, Figure 36 The block diagram illustrates an example of the hardware structure of the information processing apparatus 10 described in the above embodiments.

[0224] The information processing device 10 includes a CPU (Central Processing Unit) 361, RAM (Random Access Memory) 362, ROM (Read Only Memory) 363, a storage device 364, a display device 365, an input device 366, and a communication device 367, which are connected by a bus.

[0225] CPU 361 is a processor that performs arithmetic and control processing according to a program. CPU 361 uses a predetermined area of ​​RAM 362 as its working area and performs the processing of each part of the information processing device 10 in cooperation with the program stored in ROM 363 and storage device 364.

[0226] RAM362 is a type of memory such as SDRAM (Synchronous Dynamic Random Access Memory). RAM362 functions as the operating area of ​​CPU361. ROM363 is a non-rewritable memory that stores programs and various other information.

[0227] Storage device 364 is a device for writing and reading data into magnetic recording media such as HDD (Hard Disc Drive), semiconductor-based storage media such as flash memory, or storage media capable of magnetic recording such as HDD, or storage media capable of optical recording. Storage device 364 writes and reads data into the storage medium under the control of CPU 361.

[0228] Display device 365 is a display device such as an LCD (Liquid Crystal Display). Display device 365 displays various information based on display signals from CPU 361.

[0229] Input device 366 is an input device such as a mouse and keyboard. Input device 366 accepts information input from user operation as an indication signal and outputs the indication signal to CPU 361.

[0230] The communication device 367 communicates with external devices via a network under the control of the CPU 361.

[0231] The instructions shown in the processing sequence of the above embodiments can be executed according to a program that is software. A general-purpose computer system pre-stores this program, reads it, and thus achieves the same effect as the control operation based on the information processing device described above. The instructions described in the above embodiments are recorded as a program that can be executed by a computer on a disk (flexible optical disc, hard disk, etc.), optical disc (CD-ROM, CD-R, CD-RW, DVD-ROM, DVD±R, DVD±RW, Blu-ray Disc, etc.), semiconductor memory, or similar recording media. The storage format can be arbitrary, as long as the recording medium can be read by a computer or embedded system. As long as the computer reads the program from the recording medium, the CPU executes the instructions described in the program according to the program, and thus achieves the same operation as the control operation of the information processing device described above. Of course, the program can also be obtained or read via a network when the computer acquires or reads it.

[0232] Alternatively, according to the instructions of the program installed from the recording medium onto the computer or embedded system, the OS (operating system), database management software, network, and other MW (middleware) running on the computer may execute a portion of the processes used to implement this embodiment.

[0233] Furthermore, the recording medium in this embodiment is not limited to media independent of the computer or embedded system, but also includes recording media that store or temporarily store programs downloaded and transmitted via LAN, the Internet, etc.

[0234] Furthermore, when the recording medium is not limited to one and the processing in this embodiment is performed from multiple media, the structure of the recording medium included in this embodiment can also be arbitrary.

[0235] Furthermore, the computer or embedded system in this embodiment is used to execute the various processes in this embodiment according to the program stored in the recording medium, and can be any structure including a personal computer, a microcomputer, a system in which multiple devices are connected by a network.

[0236] Furthermore, the computer in this embodiment is not limited to personal computers, but also includes arithmetic processing devices, microcomputers, etc., included in information processing equipment, and is collectively referred to as any device or apparatus that can implement the functions of this embodiment through a program.

[0237] Several embodiments of the present invention have been described, but these embodiments are provided by way of example and are not intended to limit the scope of the invention. These new embodiments can be implemented in various other ways, and various omissions, substitutions, and modifications can be made without departing from the spirit of the invention. These embodiments and their variations are included in the scope and spirit of the invention, and are included in the scope of the invention as set forth in the patent claims and its equivalents.

Claims

1. An information processing device, comprising: The generation department generates a template related to multiple items based on the input order of the input object items selected from the multiple items, using a data table containing records containing multiple items as a reference. The voice recognition department performs voice recognition on users' speech and generates voice recognition results. as well as The decision-making unit, based on the template and the voice recognition results, determines the range of input objects related to the multiple items specified by the user's speech among the multiple items. The generation unit generates a keyword recognition dictionary for detecting specific keywords and a grammar recognition dictionary for performing voice recognition on speech based on specific grammar, according to the template. The voice recognition unit generates a first speech by the user that matches the keyword recognition dictionary as a first voice recognition result, and generates a second speech that matches the syntax recognition dictionary but is delivered later than the first speech as a second voice recognition result. The information processing device further includes a determination unit, which determines the first voice recognition result as a range-specified speech for specifying the input object range, and determines the second voice recognition result as a value speech representing a value input into the input object range.

2. The information processing apparatus according to claim 1, wherein, The information processing device further includes a determination unit that, when the sound recognition result includes a portion consistent with the template, determines that the speech related to the consistent portion is a range-specifying speech used to specify the input object range, and that the speech in the sound recognition result that is later than the consistent portion is a value speech indicating a value input into the input object range.

3. The information processing apparatus according to claim 2, wherein, If the sound recognition result does not include a portion that matches the template, the determination unit determines that the speech related to the sound recognition result is the value speech.

4. The information processing apparatus according to claim 1, wherein, The generation unit generates a grammar recognition dictionary based on the template for voice recognition of speech based on specific grammar. The voice recognition unit generates the user's speech, which is consistent with the grammar recognition dictionary, as the voice recognition result.

5. The information processing apparatus according to claim 1, wherein, The information processing device also includes a buffer unit and a judgment unit. The buffer unit buffers the user's speech as audio data. The generation unit generates a keyword recognition dictionary for detecting specific keywords and a grammar recognition dictionary for performing voice recognition on speech based on specific grammar, according to the template. The voice recognition unit generates a first speech by the user that matches the keyword recognition dictionary as a first voice recognition result. Using the cached voice data, it generates a second speech that matches the syntax recognition dictionary from voice data retrieved from a predetermined period of voice data corresponding to the first voice recognition result as a second voice recognition result. The determination unit determines the second voice recognition result as a range-specified speech used to specify the range of the input object.

6. The information processing apparatus according to claim 5, wherein, The voice recognition unit generates a third speech that matches the grammar recognition dictionary from the voice data that is located later in the voice data portion corresponding to the first voice recognition result, and uses this as the third voice recognition result. The determination unit determines the third voice recognition result as a value speech that represents a value input into the range of the input object.

7. The information processing apparatus according to claim 2, wherein, The range of input objects is the range that determines the input positions on the data table used for recording. The information processing device also includes an input unit that inputs values ​​related to the value statement into the input location.

8. The information processing apparatus according to claim 7, wherein, The determination unit determines whether the value statement is a statement intended to skip input for the range of input objects.

9. The information processing apparatus according to claim 8, wherein, If the input unit determines that the value statement is intended to skip the input, it will not input a value or will input a predetermined symbol within the input object range.

10. The information processing apparatus according to claim 1, wherein, The information processing device also includes a control unit that highlights the range of input objects on the recording data table.

11. The information processing apparatus according to claim 1, wherein, The template is an expression consistent with at least one of the following: the endpoint of the input object range, the start and end points of the input object range, the number of items included in the input object range, and the names of the items included in the input object range.

12. The information processing apparatus according to claim 1, wherein, The input order corresponds to at least one of the following: the project name, the identifier of the input position of the value for that project, and the group name of the project. The template uses at least one of the name, the identifier, and the group name that corresponds to the input order to indicate the input location.

13. An information processing method, wherein, For a data table containing records of multiple items, a template is generated based on the input order related to the input object items selected from the multiple items, and that may be specified for the multiple items. Perform voice recognition on the user's speech and generate voice recognition results. Based on the template and the voice recognition results, determine the range of input objects related to the multiple items specified by the user's speech among the multiple items. The steps of generating a template include: generating a keyword recognition dictionary for detecting specific keywords and a grammar recognition dictionary for performing voice recognition on speech based on specific grammar, according to the template. The step of generating the voice recognition result includes: generating a first speech of the user that is consistent with the keyword recognition dictionary as a first voice recognition result, and generating a second speech that is consistent with the syntax recognition dictionary but follows the first speech as a second voice recognition result. The information processing method further determines the first voice recognition result as a range-specified speech used to specify the range of the input object, and determines the second voice recognition result as a value speech representing a value input into the range of the input object.

14. A computer program product comprising an information processing program for enabling a computer to function as a unit: The generation unit generates a template related to multiple items that may be specified, based on the input order of the input object items selected from the multiple items, using a data table containing records of multiple items. The voice recognition unit performs voice recognition on the user's speech and generates voice recognition results; as well as The decision unit, based on the template and the voice recognition result, determines the range of input objects related to the multiple items specified by the user's speech among the multiple items. The generation unit generates a keyword recognition dictionary for detecting specific keywords and a grammar recognition dictionary for performing voice recognition on speech based on specific grammar, according to the template. The voice recognition unit generates a first speech by the user that matches the keyword recognition dictionary as a first voice recognition result, and generates a second speech that matches the syntax recognition dictionary but is delivered after the first speech as a second voice recognition result. It also enables the computer to function as a determination unit, which determines the first sound recognition result as a range-specified speech for specifying the range of the input object, and determines the second sound recognition result as a value speech representing the value input into the range of the input object.