Grammar creation support device and computer-readable storage medium
The grammar creation assistance device aids in creating and refining speech recognition grammars for industrial equipment by visualizing acoustic distances and linking words, enhancing accuracy and adaptability to industrial environments.
Patent Information
- Application Number
- JP2023575015
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-01-21
- Publication Date
- 2025-12-23
- Estimated Expiration
- 2042-01-21
AI Technical Summary
The accuracy of speech recognition grammars in industrial equipment is evaluated through text-based editing, lacking support for efficient creation and refinement of voice command grammars tailored to specific industrial environments.
A grammar creation assistance device and medium that includes a grammar storage unit, speech recognition unit, evaluation data storage, recognition result evaluation, and grammar processing unit, which assists in creating, evaluating, and refining grammars for industrial equipment by visualizing acoustic distances and linking words based on recognition results.
Enhances the creation and refinement of speech recognition grammars for industrial equipment, improving accuracy and adaptability to specific industrial environments by visualizing recognition results and allowing for targeted grammar modifications.
Smart Images

Figure 0007791215000001 
Figure 0007791215000002 
Figure 0007791215000003
Abstract
Description
[Technical Field]
[0001] The present invention relates to a grammar creation assistance device for speech recognition and a computer-readable storage medium. [Background technology]
[0002] Currently, various types of equipment are in operation in industrial fields such as manufacturing, including robots, conveyors, machine tools, and machinery. Many of these devices are equipped with operating units, and the devices that control them, such as PLCs (Programmable Logic Controllers), NCs (Numerical Controllers), and control panels, often also have operating units.
[0003] The operation sections of devices often have many buttons and operation screens, but their operation can be complex and require time to master. Voice input interfaces allow users to perform desired operations simply by uttering voice commands. For this reason, efforts are being made to improve operability by using voice input interfaces.
[0004] Voice commands used to operate devices can be predicted based on the type of device that uses the voice commands, the location where the device is installed, the operation of the device, etc. Therefore, predicted voice commands can be created using grammar (syntax and words). For example, see Patent Document 1. [Prior art documents] [Patent documents]
[0005] [Patent Document 1] Japanese Patent Application Publication No. 9-325787 Summary of the Invention [Problem to be solved by the invention]
[0006] The accuracy of the created grammar is evaluated using evaluation data. The creator of the speech recognition system checks the accuracy of speech recognition when the created grammar is used and edits the grammar. Speech recognition grammars are often written in text.
[0007] In the industrial field, there is a demand for technology that supports the creation of grammars for speech recognition. [Means for solving the problem]
[0008] A grammar creation assistance device according to one aspect of the present disclosure includes a grammar storage unit that stores grammars of voice commands for operating industrial equipment, a voice recognition unit that performs voice recognition based on the grammar, an evaluation data storage unit that stores evaluation data including voice data for evaluating the grammar and correct answer data for the evaluation voice data, and a storage unit that stores the results of recognition of the evaluation data by the voice recognition unit. By voice command type The system includes a recognition result evaluation unit that creates a summary, and a grammar processing unit that presents a summary of the evaluation of the recognition result in association with a grammar and accepts processing of the grammar. Furthermore, a grammar creation assistance device according to one aspect of the present disclosure includes a grammar storage unit that stores grammars for voice commands for operating industrial equipment, a speech recognition unit that performs speech recognition based on the grammar, an evaluation data storage unit that stores evaluation data including speech data for evaluating the grammar and correct answer data for the evaluation speech data, a recognition result evaluation unit that creates a summary of the recognition results of the evaluation data by the speech recognition unit, and a grammar processing unit that presents the summary of the evaluation of the recognition results in association with the grammar and accepts processing of the grammar, and the grammar processing unit visualizes the acoustic distances of the words that make up the grammar and connects the words with links. A storage medium according to one aspect of the present disclosure stores a grammar of voice commands for operating industrial equipment, and is executed by one or more processors to perform speech recognition on speech data for evaluation of the grammar based on the grammar, and based on the recognition result of the speech recognition and correct answer data of the evaluation speech data, By voice command type The apparatus stores processor-readable instructions for creating a summary of the recognition result, presenting the summary of the recognition result in association with the grammar, and accepting processing of the grammar. A storage medium according to one aspect of the present disclosure stores a grammar for voice commands for operating industrial equipment, and stores processor-readable instructions that, when executed by one or more processors, perform speech recognition of speech data for evaluating the grammar based on the grammar, create a summary of the recognition results based on the recognition results and correct answer data of the evaluation speech data, associate the summary of the recognition results with the grammar and present it, and accept grammar processing. The presentation visualizes the acoustic distances between words that make up the grammar and connects the words with links. . [Effects of the Invention]
[0009] One aspect of the present invention is to assist in creating grammars for speech recognition. [Brief explanation of the drawings]
[0010] [Figure 1] FIG. 1 is a block diagram showing the configuration of a grammar creation assistance device. [Figure 2]FIG. 10 is a diagram showing examples of syntax definitions and word definitions. [Figure 3] FIG. 10 is a diagram showing an example of a combination of speakers and recording locations of evaluation data. [Figure 4] FIG. 10 is a diagram showing an example of an evaluation result display screen. [Figure 5] FIG. 10 is a diagram illustrating an example of a history display screen. [Figure 6] FIG. 10 is a diagram showing an example of an image display of a grammar. [Figure 7] FIG. 10 is a diagram showing an example of processing a grammar. [Figure 8] 10 is a flowchart illustrating a process of the grammar creation assistance device. [Figure 9] 1 shows the hardware configuration of a grammar creation assistance device. DETAILED DESCRIPTION OF THE INVENTION
[0011] The grammar creation assistance device 100 will now be described. The grammar creation assistance device 100 is implemented in an information processing device having a calculation unit and a storage unit. Examples of such information processing devices include, but are not limited to, a PC (personal computer) and a mobile terminal.
[0012] 1 shows the basic configuration of a grammar creation assistance device 100. The grammar creation assistance device 100 is composed of an evaluation data storage unit 11, a target performance registration unit 12, a speech recognition unit 13, a grammar storage unit 14, a recognition result evaluation unit 15, a grammar processing unit 16, and an evaluation history storage unit 17.
[0013] The speech recognition unit 13 inputs speech data and outputs the recognition results of the input speech data in text format. The speech recognition unit 13 generally comprises an acoustic model, a language model, and a decoder. The acoustic model inputs speech data and outputs the phonemes (senones) that make up the speech data based on the features of the speech data. The language model outputs the occurrence probability of word strings. The language model selects hypothetical word strings based on the phonemes and outputs linguistically plausible candidates. The decoder outputs word strings with high probability as recognition results based on the outputs of the statistically created acoustic model and language model.
[0014] The grammar storage unit 14 stores the grammar of voice commands. Voice commands are voice commands for operating industrial equipment. The voice recognition unit 13 selects voice commands defined in the grammar. The grammar of a voice command consists of syntax and words. The grammar storage unit 14 includes a syntax storage unit 18 that stores syntax and a word storage unit 19 that stores words. The words include words to be recognized by voice recognition and their phonemic representations. The syntax defines the words that make up the voice command and their order. In this disclosure, a base grammar is first created. The base grammar is created comprehensively to cover as many voice commands as possible that are expected to be used in the field. The grammar creation assistance device 100 assists in creating an appropriate grammar by modifying the base grammar based on the recognition results of evaluation data. The base grammar is determined based on the type of equipment that recognizes the voice command, the work content, etc.
[0015] Figure 2 shows an example of a syntax definition and an example of a word definition. The example syntax definition defines the words that make up a voice command and the order of the words. In the first line of the syntax definition in Figure 2, "S:NS_B COMMAND NS_E," "S" is the start symbol of a voice command, and "NS_B" and "NS_E" are the silent intervals at the beginning and end of the sentence. The syntax element "COMMAND" exists between the silent intervals. The second and third lines define the "tags" that go into "COMMAND". The second line defines that the tags "ROBOT" and "INTERFACE" go into the syntactic element "COMMAND", and the third line defines that the tags "NAIGAI" and "INTERFACE" go into the syntactic element "COMMAND".
[0016] The first and second lines of the word definition define the Japanese spelling and phonemic spelling of the tag "ROBOT." The Japanese spelling of the tag "ROBOT" is "robot" and the phonemic spelling is "roboqto." Lines 3 to 5 of the word definition define the Japanese spelling and phonemic spelling of the Japanese word that goes into the tag "NAIGAI." The tag "NAIGAI" contains two Japanese words, "outside" and "inside." The phonemic spelling of "outside" is "gaibu," and the phonemic spelling of "inside" is "naibu." Lines 6 to 8 of the word definition define the Japanese spelling and phonemic spelling of the Japanese word that goes into the tag "INTERFACE." The tag "INTERFACE" contains one Japanese word, "interface." There are two phonemic spellings for "interface," "iNtafe:su" and "iNta:feisu." "%NS_B" defines the silent section [s] at the beginning of the sentence, and "%NS_E" defines the silent section [ / s] at the end of the sentence.
[0017] The evaluation data storage unit 11 stores voice data including voice commands recorded by multiple speakers at multiple recording locations in association with correct answer data, which is correct text for the voice data. For example, the evaluation data storage unit 11 stores voice data in which multiple speakers utter "external interface" at multiple recording locations in association with correct answer data (text) called "external interface." The evaluation data includes speech data recorded at different locations by speakers with different attributes (gender, age). Figure 3 is a table showing the relationship between the speakers and recording locations in the evaluation data. The evaluation data in Figure 3 includes speech recorded by Speaker A (male, 60 years old) at Factories A and B, and speech recorded by Speaker B (female, 30 years old) at Factories C and D.
[0018] The target performance registration unit 12 accepts registration of target performance for voice recognition. The target performance registration unit 12 accepts target values such as the accuracy rate of voice commands, the accuracy rate for each type of voice command, and the processing time (average value) for voice recognition. The registered contents of the target performance are reflected on the evaluation result display screen, which will be described later.
[0019] The recognition result evaluation unit 15 compares the correct text stored in the evaluation data storage unit with the recognition results of the voice data, creates a summary of the grammar evaluation results, and displays the summary on the display unit. Figure 4 shows an example of a recognition result display screen. In the example of Figure 4, an evaluation of the entire voice command and an evaluation of each voice command type are displayed. Voice command types include, for example, approval commands, numerical commands, and transition commands. Approval commands indicate approval. Approval commands include "Yes," "No," "Yes," "No," "Execute," and "Cancel." Numeric commands specify numerical values such as "0.5," "1," "2," and "100." "Transition commands" specify display screens such as the "Home screen" and the "Speed setting screen." Other examples include "machine operation commands" that instruct equipment operation, such as "Set the workpiece." The recognition result display screen may also display the voice recognition processing time. The target performance registered in the target performance registration unit may also be displayed.
[0020] The recognition result evaluation unit 15 may display the history of recognition results. FIG. 5 shows a history display screen. The history display screen allows selection of data related to past speech recognition. In the example of FIG. 5, the identification number of the evaluation result and the time when the speech recognition was performed are displayed. When the time or the identification number is selected, the evaluation of the selected speech recognition and the grammar used for the speech recognition are displayed. Note that the history display screen is not limited to the arrangement shown in FIG. 5, as long as it has a configuration that allows past recognition results to be compared and selected.
[0021] The grammar processing unit 16 accepts processing (editing) of the grammar. The creator of the grammar can process (edit) the grammar while checking the evaluation result of the speech recognition and the grammar corresponding to the evaluation result.
[0022] The grammar may be displayed as text or as an image. When the grammar is displayed as an image, the acoustic distance of the voice command is calculated and the words and paths of the words are connected by links. The acoustic distance may be calculated from the voice data or correct answer data of the evaluation data, or may be calculated from the phoneme representation of the grammar. An example of a visual display of the grammar is shown in Figure 6. Figure 6 is an example of a visual display of the syntax definition and word definition in Figure 2. In the grammar in Figure 2, the grammar element "COMMAND" contains the words "ROBOT", "INTERFACE", and "NAIGAI", "INTERFACE". The grammar processing unit 16 calculates the acoustic distance between these words. In the example of Figure 6, "naibu" and "gaibu", and "iNtafe:su" and "iNta:feisu" are acoustically close, so they are displayed in close positions. "roboqto" is acoustically distant from all other words, so it is displayed in a distant position. The grammar processing unit 16 arranges words that can be included in the syntax on the screen and connects the paths between those words with links. For example, in the example of Figure 6, the words that go into "ROBOT" and the words that go into "INTERFACE", and the words that go into "NAIGAI" and the words that go into "INTERFACE" are connected with links.
[0023] A known network visualization method is used to arrange the words. A spring model is exemplified as one of the network visualization methods. In the spring model of the present disclosure, words are regarded as nodes, and the acoustic distance between any two nodes is calculated. The acoustic distance between two nodes is regarded as the length of a spring, and the two nodes are arranged in space. After the words are arranged in a graph, syntax is used to Connect words with links.
[0024] It is also possible to visually represent areas where speech recognition errors are likely to occur, areas where phonemes are close, the accuracy rate between correct answer data and speech recognition results, the occurrence rate of words, and areas where phonemes match. An example of an area where phonemes match is the phoneme "aib" contained in "naibu" and "gaibu." An example of an area where phonemes are close is the phoneme "afe:" contained in "iNta:feisu" and the phoneme ":fei" contained in "iNta:feisu." In the example in Figure 6, these are highlighted using bold text. The high occurrence rate, high accuracy rate, etc. can also be represented by the size of the text.
[0025] Figure 7 shows an example of correcting the grammar in Figure 6. In Figure 7, the link for "naibu" has been removed. For example, if the grammar creator is experiencing misrecognition between "naibu" and "gaibu" and the specifications do not require the use of "naibu," they can remove the link for "naibu." If the specifications require the word "naibu," they can manually leave it in. In the grammar creation assistance device 100 of the present disclosure, words and sentence structures that cannot be removed from the specifications can be left in at the discretion of the creator.
[0026] Grammar modification and evaluation of recognition results are repeated. The grammar creator can check the evaluation of the recognition results (for example, accuracy rate) for the grammar modification, modify the grammar within the range that complies with the specifications, and customize the grammar.
[0027] The evaluation history storage unit 17 stores the recognition results in association with the grammar. When a grammar to be stored in the evaluation history storage unit 17 is selected, the evaluation result display screen shown in Figure 4 is displayed. The grammar creator processes the grammar while referring to summary information such as the accuracy rate of the speech recognition. As an example of a method for confirming the summary information, approval commands such as "yes" and "no" are used for final confirmation, so a high accuracy rate is required. Numeric commands that specify numerical values also require a high accuracy rate. Transition commands that specify screen transitions can have a lower accuracy rate than approval commands and numeric commands. The grammar creator can register such performance goals and modify the grammar while taking into account the needs of each site.
[0028] The processing of the grammar creation assistance device 100 will be described with reference to FIG. As a preparation step, the grammar creation assistance device 100 receives registration of target performance for speech recognition (step S1) and registration of the number of saved evaluation histories for speech recognition (step S2). The grammar creation device acquires evaluation data for the grammar (step S3).
[0029] The grammar creator creates a base grammar based on the specifications of the field. The base grammar is created as comprehensively as possible in accordance with the requests of the device user. The grammar creation assistance device 100 stores the base grammar (step S4).
[0030] The grammar creation assistance device 100 performs speech recognition on the evaluation data using the registered grammar (step S5). The grammar creation assistance device 100 summarizes the recognition result of step S5 and presents it to the creator (step S6). The creator checks the recognition result and, if they determine that the grammar is complete (step S7; YES), ends the creation of the grammar.
[0031] If the creator checks the recognition results and determines that the grammar needs to be modified (step S7; NO), the grammar creation assistance device 100 stores the previously created grammar and a summary of the recognition results in the recognition result storage unit and accepts grammar modification (step S8). The grammar creation assistance device 100 registers the grammar modified in step S8, proceeds to step S5, and performs speech recognition using the registered grammar. The grammar creator compares the grammar created in the past with the newly created grammar. The grammar creation assistance device 100 repeats the processes from step S5 to step S8 until the creator determines that the grammar is complete.
[0032] As described above, the grammar creation assistance device 100 of the present disclosure is a device that assists in the creation of grammars for voice commands, performs speech recognition of evaluation data using the created grammar, summarizes the recognition results of the evaluation data, and presents the summary results to the creator of the grammar. The recognition results for the evaluation data are calculated for all voice commands and for each type of voice command. The target performance differs for each type of voice command. The grammar creator can modify the grammar to achieve the target performance for each type of voice command.
[0033] Grammars can be displayed either as text or as images. When displayed as images, words (nodes) are linked according to the syntax, using the acoustic distance between words. Because words are arranged using acoustic distance, the grammatical structure can be visually determined.
[0034] The acoustic distance may be calculated from the speech data of the evaluation data, or may be calculated from phonemes expressed in text. Methods for calculating acoustic distance from speech data include distribution distance. Methods for calculating acoustic distance from phonemes expressed in text include cosine distance, Levenshtein distance, Jaro-Winkler distance, and Hamming distance. There are no limitations on the method for calculating acoustic distance. Cosine distance, Euclidean distance, Levenshtein distance, Jaro-Winkler distance, and Hamming distance are well known.
[0035] Industrial equipment is installed in noise-generating locations such as factories. Noise has characteristics that vary by location or time of day. In this disclosure, evaluation data is collected at the location where the equipment is installed, and evaluation is performed taking into account noise specific to the location.
[0036] When operating industrial equipment, there are technical terms specific to each site, and certain fixed terms are often used frequently. A comprehensively created grammar contains words and syntax that are not actually used, but it is difficult to know in advance the terms that will actually be used in the site. In this disclosure, a comprehensive grammar is created and words and syntax that are not used in the site are deleted to improve the accuracy rate of speech recognition. Furthermore, in this disclosure, rather than deleting all grammar that is used infrequently, it is also possible to retain words and grammar that are necessary for the specifications, even at the expense of accuracy rate. Note that words and syntax may be added as needed.
[0037] [Hardware configuration] The hardware configuration of the grammar creation assistance device 100 will be described with reference to Fig. 9. The grammar creation assistance device 100 includes a CPU 111, which is a processor that controls the entire device. The CPU 111 reads a system program stored in a ROM 112 via a bus, and controls the entire device 100 in accordance with the system program. The RAM 113 temporarily stores temporary calculation data, display data, various data input by the user via the input unit 71, and the like.
[0038] The display unit 70 is a monitor or the like attached to the grammar creation assistance device 100. The display unit 70 displays an operation screen, a setting screen, and the like of the grammar creation assistance device 100.
[0039] The input unit 71 is a keyboard, a touch panel, an operation button, or the like that is integrated with the display unit 70 or is separate from the display unit 70. The user operates the input unit 71 to input data to the screen displayed on the display unit 70. The display unit 70 and the input unit 71 may be mobile terminals.
[0040] The nonvolatile memory 114 is a memory that retains its stored state even when the power to the grammar creation assistance device 100 is turned off, for example, by being backed up by a battery (not shown). The nonvolatile memory 114 stores machining programs, system programs, available options, a charge table, etc. The nonvolatile memory 114 stores programs read from external devices via an interface (not shown), programs input via the input unit 71, and various data acquired from each unit of the grammar creation assistance device 100, machine tools, etc. (e.g., setting parameters acquired from machine tools, etc.). The programs and various data stored in the nonvolatile memory 114 may be loaded into the RAM 113 when executed / used. Furthermore, various system programs are written in the ROM 112 in advance. [Explanation of symbols]
[0041] 100 Grammar Creation Support Device 11 Evaluation data storage unit 12 Target performance registration section 13 Voice Recognition Unit 14 Grammar storage 15 Recognition result evaluation unit 16 Grammar Processing Department 17 Evaluation history memory unit 18 Syntactic memory 19 Word Memory Section 70 Display section 71 Input section 111 CPU 112 ROM 113 RAM 114 Non-volatile memory
Claims
1. a grammar storage unit that stores grammars of voice commands for operating industrial equipment; a speech recognition unit that performs speech recognition based on the grammar; an evaluation data storage unit that stores evaluation data including speech data for evaluation of the grammar and correct answer data for the speech data for evaluation; a recognition result evaluation unit that creates a summary of the recognition result of the evaluation data by the speech recognition unit for each type of voice command; a grammar processing unit that presents a summary of the evaluation of the recognition result in association with the grammar and accepts processing of the grammar; A grammar creation assistance device comprising:
2. a grammar storage unit that stores grammars of voice commands for operating industrial equipment; a speech recognition unit that performs speech recognition based on the grammar; an evaluation data storage unit that stores evaluation data including speech data for evaluation of the grammar and correct answer data for the speech data for evaluation; a recognition result evaluation unit that creates a summary of the recognition result of the evaluation data by the speech recognition unit; a grammar processing unit that presents a summary of the evaluation of the recognition result in association with the grammar and accepts processing of the grammar; Equipped with the grammar processing unit visualizes the acoustic distances of the words that make up the grammar and connects the words with links; Grammar creation aid.
3. 3. The grammar creation assistance device according to claim 2, wherein the grammar processing unit accepts deletion or addition of the words or links between the words.
4. 3. The grammar creation assistance device according to claim 1, wherein the summary includes at least one of an accuracy rate of speech recognition and a processing time of speech recognition.
5. 3. The grammar creation assistance device according to claim 1, further comprising an evaluation history storage unit that stores a history of at least one of the recognition results and the summaries.
6. 6. The grammar creation assistance device according to claim 5, wherein a plurality of recognition results or summaries stored in said evaluation history storage unit are presented in a comparable format.
7. It memorizes the grammar of voice commands to operate industrial equipment, When executed by one or more processors, performing speech recognition on speech data for evaluating the grammar based on the grammar; creating a summary of the recognition results for each type of voice command based on the recognition results of the voice recognition and correct answer data of the evaluation voice data; presenting a summary of the recognition result in association with the grammar, and accepting requests to modify the grammar; A storage medium that stores instructions readable by the processor.
8. It memorizes the grammar of voice commands to operate industrial equipment, When executed by one or more processors, performing speech recognition on speech data for evaluating the grammar based on the grammar; creating a summary of the recognition result based on the recognition result of the speech recognition and the correct answer data of the speech data for evaluation; presenting a summary of the recognition result in association with the grammar, and accepting requests to modify the grammar; a storage medium storing instructions readable by the processor, The presentation visualizes the acoustic distances between the words that make up the grammar and connects the words with links. A storage medium that stores instructions readable by the processor.
Citation Information
Patent Citations
horumuarudehidooganjusuruhaisuinojodokuhoho
JP1976061174A
Voice synthesizing method, voice synthesizing device, method and device for incorporating voice command in sentence
JP1997325787A
Method for preparing recognition grammar model and inspection method
JP2004151547A
Speech recognition device and speech recognition method
JP2009229529A
Dictionary update device and program
JP2018040906A