Grammar adjustment device and computer-readable storage medium

The grammar adjustment device optimizes speech recognition grammars for industrial devices by automating extraction, evaluation, and selection, enhancing accuracy and adaptability to site-specific conditions.

JP7865995B2Active Publication Date: 2026-05-26FANUC LTD

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
FANUC LTD
Filing Date
2022-01-21
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

Existing speech recognition systems for industrial devices require manual creation of grammars, which are time-consuming and lack objective evaluation methods for accuracy improvement.

Method used

A grammar adjustment device and storage medium that automatically extract, evaluate, and select grammars using a k-means method, speech recognition, and evaluation data to optimize voice command recognition for industrial equipment.

Benefits of technology

Automated grammar creation and selection improve speech recognition accuracy for industrial devices, adapting to site-specific noise and improving usability without relying on subjective expertise.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007865995000001
    Figure 0007865995000001
  • Figure 0007865995000002
    Figure 0007865995000002
  • Figure 0007865995000003
    Figure 0007865995000003
Patent Text Reader

Abstract

The present invention stores the grammar of voice commands that operate industrial machinery, uses one or more processors to extract a portion of the grammar, receives registration of a target for an evaluation value for voice recognition of the extracted grammar, uses the extracted grammar to perform voice recognition on voice data for evaluation, calculates an evaluation value for voice recognition of the extracted grammar on the basis of the results of the voice recognition that uses the extracted grammar and of correct answer data for the voice data for evaluation, and selects grammar that satisfies the target from the extracted grammar.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a grammar adjustment device for speech recognition and a computer-readable storage medium.

Background Art

[0002] Currently, in industrial fields such as manufacturing, various devices such as robots, conveyors, machine tools, and machinery and equipment are operating. Many of these devices have an operation unit, and many of the devices themselves for controlling each device, such as a PLC (Programmable Logic Controller), NC (Numerical Controller), and control panel, also have an operation unit.

[0003] The operation units of devices have many buttons and operation screens, but the operations can be complex and may take time to master. The voice input interface can execute the target operation simply by uttering a voice command. Therefore, attempts have been made to improve the operability by using the voice input interface.

[0004] The voice commands used for operating the device can be assumed based on the type of device using the voice command, the site where the device is installed, the operation content of the device, etc. Therefore, the assumed voice commands can be created in grammar (syntax and words). For example, refer to Patent Document 1.

Prior Art Documents

Patent Documents

[0005]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0006] The accuracy of the created grammar is evaluated using evaluation data. The creator of the speech recognition system checks the accuracy of speech recognition when using the created grammar and edits the grammar accordingly. Speech recognition grammars are often created manually.

[0007] In the industrial sector, there is a demand for technologies that support the creation of speech recognition grammars. [Means for solving the problem]

[0008] A grammar adjustment device according to one aspect of the present disclosure includes: a grammar storage unit that stores the grammar of voice commands for operating industrial equipment; a grammar extraction unit that extracts a part of the grammar; a target registration unit that accepts registration of a target for the evaluation value of the speech recognition of the extracted grammar; a speech recognition unit that performs speech recognition of evaluation voice data using the extracted grammar; an evaluation value calculation unit that calculates an evaluation value of the speech recognition of the extracted grammar based on the result of speech recognition using the extracted grammar and the correct data of the evaluation voice data; and a grammar selection unit that selects a grammar that satisfies the target from one or more extracted grammars extracted by the grammar extraction unit. The device accepts the execution time for grammar adjustment and repeats the extraction of grammar by the grammar extraction unit, speech recognition using the extracted grammar by the speech recognition unit, and calculation of the evaluation value of the speech recognition using the extracted grammar by the evaluation value calculation unit until the execution time is reached. Furthermore, a grammar adjustment device in one aspect of the present disclosure includes: a grammar storage unit that stores the grammar of voice commands for operating industrial equipment; a grammar extraction unit that extracts a representative of the clustered grammar, which is a part of the grammar that clusters the grammar stored in the grammar storage unit; a target registration unit that accepts registration of a target for the evaluation value of the speech recognition of the extracted grammar; a speech recognition unit that performs speech recognition of evaluation voice data using the extracted grammar; an evaluation value calculation unit that calculates an evaluation value for the speech recognition of the extracted grammar based on the result of speech recognition using the extracted grammar and the correct data of the evaluation voice data; and a grammar selection unit that selects a grammar that satisfies the target from among one or more extracted grammars extracted by the grammar extraction unit. A storage medium in one aspect of this disclosure stores a grammar of voice commands for operating industrial equipment, and, by execution by one or more processors, extracts a portion of the grammar, accepts registration of a target evaluation value for the speech recognition of the extracted grammar, performs speech recognition of evaluation voice data using the extracted grammar, calculates an evaluation value for the speech recognition of the extracted grammar based on the result of speech recognition using the extracted grammar and the correct data of the evaluation voice data, and selects a grammar that satisfies the target from among one or more extracted grammars, stores a processor-readable instruction, accepts the execution time for grammar adjustment, and continues until the execution time is reached. The aforementioned grammar The process of extracting the grammar, performing speech recognition using the extracted grammar, and calculating an evaluation value for the speech recognition using the extracted grammar is repeated. Furthermore, a storage medium in one aspect of the present disclosure stores a grammar of voice commands for operating industrial equipment, and stores a processor-readable instruction that, when executed by one or more processors, includes a part of the grammar, clusters the stored grammar, extracts a representative of the clustered grammar, accepts registration of a target for the speech recognition evaluation value of the extracted grammar, performs speech recognition of evaluation voice data using the extracted grammar, calculates a speech recognition evaluation value of the extracted grammar based on the result of speech recognition using the extracted grammar and the correct data of the evaluation voice data, and selects a grammar that satisfies the target from among one or more extracted grammars. [Effects of the Invention]

[0009] According to one aspect of the present invention, it is possible to support the creation of grammars for speech recognition. [Brief explanation of the drawing]

[0010] [Figure 1] This is a block diagram showing the configuration of a grammar adjustment device. [Figure 2] This figure shows examples of syntax definitions and word definitions. [Figure 3] This figure shows examples of speaker and recording location combinations for evaluation data. [Figure 4] This figure shows an example of the calculation results for the evaluation value. [Figure 5] This is a diagram showing an example of a goal registration screen. [Figure 6] This figure shows examples of correct answer rates for different grammars. [Figure 7] This is a flowchart explaining the processing of the grammar adjustment device. [Figure 8] This is the hardware configuration of the grammar adjustment device. [Modes for carrying out the invention]

[0011] The grammar adjustment device 100 will be described below. The grammar adjustment device 100 is implemented in an information processing device equipped with an arithmetic unit and a memory unit. Such an information processing device may include, but is not limited to, a PC (personal computer) or a mobile terminal.

[0012] Figure 1 shows the basic configuration of the grammar adjustment device 100. The grammar adjustment device 100 consists of an evaluation data storage unit 11, a target registration unit 12, a base grammar storage unit 13, a grammar extraction unit 14, a speech recognition unit 15, an extracted grammar storage unit 16, an evaluation value calculation unit 17, and a grammar selection unit 18.

[0013] The base grammar memory unit 13 stores the grammar of the base voice commands. Voice commands are commands used to operate industrial equipment by voice. The grammar of voice commands consists of syntax and words. The base grammar memory unit 13 includes a syntax memory unit 19 that stores syntax and a word memory unit 20 that stores words. A word includes the words that make up a voice command and the phoneme representation of those words. Syntax defines the arrangement of the words that make up a voice command. The base grammar is comprehensively created to cover as many voice commands as possible that are expected to be used on-site. For example, the syntax of the voice command to set the "override" of a numerical control device to "30" assumes "override 30", "set override to 30", "make override 30" and so on. The grammar creator constructs as many grammars as possible. The base grammar is determined by the type, specifications, work content, etc. of the device that recognizes voice commands. There may be cases where multiple phoneme sequences are assigned to one word. For example, the word "override" can be expressed by multiple phonemes such as "o:ba:raido", "oubaaraido", "oubaraido". The base grammar is created to cover as many such word phonemes as possible.

[0014] Grammar is composed of syntax and words. Figure 2 shows examples of syntax definitions and word definitions. In the example of the syntax definition, the words that make up the voice command and the order of the words are defined. In the first line of the syntax definition in Figure 2, "S:NS_B COMMAND NS_E", "S" is the start symbol of the voice command, and "NS_B" and "NS_E" are the silent intervals at the beginning and end of the sentence. The element "COMMAND" of the syntax exists between the silent intervals. The second and third lines define the "tags" that enter "COMMAND". The second line defines that the tags "ROBOT" and "INTERFACE" enter the syntax element "COMMAND", and the third line defines that the tags "NAIGAI" and "INTERFACE" enter the syntax element "COMMAND".

[0015] The first and second lines of the word definition define the Japanese notation and the phonetic notation of the tag "ROBOT". The Japanese notation of the tag "ROBOT" is "ロボット", and the phonetic notation is "roboqto". The third to fifth lines of the word definition define the Japanese notation and the phonetic notation of the Japanese words included in the tag "NAIGAI". Two Japanese words, "外部" and "内部", are included in the tag "NAIGAI". The phonetic notation of "外部" is "gaibu", and the phonetic notation of "内部" is "naibu". The sixth to eighth lines of the word definition define the Japanese notation and the phonetic notation of the Japanese words included in the tag "INTERFACE". One Japanese word, "インターフェース", is included in the tag "INTERFACE". The tag "インターフェース" has two types of phonetic notations, "iNtafe:su" and "iNta:feisu". "%NS_B" defines the silent interval [s] at the beginning of the sentence, and "%NS_E" defines the silent interval [ / s] at the end of the sentence.

[0016] The grammar extraction unit 14 extracts some grammars from the comprehensive grammars stored in the basic grammar storage unit 13. As an example of the grammar extraction method, cluster division of the k-means method is used. Other methods than the k-means method may be used for grammar extraction. In the cluster division of the k-means method, the acoustic distance of the grammar is used. There are methods such as calculating the acoustic distance from the acoustic spectrum and calculating the acoustic distance from the phoneme string for the acoustic distance. In the method of calculating the acoustic distance from the acoustic spectrum, the acoustic spectrum of the voice command is vectorized, and the cosine distance or Euclidean distance between vectors is calculated. In the method of calculating the acoustic distance from the phoneme string, the cosine distance, Levenshtein distance, Jaro-Winkler distance, and Hamming distance are used. The cosine distance, Euclidean distance, Levenshtein distance, Jaro-Winkler distance, and Hamming distance are well-known. In addition to calculating the distance of the entire voice command, it is possible to calculate the acoustic distance of the phonemes of the words included in the voice command (for example, "iNtafe:su", "iNtafeisu", etc.) and extract a part from the cluster of words with a close acoustic distance.

[0017] In the k-means method presented as an example in this disclosure, random numbers are used to set the centers of K clusters, (a) the center closest to each voice command (or word) is assigned, and (b) the center is calculated for each cluster. Steps (a) and (b) are repeated until the centers of all clusters no longer change, thereby dividing the voice command into clusters. The grammar extraction unit 14 extracts the grammar (syntax and words) of voice commands included in the same cluster and outputs it to the speech recognition unit as a grammar for evaluation. Note that the k-means method is just one example of a method for extracting grammars that are close together, and other methods may be used. Also, the results of the k-means method are affected by the initial random number and the number of clusters K. The random number and the number of clusters K may be set manually by the user or automatically by the grammar extraction unit 14.

[0018] The evaluation data storage unit 11 stores audio data containing voice commands recorded by multiple speakers at multiple recording locations, associating it with correct answer data, which is the correct text corresponding to the audio data. For example, it stores audio data in which multiple speakers utter "external interface" at multiple recording locations, associating it with the correct answer data (text) "external interface". The voice data in the evaluation data storage unit 11 is recorded by speakers with different attributes (gender, age) at different recording locations. Since the evaluation data is recorded at the site where voice commands are used, it includes background noise from the site where the voice commands are used. Figure 3 is a table showing the relationship between the speaker and the recording location of the evaluation data. The evaluation data in Figure 3 includes voice recordings by speaker A (male, 60 years old) at factories A and B, and voice recordings by speaker B (female, 30 years old) at factories C and D, etc.

[0019] The speech recognition unit 15 receives voice commands from the evaluation data storage unit 11 and performs speech recognition of the input voice commands. The speech recognition unit 15 generally consists of an acoustic model, a language model, and a decoder. The acoustic model receives voice data and outputs phonemes (senons) that make up the voice data based on the features of the voice data. The language model outputs the occurrence probability of word sequences. The language model selects hypothetical word sequences based on the phonemes and outputs linguistically plausible candidates. The decoder outputs word sequences with high probability as recognition results based on the outputs of the statistically created acoustic model and language model.

[0020] The extracted grammar memory unit 16 comprises a syntactic memory unit 21 and a word memory unit 22, and stores the grammar extracted by the grammar extraction unit 14. The speech recognition unit 15 performs speech recognition using the grammar stored in the extracted grammar memory unit 16.

[0021] The evaluation value calculation unit 17 compares the correct text from the evaluation data storage unit with the recognition result of the speech recognition unit 15 and calculates the accuracy rate of speech recognition. Figure 4 shows an example of the accuracy rate as an evaluation value. In this disclosure, not only is the overall evaluation of the speech command calculated, but the accuracy rate for each type of speech command is also calculated. Examples of speech command types include approval commands, numerical commands, and transition commands. Approval commands are commands that indicate approval. Examples of approval commands include "yes," "no," "yes," "no," "execute," and "cancel." Numerical commands are commands that specify numerical values ​​such as "0.5," "1," "2," and "100." "Transition commands" are commands that specify display screens such as "home screen" and "speed setting screen." In addition, "machine operation commands" that instruct the movement of equipment, such as "set the workpiece," are also conceivable.

[0022] The target registration unit 12 accepts registration of target values ​​for speech recognition. The target registration unit 12 accepts target values ​​such as the target accuracy rate for the entire speech command, the target accuracy rate for each type of speech command, and the target time for the search. Figure 5 shows an example of a goal registration screen. In Figure 5, the target accuracy rates for each type of voice command are set as follows: "Approval command: 95% or higher", "Numerical command: 90% or higher", and "Transition command: 80% or higher", with a maximum execution time of "within 30 minutes".

[0023] The grammar selection unit 18 compares the speech recognition results with the target accuracy rate, and if it determines that any grammar meets the target accuracy rate based on the speech recognition results, it selects that grammar as the appropriate grammar. The processing of the grammar selection unit 18 is repeated until the target time for grammar adjustment has elapsed or the target accuracy rate has been achieved. If the target time for grammar adjustment has elapsed, it selects an appropriate grammar from the grammars for which speech recognition has been performed so far. The grammar selection unit 18 may also present the accuracy rate of each grammar to the grammar creator, who may then select a grammar.

[0024] An example of a grammar selection method is described below. In this disclosure, the accuracy rate of voice commands was calculated for each type: approval commands, transition commands, and numerical commands. Of these voice commands, approval commands require a high accuracy rate because they are used to confirm operations. Numerical commands, which specify numerical values, also require a high accuracy rate. Transition commands, which instruct screen transitions, may have a lower accuracy rate compared to approval commands and numerical commands. The grammar adjustment device in this disclosure allows setting a target accuracy rate for each voice command or for each type of voice command. The grammar that achieves the target accuracy rate is automatically selected. For example, Figure 6 shows the accuracy rates for grammar A and grammar B. The accuracy rates for "approval commands," "numerical commands," and "transition commands" in grammar A satisfy the target accuracy rate registered in the target registration unit 12, so grammar A is selected as the appropriate grammar. If multiple grammars satisfy the target accuracy rate, further conditions may be set.

[0025] The processing of the grammar adjustment device 100 of this disclosure will be described with reference to Figure 7. As a preparation step, the grammar adjustment device 100 accepts registration of the target accuracy rate of voice commands (step S1), registration of the maximum execution time for grammar adjustment (step S2), and registration of cluster partitioning criteria (step S3). When using the k-means method, it accepts registration of the initial random number and the number of clusters K as the cluster partitioning criteria.

[0026] The grammar adjustment device 100 clusters the voice commands (or words included in the voice commands) stored in the base grammar memory unit 13 (step S4), extracts one or more representative voice commands (or words included in the voice commands) from each cluster (step S5), and reconstructs the grammar using the extracted voice commands (or words included in the voice commands) (step S6). The grammar adjustment device 100 performs speech recognition of the evaluation data using the grammar reconstructed in step S5 (step S7). The grammar adjustment device 100 calculates an evaluation value for speech recognition (step S8). The grammar adjustment device 100 compares the evaluation result with the target accuracy rate, and if the evaluation result meets the conditions for the target accuracy rate (step S9; Yes), it selects the grammar (step S10).

[0027] In step S9, if the evaluation result does not meet the target accuracy criteria (step S9; No), it is determined whether the maximum execution time has been reached. If the target time for grammar adjustment has been reached (step S11; Yes), the grammar adjustment device 100 presents the grammars that have been recognized so far to the user and accepts the selection of a grammar (step S10).

[0028] If the target time for grammar adjustment has not been reached in step S9 (step S11; No), the process moves to step S4, and the process from step S4 to step S9 is repeated. In this flowchart, grammar adjustments are terminated once the target accuracy rate is met, but adjustments may be continued until the maximum execution time is reached.

[0029] As described above, the grammar adjustment device 100 of this disclosure is a device that assists in creating grammars for voice commands, and extracts a portion of the comprehensively created grammar, reconstructs the grammar, and selects the grammar with a high accuracy rate.

[0030] Since the grammatical accuracy is calculated for each type of voice command, the grammar can be adjusted to achieve an accuracy suitable for the specific situation in which voice recognition is used.

[0031] Since the data used for grammar evaluation is recorded at the site where voice commands are used, it is possible to construct a grammar suitable for recognizing voice data that includes noise specific to the site or time of day. Furthermore, because the grammar adjustment device of this disclosure registers site-specific technical terms and expressions as grammar, the accuracy rate is improved because it selects recognition candidates from the words and sentence structures registered in the grammar even if noise is present.

[0032] The grammar adjustment device 100 of this disclosure automatically adjusts grammar, and therefore can optimize grammar based on objective criteria, without relying on the subjective opinion or know-how of the grammar creator. Furthermore, because it automatically adjusts grammar, even engineers with little experience can adjust the grammar.

[0033] [Hardware configuration] Referring to Figure 8, the hardware configuration of the grammar adjustment device 100 will be described. The CPU 111 of the grammar adjustment device 100 is a processor that controls the grammar adjustment device 100 as a whole. The CPU 111 reads the system program processed in the ROM 112 via the bus and controls the entire grammar adjustment device 100 according to the system program. The RAM 113 temporarily stores temporary calculation data, display data, and various data entered by the user via the input unit 71.

[0034] The display unit 70 is a monitor or similar device attached to the grammar adjustment device 100. The display unit 70 displays the operation screen, settings screen, etc., of the grammar adjustment device 100.

[0035] The input unit 71 is either integrated with the display unit 70 or separate from it, such as a keyboard, touch panel, or operation buttons. The user operates the input unit 71 to input information onto the screen displayed on the display unit 70. The display unit 70 and the input unit 71 may also be portable terminals.

[0036] The non-volatile memory 114 is a memory that retains its stored state even when the power to the grammar adjustment device 100 is turned off, for example, by being backed up by a battery (not shown). The non-volatile memory 114 stores processing programs, system programs, available options, billing tables, etc. The non-volatile memory 114 stores programs read from external devices via an interface (not shown), programs input via the input unit 71, and various data acquired from various parts of the grammar adjustment device 100 and machine tools (for example, setting parameters acquired from machine tools). The programs and various data stored in the non-volatile memory 114 may be expanded into the RAM 113 during execution / use. In addition, various system programs are pre-written to the ROM 112. [Explanation of Symbols]

[0037] 100 Grammar adjuster 11 Evaluation data storage unit 12. Goal Registration Section 13. Basic Grammar Memory Unit 14 Grammar extraction part 15. Voice Recognition Unit 16 Extracted grammar storage 17. Evaluation Value Calculation Unit 18. Grammar Selection Section 19 Syntax Memory Unit 20 Word Memory Section 21 Syntax Memory Unit 22 Word Memory Section 70 Display section 71 Input section 111 CPU 112 ROM 113 RAM 114 Non-volatile memory

Claims

1. A grammar memory unit that stores the grammar of voice commands for operating industrial equipment, A grammar extraction unit that extracts a part of the aforementioned grammar, A target registration unit that accepts registration of target values ​​for speech recognition of the extracted grammar, A speech recognition unit performs speech recognition on the evaluation speech data using the extracted grammar, An evaluation value calculation unit calculates an evaluation value for speech recognition of the extracted grammar based on the results of speech recognition using the extracted grammar and the correct data of the evaluation speech data. A grammar selection unit selects a grammar that satisfies the objective from among one or more grammars extracted by the grammar extraction unit, A grammar adjustment device comprising, The system accepts the execution time for grammar adjustment and, until the execution time is reached, repeats the following steps: grammar extraction by the grammar extraction unit, speech recognition using the extracted grammar by the speech recognition unit, and calculation of an evaluation value for speech recognition using the extracted grammar by the evaluation value calculation unit. Grammar adjuster.

2. A grammar memory unit that stores the grammar of voice commands for operating industrial equipment, A grammar extraction unit that clusters a part of the grammar stored in the grammar memory unit and extracts a representative of the clustered grammar, A target registration unit that accepts registration of target values ​​for speech recognition of the extracted grammar, A speech recognition unit performs speech recognition on the evaluation speech data using the extracted grammar, An evaluation value calculation unit calculates an evaluation value for speech recognition of the extracted grammar based on the results of speech recognition using the extracted grammar and the correct data of the evaluation speech data. A grammar selection unit selects a grammar that satisfies the objective from among one or more grammars extracted by the grammar extraction unit, A grammar adjustment device equipped with the following features.

3. The grammar adjustment device according to claim 2, wherein the grammar extraction unit clusters the grammar using the acoustic distance of the voice commands defined in the grammar.

4. The grammar adjustment device according to claim 3, wherein the grammar extraction unit clusters the grammar using the acoustic distance of words included in the voice commands defined by the grammar.

5. The grammar adjustment device according to claim 1 or 2, wherein the evaluation value is the accuracy rate of the speech recognition.

6. Memorize the grammar of voice commands for operating industrial equipment. One or more processors execute, Extracting a portion of the aforementioned grammar, The system accepts registration of target values ​​for the speech recognition evaluation of the extracted grammar. Using the extracted grammar, speech recognition is performed on the evaluation audio data. Based on the speech recognition results using the extracted grammar and the correct data of the evaluation speech data, the evaluation value of the extracted grammar is calculated. From among the one or more extracted grammars, select the grammar that satisfies the objective. A storage medium for storing instructions that the processor can read, The system accepts the execution time for grammar adjustment and repeats the following steps until the execution time is reached: extraction of the grammar, speech recognition using the extracted grammar, and calculation of an evaluation value for the speech recognition using the extracted grammar. A storage medium for storing instructions that the processor can read.

7. Memorize the grammar of voice commands for operating industrial equipment. One or more processors execute, A part of the aforementioned grammar, the stored grammar is clustered, and a representative of the clustered grammar is extracted. The system accepts registration of target values ​​for the speech recognition evaluation of the extracted grammar. Using the extracted grammar, speech recognition is performed on the evaluation audio data. Based on the speech recognition results using the extracted grammar and the correct data of the evaluation speech data, the evaluation value of the extracted grammar is calculated. From among the one or more extracted grammars, select the grammar that satisfies the objective. A storage medium for storing instructions that the processor can read.