Voice recognition device

The speech recognition device simplifies the model selection process and enhances accuracy by using a provisional parameter selection mechanism based on narrowed-down information, optimizing the choice of acoustic and grammar models.

JP7750950B2Active Publication Date: 2025-10-07FANUC LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
JP2023529280
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-06-22
Publication Date
2025-10-07
Estimated Expiration
2041-06-22

AI Technical Summary

Technical Problem

Selecting an appropriate speech recognition model in advance places a heavy burden on the user.

Method used

A speech recognition device that includes a reception unit, parameter storage unit, provisional setting parameter selection unit, recognition unit, and parameter selection unit to automatically select an appropriate speech recognition model based on narrowed-down information, using acoustic and grammar setting parameters.

Benefits of technology

Simplifies the task of selecting a speech recognition model and improves recognition accuracy by automatically choosing the most reliable model for a given context.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007750950000001
    Figure 0007750950000001
  • Figure 0007750950000002
    Figure 0007750950000002
  • Figure 0007750950000003
    Figure 0007750950000003
Patent Text Reader

Abstract

This speech recognition device is provided with: an acceptance unit that accepts input of speech information; a parameter storage unit that stores a plurality of parameters for setting a speech recognition model; a temporary setting parameter selection unit that selects, on the basis of narrowing information, a temporary setting parameter to be temporarily set, from the plurality of parameters; a recognition unit that recognizes speech information on the basis of the selected temporary setting parameter; and a parameter selection unit that selects, on the basis of information indicating a recognition result of the recognized speech information, one of the temporary setting parameters.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to a speech recognition device. [Background technology]

[0002] In recent years, there have been attempts to utilize voice recognition technology in the field of industrial machinery (for example, Patent Document 1). In order to improve the recognition accuracy of voice recognition, it is necessary to select an appropriate voice recognition model in advance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2020-160586 Summary of the Invention [Problem to be solved by the invention]

[0004] However, selecting an appropriate speech recognition model in advance places a heavy burden on the user.

[0005] An object of the present disclosure is to provide a speech recognition device that can simplify the task of selecting a speech recognition model. [Means for solving the problem]

[0006] The speech recognition device includes a reception unit that receives input of speech information, a parameter storage unit that stores a plurality of parameters for setting a speech recognition model, and a parameter storage unit that stores a plurality of parameters for setting a speech recognition model based on the narrowed-down information. The aforementioned a temporary setting parameter selection unit that selects a temporary setting parameter that is temporarily set from a plurality of parameters; The aforementioned Provisional parameters The speech recognition model stored in association with A recognition unit that recognizes voice information and The aforementioned Voice recognition results Information including the reliability of Based on The aforementioned a parameter selection unit that selects one of the provisionally set parameters; The plurality of parameters include at least one of a plurality of acoustic setting parameters for setting an acoustic model and a plurality of grammar setting parameters for setting a grammar model.. [Effects of the Invention]

[0007] One aspect of the present disclosure makes it possible to simplify the task of selecting a speech recognition model. [Brief explanation of the drawings]

[0008] [Figure 1] FIG. 2 is a block diagram showing an example of a hardware configuration of a numerical control device. [Figure 2] FIG. 2 is a block diagram showing an example of functions of a voice recognition device. [Figure 3] FIG. 2 is a block diagram showing an example of the function of a recognition unit. [Figure 4] FIG. 10 is a diagram illustrating an example of an image displayed on a display screen of the input / output device. [Figure 5] 10 is a flowchart illustrating an example of processing performed in a preparation stage. [Figure 6] 10 is a flowchart illustrating an example of processing performed at a parameter setting stage. [Figure 7] 10 is a flowchart illustrating an example of processing performed after parameter setting. [Figure 8] FIG. 10 is a block diagram showing an example of functions of a narrowed-down information acquisition unit. [Figure 9] FIG. 2 is a block diagram showing an example of functions of a voice recognition device. [Figure 10] FIG. 2 is a block diagram showing an example of functions of a voice recognition device. [Figure 11] FIG. 10 is a diagram illustrating an example of an image displayed on a display screen of the input / output device. DETAILED DESCRIPTION OF THE INVENTION

[0009] An embodiment of the present disclosure will be described below with reference to the drawings. Note that not all combinations of features described in the following embodiment are necessarily required to solve the problem. In addition, more detailed description than necessary may be omitted. Furthermore, the following description of the embodiment and the drawings are provided to enable those skilled in the art to fully understand the present disclosure, and are not intended to limit the scope of the claims.

[0010] A speech recognition device is a device that recognizes speech uttered by a speaker. The speech recognized by the speech recognition device is converted into a command that instructs the operation of, for example, industrial machinery. The industrial machinery operates based on this command.

[0011] The voice recognition device is implemented, for example, in a numerical control device that controls industrial machinery. The voice recognition device may be implemented in a server connected to the numerical control device via a LAN (Local Area Network). Furthermore, if security is ensured, the voice recognition device may be implemented in a server connected to the numerical control device via the Internet. An example in which the voice recognition device is implemented in a numerical control device of industrial machinery will be described below.

[0012] 1 is a block diagram showing an example of the hardware configuration of an industrial machine. The industrial machine 1 is, for example, a machine tool, a wire electric discharge machine, or an industrial robot. The machine tool includes a lathe, a machining center, and a multi-tasking machine. The industrial robot includes a manipulator.

[0013] The industrial machine 1 includes a numerical control device 2, an input / output device 3, a servo amplifier 4 and a servo motor 5, a spindle amplifier 6 and a spindle motor 7, an auxiliary device 8, and a microphone 9.

[0014] The numerical control device 2 is a device that controls the entire industrial machine 1. The numerical control device 2 includes a hardware processor 201, a bus 202, a ROM (Read Only Memory) 203, a RAM (Random Access Memory) 204, and a non-volatile memory 205.

[0015] The hardware processor 201 is a processor that controls the entire numerical control device 2 in accordance with a system program. The hardware processor 201 reads the system program stored in the ROM 203 via the bus 202 and performs various processes based on the system program. The hardware processor 201 also controls the servo motor 5 and the spindle motor 7 based on the machining program. The hardware processor 201 is, for example, a CPU (Central Processing Unit) or an electronic circuit.

[0016] The hardware processor 201, for example, analyzes the machining program and outputs control commands to the servo motor 5 and the spindle motor 7 for each control period.

[0017] The bus 202 is a communication path that connects the various pieces of hardware within the numerical control device 2. The various pieces of hardware within the numerical control device 2 exchange data via the bus 202.

[0018] The ROM 203 is a storage device that stores a system program for controlling the entire numerical control device 2. The ROM 203 is a computer-readable storage medium.

[0019] The RAM 204 is a storage device that temporarily stores various data and functions as a work area for the hardware processor 201 to process various data.

[0020] The nonvolatile memory 205 is a storage device that retains data even when the industrial machine 1 is turned off and no power is supplied to the numerical control device 2. The nonvolatile memory 205 stores, for example, machining programs and various parameters. The nonvolatile memory 205 is a computer-readable storage medium. The nonvolatile memory 205 is, for example, configured as an SSD (Solid State Drive).

[0021] The numerical control device 2 further includes a first interface 206 , an axis control circuit 207 , a spindle control circuit 208 , a PLC (Programmable Logic Controller) 209 , an I / O unit 210 , and a second interface 211 .

[0022] The first interface 206 connects the bus 202 and the input / output device 3. The first interface 206 sends various data processed by the hardware processor 201 to the input / output device 3, for example.

[0023] The input / output device 3 is a device that receives various data via the first interface 206 and displays the various data. The input / output device 3 also accepts input of various data and sends the various data to the hardware processor 201 via the first interface 206. The input / output device 3 is, for example, a touch panel. When the input / output device 3 is a touch panel, the touch panel is, for example, a capacitive touch panel. Note that the touch panel is not limited to a capacitive touch panel and may be a touch panel of another type. The input / output device 3 is attached, for example, to an operation panel (not shown) in which the numerical control device 2 is housed.

[0024] The axis control circuit 207 is a circuit that controls the servo motor 5. The axis control circuit 207 receives a control command from the hardware processor 201 and outputs a command to drive the servo motor 5 to the servo amplifier 4. The axis control circuit 207 sends, for example, a torque command to control the torque of the servo motor 5 to the servo amplifier 4.

[0025] The servo amplifier 4 receives a command from the axis control circuit 207 and supplies a current to the servo motor 5 .

[0026] The servo motor 5 is driven by receiving a current supply from the servo amplifier 4. The servo motor 5 is connected to, for example, a ball screw that drives a tool post. When the servo motor 5 is driven, a structure of the industrial machine 1, such as the tool post, moves, for example, in the X-axis direction, Y-axis direction, or Z-axis direction. The servo motor 5 may also have a built-in speed detector (not shown) that detects the feed speed of each feed axis.

[0027] The spindle control circuit 208 is a circuit for controlling the spindle motor 7. The spindle control circuit 208 receives a control command from the hardware processor 201 and outputs a command to the spindle amplifier 6 for driving the spindle motor 7. The spindle control circuit 208 sends, for example, a torque command for controlling the torque of the spindle motor 7 to the spindle amplifier 6.

[0028] The spindle amplifier 6 receives a command from the spindle control circuit 208 and supplies a current to the spindle motor 7 .

[0029] The spindle motor 7 is driven by receiving a current supplied from the spindle amplifier 6. The spindle motor 7 is connected to the main shaft and rotates the main shaft.

[0030] The PLC 209 is a device that executes a ladder program to control the auxiliary device 8. The PLC 209 sends commands to the auxiliary device 8 via an I / O unit 210.

[0031] The I / O unit 210 is an interface that connects the PLC 209 and the auxiliary device 8. The I / O unit 210 sends a command received from the PLC 209 to the auxiliary device 8.

[0032] The auxiliary device 8 is a device that is installed in the industrial machine 1 and performs auxiliary operations in the industrial machine 1. The auxiliary device 8 operates based on commands received from the I / O unit 210. The auxiliary device 8 may be a device that is installed in the periphery of the industrial machine 1. The auxiliary device 8 is, for example, a tool changer, a cutting fluid injection device, or an opening / closing door drive device.

[0033] The second interface 211 connects the bus 202 and the microphone 9. The second interface 211 sends, for example, audio information output from the microphone 9 to the hardware processor 201.

[0034] The microphone 9 is an acoustic device that captures sound and converts it into audio information. Here, the audio information is an electrical signal. The microphone 9 sends the audio information to the hardware processor 201 via the second interface 211.

[0035] Next, an overview of the voice recognition device 20 will be described.

[0036] FIG. 2 is a block diagram showing an example of the functions of the voice recognition device 20 implemented in the numerical control device 2. As shown in FIG.

[0037] The speech recognition device 20 includes a receiving unit 21, a parameter storage unit 22, a narrowed-down information acquisition unit 23, a provisional setting parameter selection unit 24, a recognition unit 25, a parameter selection unit 26, an output unit 27, and a setting parameter storage unit 28.

[0038] The reception unit 21, the narrowed-down information acquisition unit 23, the provisional setting parameter selection unit 24, the recognition unit 25, the parameter selection unit 26, and the output unit 27 are realized, for example, by the hardware processor 201 performing arithmetic processing using the system program stored in the ROM 203 and various data stored in the non-volatile memory 205.

[0039] The parameter storage unit 22 and the set parameter storage unit 28 are realized by storing data input from the input / output device 3 and various parameters in the RAM 204 or the nonvolatile memory 205, for example.

[0040] The speech recognition device 20 selects an appropriate speech recognition model from a plurality of pre-stored speech recognition models and performs speech recognition. Parameters are selected so that the speech recognition device 20 can select a speech recognition model. By selecting parameters for setting an appropriate speech recognition model for the speech recognition device 20 and using the appropriate speech recognition model, the speech recognition device 20 can recognize speech information with high accuracy.

[0041] In order to set an appropriate speech recognition model, first, the provisional parameter selection unit 24 selects parameters to be provisionally set from the parameters stored in the parameter storage unit 22. At this time, the provisional parameter selection unit 24 selects the parameters to be provisionally set from the parameters narrowed down based on the narrowing conditions.

[0042] The recognition unit 25 recognizes speech information using a speech recognition model stored in association with the parameters selected by the provisional parameter selection unit 24. The recognition unit 25 recognizes speech information using, for example, a plurality of speech recognition models, and derives recognition results for the speech information for each of the speech recognition models. The parameter storage unit 28 stores, for example, provisional parameters that have derived the most reliable recognition result from among the plurality of recognition results.

[0043] In this way, the provisionally set parameters that have led to a highly reliable recognition result can be set as parameters to be used in subsequent speech recognition, thereby improving the accuracy of speech recognition. Next, each unit of the speech recognition device 20 will be described in detail.

[0044] The reception unit 21 receives input of voice information transmitted from the microphone 9. The voice information is, for example, an analog signal representing the voice spoken by a speaker. The voice information may also be a digital signal converted from the analog signal representing the voice. The voice of the speaker is acquired, for example, by a microphone 9 installed in the industrial machine 1 or a microphone 9 placed at a predetermined position in a factory.

[0045] The parameter storage unit 22 stores a plurality of parameters for setting a speech recognition model. The speech recognition model is, for example, an acoustic model and a grammar model. That is, the parameter storage unit 22 stores a plurality of acoustic setting parameters for setting the acoustic model and a plurality of grammar setting parameters for setting the grammar model. Speech recognition models such as acoustic models and grammar models will be described in detail later.

[0046] The audio setting parameters include, for example, Japanese setting parameters, English setting parameters, Chinese setting parameters, and German setting parameters, and the grammar setting parameters include, for example, network setting parameters, tool setting parameters, power setting parameters for general users, and power setting parameters for administrators.

[0047] The narrowing-down information acquisition unit 23 acquires narrowing-down information for narrowing down the multiple parameters stored in the parameter storage unit 22.

[0048] The narrowed-down information includes, for example, speaker information that identifies the speaker who uttered the voice. The speaker information includes, for example, language information that indicates the language spoken by the speaker and job information that indicates the job performed by the speaker.

[0049] The language information is information indicating at least one of Japanese, English, Chinese, and German, for example, and the job information is information indicating at least one of network settings, machining, user power settings, and administrator power settings, for example.

[0050] When the reception unit 21 receives voice information, the narrowed-down information acquisition unit 23 first analyzes the voice information to, for example, identify the speaker. Next, the narrowed-down information acquisition unit 23 acquires language information and job information stored in association with the identified speaker. The narrowed-down information acquisition unit 23 identifies the speaker, for example, by comparing the voice information received by the reception unit 21 with narrowed-down criteria information registered in advance. The narrowed-down criteria information is, for example, voice information indicating the voice uttered by the speaker.

[0051] The narrowed-down information may include time information that specifies the reception time at which the voice information was received by the reception unit 21. The narrowed-down information may include location information that specifies the location at which the microphone 9 that captures the voice is installed.

[0052] The provisional parameter selection unit 24 selects provisional parameters to be provisionally set from the plurality of parameters based on the narrowing down information. That is, the provisional parameter selection unit 24 narrows down the parameters to be provisionally set based on the narrowing down information.

[0053] The provisional parameter selection unit 24 may select one parameter from the plurality of parameters as the provisional parameter. The provisional parameter selection unit 24 may select a set of parameters including a plurality of parameters from the plurality of parameters. The provisional parameter selection unit 24 may select a plurality of sets of parameters.

[0054] For example, if the narrowed-down information acquisition unit 23 acquires language information indicating Japanese as speaker information and job information indicating network settings, the provisional setting parameter selection unit 24 selects Japanese setting parameters and network setting parameters as parameters to be provisionally set.

[0055] Furthermore, when the narrowed-down information acquisition unit 23 acquires only language information indicating Japanese as the speaker information, the provisional parameter selection unit 24 selects Japanese setting parameters as the parameters to be provisionally set. In this case, the provisional parameter selection unit 24 may select, for example, two pairs of Japanese setting parameters and network setting parameters, and two pairs of Japanese setting parameters and tool setting parameters as the provisional setting parameters.

[0056] The recognition unit 25 recognizes the voice information based on the provisional parameters selected by the provisional parameter selection unit 24. That is, the recognition unit 25 recognizes the voice information using a voice recognition model stored in association with the provisional parameters. When the provisional parameter selection unit 24 selects multiple sets of provisional parameters, the recognition unit 25 recognizes the voice information using each set of provisional parameters.

[0057] 3 is a block diagram showing an example of the functions of the recognition unit 25. The recognition unit 25 includes a model storage unit 251, a dictionary storage unit 252, and a recognition processing unit 253.

[0058] The model storage unit 251 stores a plurality of speech recognition models respectively corresponding to a plurality of parameters stored in the parameter storage unit 22. As described above, the speech recognition models include, for example, an acoustic model and a grammar model.

[0059] The acoustic model is used to distinguish phonemes included in the speech information. For example, the acoustic model includes a Japanese model, an English model, a Chinese model, and a German model. The Japanese model, the English model, the Chinese model, and the German model are used by the recognition unit 25 when Japanese setting parameters, English setting parameters, Chinese setting parameters, and German setting parameters are set, respectively. The acoustic model is generated, for example, by machine learning using speech information of speech uttered by speakers of each language as training data.

[0060] The grammar model is used to determine phoneme sequences, i.e., strings and word sequences that match phoneme patterns, and to derive appropriate strings and word sequences for the language. The grammar model includes, for example, a network configuration model, a tool configuration model, a general user power configuration model, and an administrator power configuration model.

[0061] The network setting model is a grammar model for determining with high accuracy the words, character strings, commands, etc. used when setting up a network. The tool setting model is a grammar model for determining with high accuracy the words, character strings, commands, etc. used when setting up a tool. The power setting model for general users is a grammar model for determining with high accuracy the words, character strings, commands, etc. used when general users configure power settings. The power setting model for administrators is a grammar model for determining with high accuracy the words, character strings, commands, etc. used when administrators configure power settings.

[0062] The network setting model, tool setting model, power setting model for general users, and power setting model for administrators each include a model compatible with Japanese, a model compatible with English, a model compatible with Chinese, and a model compatible with German.

[0063] Furthermore, each grammar model is generated by machine learning using text information in the language spoken by speakers of that language as training data. For example, the training data for a network configuration model is text information indicating the character strings used by network configuration workers when configuring the network, as well as word sequences. The grammar model formulates sentence structures such as the order of parts of speech, the relationships between words, and so on.

[0064] The network setting model, the tool setting model, the power setting model for general users, and the power setting model for administrators are used by the recognition unit 25 when the network setting parameters, the tool setting parameters, the power setting parameters for general users, and the power setting parameters for administrators are set, respectively.

[0065] The dictionary storage unit 252 stores dictionaries, including, for example, a Japanese dictionary, an English dictionary, a Chinese dictionary, and a German dictionary.

[0066] The recognition processing unit 253 executes speech recognition processing using a speech recognition model. First, the recognition processing unit 253 extracts features from speech information. For example, the recognition processing unit 253 extracts the strength and frequency characteristics of a speech from the speech information, which is an analog signal, as features.

[0067] The recognition processing unit 253 uses an acoustic model to determine the phonemes contained in the speech information from the extracted features. For example, if a Japanese model is selected by the provisional parameter selection unit 24, the recognition processing unit 253 uses the Japanese model to determine the Japanese phonemes and phoneme sequences contained in the speech information from the features.

[0068] The recognition processing unit 253 uses a grammar model to determine character strings and word sequences that match the phoneme patterns and derives appropriate character strings and word sequences for the language. In other words, the recognition processing unit 253 uses a grammar model to convert speech information into text so that the meaning of the language indicated by the speech information can be understood. The recognition processing unit 253 uses a dictionary stored in the dictionary storage unit 252 when determining character strings and word sequences that match the phoneme patterns.

[0069] For example, when the provisional setting parameter selection unit 24 selects Japanese setting parameters and network setting parameters, the recognition processing unit 253 converts the voice information into text using a network setting model compatible with Japanese.

[0070] When converting speech information into text, the recognition processing unit 253 stores multiple candidates for character strings and word sequences that match the phoneme pattern and searches for the optimal character string and word sequence. The maximum number of candidates is called the beam width, and the beam width is set to, for example, 1, 3, or 5.

[0071] The recognition processing unit 253 calculates information indicating the recognition result of the voice recognition. The information indicating the recognition result includes a reliability indicating the accuracy of the recognition of the voice information. When there are many candidates for the converted text character string and word sequence, the recognition processing unit 253 sets a low value for the reliability of each recognition result. On the other hand, when there are few candidates for the converted text character string and word sequence, the recognition processing unit 253 sets a high value for the reliability of each recognition result. The recognition processing unit 253 calculates the reliability as a value greater than or equal to 0 and less than or equal to 1. Note that the method for calculating the reliability is not limited to this, and other calculation methods may be used.

[0072] For example, when the recognition unit 25 recognizes the speech information of a voice saying "I want to check the IP address" spoken in a specified language such as Japanese, English, Chinese, or German, information indicating the recognition result is derived, for example, "I want to check the IP address (reliability 0.8)" and "PowerOff (reliability 0.1)".

[0073] The parameter selection unit 26 selects one of the provisional setting parameters based on information indicating the recognition result of the recognized speech information. The parameter selection unit 26 selects one of the provisional setting parameters based on, for example, reliability. Hereinafter, the provisional setting parameter selected by the parameter selection unit 26 will be referred to as a setting parameter.

[0074] The parameter selection unit 26 selects, for example, the setting parameters with the highest reliability. That is, it selects the set of setting parameters with the highest reliability from among the multiple sets of parameters selected by the temporary setting parameter selection unit 24. Alternatively, the parameter selection unit 26 may select setting parameters with a reliability equal to or higher than a predetermined threshold value.

[0075] The output unit 27 outputs the setting parameters selected by the parameter selection unit 26. In other words, when the setting parameter storage unit 28 stores the setting parameters, the output unit 27 outputs a message that the setting parameter storage unit 28 has stored the setting parameters.

[0076] The output unit 27 outputs information indicating that the setting parameter storage unit 28 has stored the temporary setting parameters to, for example, an indicator light, the input / output device 3, or a speaker installed in the industrial machine 1. When the setting parameter storage unit 28 stores the setting parameters, the output unit 27 may output information that the setting parameter storage unit 28 has stored the setting parameters.

[0077] Furthermore, the output unit 27 outputs the recognition result of the speech recognition derived from the setting parameters selected by the parameter selection unit 26.

[0078] 4 is a diagram showing an example of an image displayed on the display screen of the input / output device 3. When the setting parameters are selected by the parameter selection unit 26, for example, a pop-up screen showing the recognition result is displayed on the display screen.

[0079] The setting parameter storage unit 28 stores the setting parameters selected by the parameter selection unit 26. That is, the setting parameters selected by the parameter selection unit 26 are stored in the setting parameter storage unit 28, thereby setting the parameters.

[0080] Once the setting parameters are stored in the setting parameter storage unit 28, that is, once the parameters are set, the recognition unit 25 recognizes the voice information based on the setting parameters.

[0081] Next, we will explain an example of the flow of processing performed in the voice recognition device 20. The processing performed in the voice recognition device 20 includes processing performed in a preparation stage, processing performed in a parameter setting stage, and processing performed after parameter setting.

[0082] 5 is a flowchart showing an example of processing performed in the preparation stage. In the speech recognition device 20, parameters and a speech recognition model are registered (step SA1). That is, a plurality of parameters and a plurality of speech recognition models corresponding to the plurality of parameters are stored in the parameter storage unit 22 and the model storage unit 251, respectively. At this time, other processing such as dictionary registration may also be performed.

[0083] Next, a reliability threshold is registered (step SA2). That is, the reliability threshold, which is used as a reference when the parameter selection unit 26 selects a provisional parameter, is stored in a predetermined storage unit (not shown). Note that this process is not performed when it is not necessary to register a threshold, such as when the parameter selection unit 26 selects a provisional parameter that has derived the maximum reliability.

[0084] Next, the narrowing down criteria information is registered (step SA3), and the process ends.

[0085] Next, the processing performed at the parameter setting stage will be described.

[0086] 6 is a flowchart showing an example of processing performed in the parameter setting stage. In the parameter setting stage, first, the receiving unit 21 receives voice information (step SB1).

[0087] Next, the narrowed-down information acquisition unit 23 acquires the narrowed-down information (step SB2).

[0088] Next, the provisional parameter selection unit 24 selects provisional parameters (step SB3).

[0089] Next, the recognition unit 25 recognizes the voice information (step SB4). The recognition unit 25 recognizes the voice information based on the plurality of sets of provisional parameters selected by the provisional parameter selection unit 24, for example.

[0090] Next, the parameter selection unit 26 selects one of the provisional setting parameters as a setting parameter based on the recognition result by the recognition unit 25 (step SB5).

[0091] Next, the output unit 27 outputs the setting parameters selected by the parameter selection unit 26 (step SB6).

[0092] Next, the setting parameter storage unit 28 stores the setting parameters (step SB7), and the process ends.

[0093] Next, the processing performed after the parameters are set will be described.

[0094] 7 is a flowchart showing an example of processing that is performed after the parameters are set. After the parameters are set, the receiving unit 21 receives voice information (step SC1).

[0095] Next, the recognition unit 25 recognizes the voice information (step SC2). At this time, the recognition unit 25 recognizes the voice information using a voice recognition model associated with the setting parameters. As a result, for example, a command for the numerical control device 2 is generated, and the numerical control device 2 is controlled or various settings are made based on the generated command. Note that when processing such as network setting using the voice recognition device 20 is completed, the voice recognition device 20 ends this processing.

[0096] 6, the voice information received by the reception unit 21 is used only for parameter setting. However, if the voice information received by the reception unit 21 is recognized as, for example, a command to be executed by the numerical control device 2, a control unit (not shown) of the numerical control device 2 may execute this command.

[0097] As described above, the voice recognition device 20 includes a receiving unit 21 that receives input of voice information, a parameter storage unit 22 that stores a plurality of parameters for setting a voice recognition model, a provisional setting parameter selection unit 24 that selects provisional setting parameters to be provisionally set from the plurality of parameters based on narrowing-down information, a recognition unit 25 that recognizes voice information based on the selected provisional setting parameters, and a parameter selection unit 26 that selects one of the provisional setting parameters as a setting parameter based on information indicating the recognition result of the recognized voice information.

[0098] Therefore, appropriate parameters are automatically selected in the speech recognition device 20. In other words, an optimum speech recognition model for performing speech recognition is automatically selected in the speech recognition device 20. As a result, the workload of selecting a speech recognition model can be reduced.

[0099] Furthermore, the speech recognition device 20 narrows down the parameters to be provisionally set based on the narrowing-down information. Therefore, the number of parameters to be provisionally set or the number of parameter sets is reduced. As a result, the processing load related to speech recognition is reduced in the speech recognition device 20.

[0100] The plurality of parameters includes at least one of a plurality of acoustic setting parameters for setting an acoustic model and a plurality of grammar setting parameters for setting a grammar model. The information indicating the recognition result includes information indicating the reliability of the recognition result. The parameter selection unit 26 selects setting parameters that maximize the reliability. Alternatively, the parameter selection unit 26 selects setting parameters that maximize the reliability. Alternatively, the parameter selection unit 26 selects setting parameters that maximize the reliability. Therefore, the speech recognition device 20 can improve the accuracy of speech recognition.

[0101] The narrowed-down information also includes speaker information that identifies the speaker who uttered the voice. In this case, the provisional parameter selection unit 24 selects provisional parameters based on the speaker information. Therefore, optimal parameter settings for the speaker are performed. In other words, the optimal model is selected from many speech recognition models.

[0102] For example, in the field of industrial machinery such as factories, the scope of each worker's job is limited. Therefore, even if various grammar models corresponding to each job are pre-registered, the most appropriate grammar model is selected. In this case, each grammar model is a grammar model that recognizes speech information specific to each job. Therefore, the speech recognition device 20 can improve the accuracy of speech recognition.

[0103] The narrowed-down information also includes time information that specifies the time when the receiving unit 21 received the voice information. For example, continuous machining is performed in the industrial machine 1 at night, so the possibility of tool setting being performed in the numerical control device 2 is low. Therefore, even if the voice recognition device 20 receives voice information at night, the temporary setting parameter selection unit 24 can prevent the tool setting parameters from being selected as temporary setting parameters. In other words, the voice recognition device 20 can efficiently narrow down the temporary setting parameters.

[0104] The narrowed-down information also includes location information that identifies the location where the microphone 9 that captures the voice is installed. For example, in a factory, workers are assigned in advance to industrial machines 1. Therefore, when the microphone 9 is installed in the industrial machine 1, if the location information is identified, the speaker can generally be identified. Therefore, the voice recognition device 20 can efficiently narrow down the provisional setting parameters based on the location information.

[0105] Moreover, the speech recognition device 20 further includes a setting parameter storage unit 28 that stores the setting parameters selected by the parameter selection unit 26, and the recognition unit 25 recognizes speech information based on the setting parameters stored in the setting parameter storage unit 28. Therefore, the speech recognition device 20 can recognize speech information based on automatically set parameters.

[0106] The voice recognition device 20 further includes an output unit 27 that outputs information indicating that the setting parameter storage unit 28 has stored the setting parameter when the setting parameter storage unit 28 stores the setting parameter. Therefore, the voice recognition device 20 can notify the operator whether or not parameter setting has been performed.

[0107] In the above-described embodiment, as an example, the narrowed-down information acquisition unit 23 acquires speaker information by comparing the voice information of pre-registered speakers with the voice information received by the reception unit 21. However, the speaker information may be acquired by other methods. For example, the narrowed-down information acquisition unit 23 may include a trained model for inferring speaker information.

[0108] 8 is a block diagram showing an example of the function of the narrowed-down information acquisition unit 23. The narrowed-down information acquisition unit 23 includes a trained model M for identifying a speaker. The narrowed-down information acquisition unit 23 inputs the speech information received by the reception unit 21 into the trained model M, and obtains an output indicating an inference result of speaker information from the trained model M.

[0109] The trained model M is generated, for example, by having a machine learning machine learn speech information of various speeches uttered by multiple speakers who use the speech recognition device 20. The machine learning machine generates the trained model M by, for example, performing deep learning. Note that the machine learning machine may be provided in the speech recognition device 20 or in a device other than the speech recognition device 20.

[0110] The voice recognition device 20 may further include a history information storage unit that stores history information of the setting parameters selected by the parameter selection unit .

[0111] Fig. 9 is a block diagram showing an example of the functions of the voice recognition device 20. The voice recognition device 20 shown in Fig. 9 differs from the voice recognition device 20 shown in Fig. 2 in that the setting parameter storage unit 28 includes a history information storage unit 281.

[0112] The history information storage unit 281 stores history information of the setting parameters selected by the parameter selection unit 26. Therefore, every time a new setting parameter is selected by the parameter selection unit 26, the history information storage unit 281 accumulates and stores the selected setting parameter.

[0113] In this case, the provisional parameter selection unit 24 may select a provisional parameter to be provisionally set from a plurality of parameters by using the history information stored in the history information storage unit 281 as narrowing-down information. In other words, the narrowing-down information includes the history information stored in the history information storage unit 281.

[0114] The provisional setting parameter selection unit 24 selects provisional setting parameters from, for example, setting parameters stored in the history information storage unit 281, i.e., parameters that have already been used when the recognition unit 25 performs speech recognition. This reduces the processing load when selecting provisional setting parameters and increases the processing speed. In other words, it is possible to efficiently narrow down the provisional setting parameters.

[0115] In the above-described embodiment, the setting parameters selected by the parameter selection unit 26 are stored in the setting parameter storage unit 28. However, the setting parameters selected by the parameter selection unit 26 do not necessarily have to be stored in the setting parameter storage unit 28. In this case, the voice recognition device 20 may further include an instruction information receiving unit that receives instruction information instructing the setting parameter storage unit 28 not to store the setting parameters, and if the instruction information receiving unit does not receive the instruction information within a predetermined reception period, the setting parameter storage unit 28 may store the setting parameters selected by the parameter selection unit 26.

[0116] 10 is a block diagram showing an example of the functions of the voice recognition device 20. The voice recognition device 20 includes an instruction information receiving unit 29 in addition to the units of the voice recognition device 20 shown in FIG.

[0117] The instruction information receiving unit 29 receives instruction information indicating that the setting parameters selected by the parameter selecting unit 26 are not to be stored in the setting parameter storage unit 28. For example, when a setting parameter is selected by the parameter selecting unit 26, the output unit 27 outputs information indicating the recognition result derived from the selected setting parameter to the input / output device 3. The input / output device 3 displays the information indicating the recognition result output by the output unit 27 on a display screen.

[0118] 11 is a diagram showing an example of an image displayed on the display screen of the input / output device 3. An image for instructing not to store the setting parameters in the setting parameter storage unit 28 is displayed on the display screen. Specifically, the recognition result of the voice information accepted by the acceptance unit 21 and a button image for instructing not to accept the recognition result are displayed on the display screen. Here, the recognition result is a character string "I would like to check the IP address" which is a text conversion of the voice information, and a character string "confidence level 0.8" which indicates the accuracy of the recognition of the voice information.

[0119] For example, the input / output device 3 displays, for three seconds, information indicating the recognition result derived from the setting parameters selected by the parameter selection unit 26 on the display screen. If the button image is touched within three seconds, the instruction information receiving unit 29 receives instruction information instructing not to store the setting parameters in the setting parameter storage unit 28. In this case, the setting parameter storage unit 28 does not store the setting parameters.

[0120] On the other hand, if the button image is not touched within three seconds, the instruction information receiving unit 29 does not receive instruction information instructing not to store the setting parameters in the setting parameter storage unit 28. In this case, the setting parameter storage unit 28 stores the setting parameters selected by the parameter selection unit 26.

[0121] This allows the operator to select whether or not to set parameters depending on whether or not appropriate setting parameters have been selected by the parameter selection unit 26.

[0122] The present disclosure is not limited to the above-described embodiments, and can be appropriately modified without departing from the spirit and scope of the present disclosure. In the present disclosure, any of the components of the embodiments can be modified or omitted. [Explanation of symbols]

[0123] 1. Industrial machinery 2. Numerical control device 20 Voice recognition device 201 Hardware Processor 202 Bus 203 ROM 204 RAM 205 Non-volatile memory 206 First Interface 207 Axis control circuit 208 Spindle control circuit 209 PLC 210 I / O units 211 Second Interface 21 Reception 22 Parameter storage section 23 Filtering information acquisition unit 24 Provisional parameter selection section 25 Recognition part 251 Model Memory Unit 252 Dictionary storage unit 253 Recognition processing section 26 Parameter selection section 27 Output section 28 Setting parameter memory section 281 History information storage unit 29 Instruction Information Reception Department 3 Input / Output Devices 4 Servo amplifiers 5 Servo motors 6 Spindle amplifier 7 Spindle motor 8 Auxiliary equipment 9 Microphone M trained models

Claims

1. a reception unit that receives input of voice information; a parameter storage unit that stores a plurality of parameters for setting a speech recognition model; a provisional parameter selection unit that selects a provisional parameter to be provisionally set from the plurality of parameters based on the narrowed-down information; a recognition unit that recognizes the speech information using the speech recognition model stored in association with the selected provisional setting parameters; a parameter selection unit that selects one of the provisional parameters based on information including a reliability of a recognition result of the recognized speech information; Equipped with the plurality of parameters include at least one of a plurality of acoustic setting parameters for setting an acoustic model and a plurality of grammar setting parameters for setting a grammar model; Voice recognition device.

2. The speech recognition device according to claim 1 , wherein the parameter selection unit selects the provisionally set parameters that maximize the reliability.

3. The speech recognition device according to claim 1 , wherein the parameter selection unit selects the provisionally set parameters such that the reliability is equal to or greater than a predetermined threshold value.

4. 4. The speech recognition device according to claim 1, wherein the narrowed-down information includes speaker information for identifying a speaker who utters a voice.

5. The speech recognition device according to claim 4 , wherein the narrowed-down information is obtained by a trained model for inferring the speaker information.

6. The speech recognition device according to any one of claims 1 to 5, wherein the narrowed-down information includes either time information specifying a reception time at which the reception unit received the speech information, or location information specifying a location at which a microphone is installed.

7. a history information storage unit that stores history information of the provisional setting parameters selected by the parameter selection unit, 7. The speech recognition device according to claim 1, wherein the narrowed-down information includes the history information stored in the history information storage unit.

8. a setting parameter storage unit that stores the temporary setting parameters selected by the parameter selection unit, 8. The speech recognition device according to claim 1, wherein the recognition unit recognizes the speech information based on the provisional setting parameters stored in the setting parameter storage unit.

9. an instruction information receiving unit that receives instruction information that instructs the setting parameter storage unit not to store the temporary setting parameters; 9. The speech recognition device according to claim 8, wherein if the instruction information receiving unit does not receive the instruction information within a predetermined reception period, the setting parameter storage unit stores the provisional setting parameters selected by the parameter selection unit.

10. 10. The speech recognition device according to claim 8, further comprising an output unit that outputs, when the setting parameter storage unit stores the provisional setting parameters, information indicating that the setting parameter storage unit stores the provisional setting parameters.

Citation Information

Patent Citations

  • Word voice recognition equipment

    JP1985208800A

  • Voice recognition device, voice recognition method and voice recognition program

    JP2008064885A

  • Speech processor and program

    JP2009020352A

  • Voice recognition method and voice recognition device

    JP2020086437A

  • Machine-tool and control system

    JP2020160586A