Information processing device, information processing method, and computer program
Patent Information
- Application Number
- JP2025505089
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Filing Date
- 2025-08-19
- Publication Date
- 2025-10-31
AI Technical Summary
Existing information processing devices face challenges in accurately recognizing character strings, especially in specific contexts, due to insufficient learning and the need for time-consuming model generation, which affects speech recognition accuracy.
The implementation of an information processing device that extracts feature amounts related to speech recognition from character strings and estimates weights based on these features, using a combination of general-purpose and user-specific dictionaries to improve recognition accuracy and adaptability.
This approach enhances speech recognition accuracy by allowing quick adaptation to user-specific contexts and improving the recognition of character strings relevant to individual users, reducing the need for extensive model generation and enhancing user dictionary management.
Abstract
Description
Information processing device, information processing method, and recording medium
[0001] The present disclosure relates to the technical fields of an information processing device, an information processing method, and a recording medium.
[0002] Patent Literature 1 describes a technology that reads character information in a document to be subjected to speech recognition, tallying the frequency of appearance of words included in the character information, setting a priority for each word based on the tally of the frequency of appearance of the words, managing the registration or update of the prioritized words in a speech recognition dictionary, and recognizing newly input speech information while referring to the updated priority for each word, converting it into character information corresponding to the speech information, and outputting it. A plurality of types of speech recognition dictionaries are provided corresponding to the range of use of words, and registration or update is performed in the speech recognition dictionary corresponding to the range of use of the word by referring to the plurality of types of speech recognition dictionaries.
[0003] JP 2019-120763 A
[0004] An object of this disclosure is to provide an information processing device, an information processing method, and a recording medium that aim to improve upon the techniques described in prior art documents.
[0005] One aspect of the information processing device includes an extraction means for extracting features related to speech recognition from a character string, and an estimation means for estimating a weight corresponding to the character string based on the features, the weight being a quantity related to the speech recognition.
[0006] One aspect of the information processing method extracts features related to speech recognition from a character string, and estimates a weight corresponding to the character string based on the features.
[0007] One aspect of the recording medium has recorded thereon a computer program for causing a computer to execute an information processing method for extracting features related to speech recognition from a character string and estimating weights corresponding to the character string based on the features.
[0008] FIG. 1 is a block diagram showing the configuration of an information processing device in a first embodiment. FIG. 2 is a block diagram showing the configuration of a voice recognition engine. FIG. 3 is a block diagram showing the configuration of an information processing device in a second embodiment. FIG. 4 is a flowchart showing the flow of information processing operations of the information processing device in the second embodiment. FIG. 5 is a block diagram showing the configuration of an information processing device in a third embodiment. FIG. 6 is a block diagram showing the configuration of an information processing device in a fourth embodiment. FIG. 7 is a block diagram showing the configuration of an information processing device in a fifth embodiment. FIG. 8(a) is a block diagram showing the configuration of a weight determination unit in the fifth embodiment, and FIG. 8(b) is a flowchart showing the flow of weight determination operations of the weight determination unit. FIG. 9 is a flowchart showing the flow of information processing operations of the information processing device in the fifth embodiment. FIG. 10 is a block diagram showing the configuration of an information processing device in a sixth embodiment. FIG. 11 is a flowchart showing the flow of information processing operations of the information processing device in the sixth embodiment. FIG. 12 is a block diagram showing the configuration of an information processing device in a seventh embodiment. FIG. 13 is a block diagram showing the configuration of an information processing device in an eighth embodiment.
[0009] Hereinafter, embodiments of an information processing device, an information processing method, and a recording medium will be described with reference to the drawings. [1: First Embodiment]
[0010] A first embodiment of an information processing device, an information processing method, and a recording medium will be described below. Hereinafter, the first embodiment of the information processing device, the information processing method, and the recording medium will be described using an information processing device 1 to which the first embodiment of the information processing device, the information processing method, and the recording medium is applied. [1-1: Configuration of Information Processing Device 1]
[0011] The configuration of an information processing device 1 in the first embodiment will be described with reference to Fig. 1. Fig. 1 is a block diagram showing the configuration of the information processing device 1 in the first embodiment.
[0012] As shown in FIG. 1 , the information processing device 1 includes an extraction unit 11 and an estimation unit 12. The extraction unit 11 extracts features related to speech recognition from a character string. A character string is a word representing a group of arranged characters, but in this embodiment, a character that makes sense as a single character may also be included in a character string. Speech recognition in this embodiment is an operation of estimating, from speech, a character string corresponding to the speech.
[0013] The estimation unit 12 estimates weights corresponding to character strings based on the feature quantities. The weights are used for speech recognition. In speech recognition, the weights corresponding to character strings are used to estimate character strings corresponding to speech from the speech.
[0014] The above-described process can be said to be a process in which the extraction unit 11 extracts features related to speech recognition from a character string, and the estimation unit 12 adjusts the accuracy of speech recognition for the character string based on the features. [1-2: Technical Effects of the Information Processing Device 1]
[0015] The information processing device 1 according to the first embodiment estimates weights for speech recognition of a character string based on feature quantities for speech recognition extracted from the character string, thereby achieving highly accurate speech recognition. [2: Second Embodiment]
[0016] A second embodiment of an information processing device, an information processing method, and a recording medium will be described below. Hereinafter, the second embodiment of the information processing device, the information processing method, and the recording medium will be described using an information processing device 2 to which the second embodiment of the information processing device, the information processing method, and the recording medium is applied. [2-1: Speech Recognition]
[0017] As described above, the speech recognition in this embodiment is an operation of estimating a character string corresponding to a speech from the speech. For example, speech recognition may be realized by dividing time-series speech information into a predetermined number of frames and estimating a character string for each frame.
[0018] In this embodiment, Japanese speech recognition will be described as an example. However, speech recognition other than Japanese can also be performed in the same manner as in this embodiment, as long as the speech recognition is performed using a mechanism similar to that of this embodiment.
[0019] As described above, a character string is a word representing a collection of arranged characters. In this embodiment, however, characters that have meaning by a single character, such as "sentence", "character", "column", etc., may be included in the character string. A character string includes an arrangement of characters that includes a single language element such as a word. A character string includes an arrangement of characters that includes two or more language elements such as compound words and compound terms.
[0020] The "character string" in this embodiment may include a unit of a collection of characters to be recognized by speech recognition. The unit of a collection of characters to be recognized by speech recognition may include a phrase that is a delimited character string with meaning.
[0021] FIG. 2 is a block diagram illustrating the configuration of the speech recognition engine SE. When voice is input, the speech recognition engine SE uses a speech recognition model SM and a general-purpose dictionary GD to estimate a character string corresponding to the input voice from the input voice.
[0022] When voice is input, the speech recognition model SM outputs a character string corresponding to the voice. The speech recognition model SM used by the speech recognition engine SE may be a model that is not specialized for a specific field, application, etc. and can be used generally. The speech recognition model SM may be a model that is learned using document data in which the included character strings do not tend to be character strings used in a specific field, application, etc. and that is generated by learning general-purpose speech recognition.
[0023] The general-purpose dictionary GD is a dictionary in which the character strings recognized by the speech recognition engine SE are registered. The general-purpose dictionary GD may be a dictionary that is not specialized for a specific field, application, etc. and can be used generally.
[0024] In addition to the speech recognition model SM and the general-purpose dictionary GD, the speech recognition engine SE may use a user dictionary UD to estimate a character string corresponding to the voice from the voice. The information processing apparatus 2 in the second embodiment generates the user dictionary UD used by the speech recognition engine SE.
[0025] The user dictionary UD is a dictionary in which character strings to be recognized by the voice recognition engine SE are registered. A user dictionary UD is provided for each user who uses the voice recognition engine SE. The voice recognition engine SE uses different user dictionaries UD depending on the user.
[0026] A user may be a single person or a group including multiple people. In this embodiment, people who use the same user dictionary UD may be referred to as users. For example, a user may be people who use a speech recognition engine SE for the same purpose.
[0027] A character string registered in the user dictionary UD may be a sequence of characters that is meaningful to the user. A character string registered in the user dictionary UD may also be a character string that is not registered in the general-purpose dictionary GD. For example, if "A" and "B" are registered in the general-purpose dictionary GD but "AB" is not, and the user wishes to be recognized as a character string that is a sequence of "AB," "AB" may be registered in the user dictionary UD. Furthermore, a character string registered in the user dictionary UD may be a character string that is registered in the general-purpose dictionary GD but that the user wishes to be more easily recognized when used by the user than when used by other users.
[0028] The character strings registered in the user dictionary UD may be, for example, character strings that are often used in a particular field, application, etc. The character strings registered in the user dictionary UD may include words and phrases that are not included in general dictionaries.
[0029] The users may be people who work in the same industry. The users may be people who perform the same job duties. For example, the speech recognition engine SE may be applied to situations where words and phrases that are only used internally within a company are frequently used, such as transcribing company minutes. In this case, the character strings registered in the user dictionary UD may include words and phrases that are only used internally within the company. The character strings registered in the user dictionary UD may include colloquial language. The character strings registered in the user dictionary UD may include product names, industry terms, etc.
[0030] By using the user dictionary UD in the speech recognition engine SE, it is possible to make it easier for a character string that the user wants to recognize to appear in the recognition results of speech recognition by the speech recognition engine SE. By using the user dictionary UD in the speech recognition engine SE, it is possible to increase the possibility of estimating the character string that the user wants to recognize from speech.
[0031] Weights corresponding to character strings are registered in the general-purpose dictionary GD and the user dictionary UD in association with the character strings. The weights corresponding to character strings are used to appropriately recognize the character strings. By associating appropriate weights with character strings, the accuracy of character string estimation in speech recognition can be improved. The weights corresponding to character strings may indicate the probability of occurrence of the character strings in speech recognition. The weights corresponding to character strings may be expressed as the logarithm of the probability of occurrence of the character strings in speech recognition. The weights corresponding to character strings registered in the user dictionary UD may be used in the speech recognition engine SE to realize speech recognition suitable for the user.
[0032] Although a user can specify a character string to be speech-recognized, it is often difficult to specify an appropriate weight for the character string. Therefore, the information processing device 2 in this embodiment automatically assigns an appropriate weight to the character string specified by the user. The information processing device 2 in the second embodiment generates a user dictionary UD in which character strings and weights are registered in association with each other. [2-4: Configuration of the information processing device 2]
[0033] 3 is a block diagram showing the configuration of an information processing device 2 in the second embodiment. As shown in FIG. 3, the information processing device 2 includes a calculation device 21 and a storage device 22. The information processing device 2 may further include a communication device 23, an input device 24, and an output device 25. However, the information processing device 2 does not necessarily have to include at least one of the communication device 23, the input device 24, and the output device 25. The calculation device 21, the storage device 22, the communication device 23, the input device 24, and the output device 25 may be connected via a data bus 26.
[0034] The arithmetic device 21 includes, for example, at least one of a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), and an FPGA (Field Programmable Gate Array). The arithmetic device 21 reads a computer program. For example, the arithmetic device 21 may read a computer program stored in the storage device 22. For example, the arithmetic device 21 may read a computer program stored in a computer-readable, non-transitory recording medium using a recording medium reading device (e.g., an input device 24 described later) not shown in the drawings that is included in the information processing device 2. The arithmetic device 21 may acquire (i.e., download or read) the computer program from a device (not shown) located outside the information processing device 2 via the communication device 23 (or another communication device). The arithmetic device 21 executes the read computer program. As a result, logical functional blocks for executing the operations to be performed by the information processing device 2 are realized within the arithmetic device 21. In other words, the arithmetic device 21 can function as a controller for realizing logical functional blocks for executing the operations (in other words, processing) to be performed by the information processing device 2.
[0035] 3 shows an example of logical functional blocks implemented within the arithmetic device 21 to execute information processing operations. As shown in FIG. 3, the arithmetic device 21 implements an extraction unit 211, which is a specific example of the "extraction means" described in the appendix below, an estimation unit 212, which is a specific example of the "estimation means" described in the appendix below, an input acceptance unit 213, and a registration unit 214, which is a specific example of the "registration means" described in the appendix below. However, at least one of the input acceptance unit 213 and the registration unit 214 does not necessarily have to be implemented within the arithmetic device 21. Details of the operations of the extraction unit 211, the estimation unit 212, the input acceptance unit 213, and the registration unit 214 will be described later with reference to FIG. 4.
[0036] The storage device 22 can store desired data. For example, the storage device 22 may temporarily store a computer program executed by the arithmetic device 21. The storage device 22 may temporarily store data temporarily used by the arithmetic device 21 when the arithmetic device 21 is executing a computer program. The storage device 22 may store data to be stored long-term by the information processing device 2. The storage device 22 may include at least one of a RAM (Random Access Memory), a ROM (Read Only Memory), a hard disk device, a magneto-optical disk device, an SSD (Solid State Drive), and a disk array device. In other words, the storage device 22 may include a non-transitory recording medium. A user dictionary UD, which is a specific example of the "user dictionary" described in the appendix below, may be implemented within the storage device 22. If the user dictionary UD is not implemented within the storage device 22, the user dictionary UD may be implemented in a storage device external to the information processing device 2.
[0037] The communication device 23 is capable of communicating with devices external to the information processing device 2 via a communication network (not shown). The communication device 23 may be a communication interface based on standards such as Ethernet (registered trademark), Wi-Fi (registered trademark), Bluetooth (registered trademark), or USB (Universal Serial Bus).
[0038] The input device 24 is a device that accepts information input to the information processing device 2 from outside the information processing device 2. For example, the input device 24 may include an operation device (e.g., at least one of a keyboard, a mouse, and a touch panel) that can be operated by an operator of the information processing device 2. For example, the input device 24 may include a reading device that can read information recorded as data on a recording medium that can be externally attached to the information processing device 2. In the second embodiment, the information processing device 2 may include, as the input device 24, a user dictionary interface UIF, which is a specific example of a "user interface" described in the appendix below.
[0039] The output device 25 is a device that outputs information to the outside of the information processing device 2. For example, the output device 25 may output information as an image. That is, the output device 25 may include a display device (a so-called display) that can display an image showing the information to be output. For example, the output device 25 may output information as sound. That is, the output device 25 may include an audio device (a so-called speaker) that can output sound. For example, the output device 25 may output information on paper. That is, the output device 25 may include a printing device (a so-called printer) that can print desired information on paper. [2-3: Information Processing Operation Performed by the Information Processing Device 2]
[0040] The flow of the information processing operation performed by the information processing device 2 will be described with reference to Fig. 4. Fig. 4 is a flowchart showing the information processing operation performed by the information processing device 2.
[0041] The input accepting unit 213 accepts character strings to be registered in the user dictionary UD. The input accepting unit 213 may accept character strings that the user wishes to register in the user dictionary UD. The user dictionary interface UIF is a user interface that the user uses to edit the user dictionary UD. When the user wants to register a new character string in the user dictionary UD, the user can operate the user dictionary interface UIF to input the character string. The user can register a new character string in the user dictionary UD by operating the user dictionary interface UIF to input the character string. The user may perform an input operation for each character string, or may be able to input multiple character strings at once. The input accepting unit 213 may accept character strings input by the user using the user dictionary interface UIF. On the other hand, the input accepting unit 213 may accept multiple character strings at once, for example, during a time period when the user does not use the speech recognition engine SE.
[0042] 4, the input receiving unit 213 determines whether or not a character string has been received (step S20). The input receiving unit 213 may determine whether or not the user has input a character string into the user dictionary interface UIF.
[0043] If a character string is received (step S20: Yes), the extraction unit 211 extracts features related to speech recognition from the received character string (step S21). In this embodiment, the "feature" related to speech recognition may be, for example, a quantity representing features related to speech recognition of a speech recognition engine SE that uses a user dictionary UD. Furthermore, the "feature" related to speech recognition may be, for example, a quantity representing features related to the speech when the character string is converted into speech, such as the reading or pronunciation of the character string. Note that a specific method for extracting features will be described in other embodiments below.
[0044] The estimation unit 212 estimates a weight corresponding to the character string based on the feature amount related to speech recognition (step S22). The estimation unit 212 may estimate, for example, an arbitrary numerical value as the weight. Alternatively, the estimation unit 212 may estimate one of a plurality of predetermined identifiers as the weight. For example, the estimation unit 212 may estimate one of ten integers from 1 to 10 as the weight. In this case, a larger number may indicate a larger weight. Alternatively, the estimation unit 212 may estimate one of three classes, for example, "large," "medium," and "small," as the weight. "Large," "medium," and "small" represent the magnitude of the weight. In other words, if the estimation unit 212 estimates the "large" class, a larger weight is estimated than if the estimation unit 212 estimates the "medium" class. Also, if the estimation unit 212 estimates the "medium" class, a larger weight is estimated than if the estimation unit 212 estimates the "small" class.
[0045] The estimation unit 212 may also estimate a weight to be assigned to a target character string using a weight estimation model WM. The weight estimation model WM is a model generated so as to output a weight corresponding to a character string when a feature related to speech recognition extracted from the character string is input. The generation of the weight estimation model WM will be described in another embodiment described later.
[0046] The registration unit 214 associates the character string with the weight estimated by the estimation unit 212 and registers the associated character string in the user dictionary UD (step S23). The user dictionary UD registers the character string that the user wants to have speech recognized and an appropriate weight corresponding to the character string. [2-4: Technical Effects of the Information Processing Device 2]
[0047] The character strings to be recognized by speech differ depending on the business area, organization, etc., but if the speech recognition engine SE has not yet learned enough about the character string, it is difficult for the speech recognition engine SE to recognize the character string. Furthermore, a speech recognition model SM tailored to a user requires a relatively large amount of time and effort to generate. On the other hand, a user dictionary UD tailored to a user often requires less time and effort to generate than a speech recognition model SM tailored to a user.
[0048] Therefore, in the second embodiment, a common speech recognition model SM and a user dictionary UD tailored to the user are used to generate a user-specific user dictionary UD to achieve speech recognition tailored to the user. The information processing device 2 can quickly meet user needs by preparing a common speech recognition model SM and generating only a user dictionary UD for each user. The information processing device 2 can meet user requirements while using a general-purpose speech recognition model SM. Since the information processing device 2 can generate an appropriate user dictionary UD, it can provide a speech recognition engine SE tailored to the user without preparing a user-specific speech recognition model SM. Furthermore, since the information processing device 2 estimates appropriate weights corresponding to character strings, it can improve the accuracy of speech recognition.
[0049] Furthermore, since the information processing device 2 is provided with a user dictionary interface UIF, the user can register character strings in the user dictionary UD as needed, and adjustments according to the user's needs can be easily realized. [3: Third Embodiment]
[0050] A third embodiment of an information processing device, an information processing method, and a recording medium will be described below. Hereinafter, the third embodiment of the information processing device, the information processing method, and the recording medium will be described using an information processing device 3 to which the third embodiment of the information processing device, the information processing method, and the recording medium is applied. [3-1: Configuration of Information Processing Device 3]
[0051] FIG. 5 is a block diagram showing the configuration of an information processing device 3 according to the third embodiment. As shown in FIG. 5, the information processing device 3 according to the third embodiment includes a calculation device 21 and a storage device 22, similar to the information processing device 2 according to the second embodiment. Furthermore, the information processing device 3 according to the third embodiment may include a communication device 23, an input device 24, and an output device 25, similar to the information processing device 2 according to the second embodiment. However, the information processing device 3 does not necessarily include at least one of the communication device 23, the input device 24, and the output device 25. The information processing device 3 according to the third embodiment differs from the information processing device 2 according to the second embodiment in that the extraction unit 311 extracts features related to speech recognition from character strings using a feature extraction mechanism FE. Other features of the information processing device 3 may be the same as those of the information processing device 2 according to the second embodiment. Therefore, the following will describe in detail the differences from the previously described embodiments, and will omit appropriate descriptions of other overlapping features. [3-2: Frequency Features]
[0052] The extraction unit 311 in the third embodiment extracts features according to the degree of learning of speech recognition of a character string. The features related to speech recognition in the third embodiment are features related to speech recognition performed by a speech recognition engine SE that uses a speech recognition model SM. The features related to speech recognition in the third embodiment are features related to the degree of learning of the speech recognition model SM for the corresponding character string. The features related to speech recognition in the third embodiment may also be features related to the speech recognition performance of the speech recognition model SM for the corresponding character string.
[0053] Character strings that frequently appear in the training data (referred to as "speech recognition training data") used in generating the speech recognition model SM can be considered to have undergone more training to recognize the character string from speech. Therefore, the character string in question is often one for which training to recognize the character string from speech has progressed. Note that the speech recognition training data used in generating the speech recognition model SM may be sentences containing character strings that are not biased (not specialized) toward a particular field, purpose, etc. The speech recognition training data may include, for example, articles contained in general newspapers that are not biased toward a particular field, purpose, etc.
[0054] The extraction unit 311 in the third embodiment uses a feature extraction mechanism FE. The feature extraction mechanism FE is generated using speech recognition training data. The feature extraction mechanism FE is a mechanism generated according to the frequency of appearance of character strings included in the speech recognition training data. When a character string is input, the feature extraction mechanism FE may output a feature representing the frequency of appearance of the character string in the speech recognition training data. When a character string is input, the feature extraction mechanism FE may output a feature representing the degree of speech recognition training of the character string.
[0055] For example, the feature extraction mechanism FE may be a mechanism that outputs a feature that estimates a relatively large weight for a character string that has been trained relatively little in speech recognition, i.e., a character string that appears relatively infrequently in speech recognition training data. Alternatively, the feature extraction mechanism FE may be a mechanism that outputs a feature that estimates a relatively small weight for a character string that has been trained relatively much in speech recognition, i.e., a character string that appears relatively frequently in speech recognition training data. The feature extracted by the extraction unit 311 in the third embodiment is referred to as a "frequency feature." The extraction unit 311 extracts the frequency feature using the feature extraction mechanism FE. [3-3: Example 1 of feature extraction mechanism FE]
[0056] The feature extraction mechanism FE used by the extraction unit 311 may be a feature extraction model. The feature extraction model is a first form of the feature extraction mechanism FE used by the extraction unit 311. The feature extraction model is a model generated by learning using speech recognition training data. When a character string is input, the feature extraction model may output, as a feature, any quantity that represents the degree of learning of speech recognition for the character string. The feature extraction model may output, for example, any number between 0 and 1. [3-4: Example 2 of Feature Extraction Mechanism FE]
[0057] The feature extraction mechanism FE used by the extraction unit 311 may be a frequency feature table. The frequency feature table is a second form of the feature extraction mechanism FE used by the extraction unit 311. The frequency feature table may be a table generated by measuring the appearance frequency of each character string included in the speech recognition training data. In response to an input character string, the frequency feature table may return a predetermined quantity representing the degree of speech recognition training for the character string as a feature. The feature extraction model may return, for example, one of three types of classes as a feature.
[0058] In the third embodiment, for example, the character string desired to be registered in the user dictionary UD may be a character string that is registered in the general-purpose dictionary GD but is not commonly used and is difficult to recognize. In other words, in the third embodiment, the character string desired to be registered in the user dictionary UD may be a character string that has already been registered in the general-purpose dictionary GD.
[0059] Furthermore, although the feature quantities extracted by the extraction unit 311 in this embodiment are feature quantities related to speech recognition, the estimation unit 312 may estimate the weights by further using feature quantities such as notation feature quantities and part-of-speech feature quantities that can be obtained from character strings regardless of speech recognition. The estimation unit 312 may use feature quantities that can be obtained from character strings regardless of speech recognition in addition to feature quantities related to speech recognition. [3-5: Technical Effects of Information Processing Device 3]
[0060] The information processing device 3 in the third embodiment can extract feature quantities according to the characteristics of speech recognition using a general-purpose speech recognition model SM. Based on the feature quantities, the information processing device 3 can estimate weights suitable for speech recognition using the speech recognition model SM. [4: Fourth Embodiment]
[0061] A fourth embodiment of an information processing device, an information processing method, and a recording medium will be described below. Hereinafter, the fourth embodiment of the information processing device, the information processing method, and the recording medium will be described using an information processing device 4 to which the fourth embodiment of the information processing device, the information processing method, and the recording medium is applied. [4-1: Configuration of Information Processing Device 4]
[0062] FIG. 6 is a block diagram showing the configuration of an information processing device 4 according to the fourth embodiment. As shown in FIG. 6 , the information processing device 4 according to the fourth embodiment includes a calculation device 21 and a storage device 22, similar to the information processing device 2 according to the second embodiment and the information processing device 3 according to the third embodiment. Furthermore, the information processing device 4 according to the fourth embodiment may include a communication device 23, an input device 24, and an output device 25, similar to the information processing device 2 according to the second embodiment and the information processing device 3 according to the third embodiment. However, the information processing device 4 does not necessarily include at least one of the communication device 23, the input device 24, and the output device 25. The information processing device 4 according to the fourth embodiment differs from the information processing device 2 according to the second embodiment and the information processing device 3 according to the third embodiment in that the extraction unit 411 extracts features related to speech recognition from a character string using a pronunciation dictionary PD. Other features of the information processing device 4 may be the same as other features of at least one of the information processing device 2 according to the second embodiment and the information processing device 3 according to the third embodiment. Therefore, in the following, only the differences from the embodiments already described will be described in detail, and the description of other overlapping parts will be omitted as appropriate. [4-1: Pronunciation Features]
[0063] The extraction unit 411 in the fourth embodiment extracts features according to the pronunciation of a character string. The features extracted by the extraction unit 411 in the fourth embodiment are referred to as "pronunciation features."
[0064] Even when the same character string is used for notation, the appropriate weight may differ depending on the reading and pronunciation. Even when the same character string is used for notation, it may be desirable to assign different weights depending on the reading and pronunciation. The extraction unit 411 may be, for example, a mechanism for extracting pronunciation features according to the relationship between the notation of a character string and the pronunciation of the character string. [4-2: Pronunciation Dictionary PD]
[0065] The extraction unit 411 extracts features using a pronunciation dictionary PD that associates the notation of a character string with the pronunciation of the character string. The pronunciation dictionary PD used by the extraction unit 411 in the fourth embodiment may register notation of a character string and katakana characters indicating the pronunciation of the character string, as shown in Table 1 below. [Table 1: Example 1 of the Pronunciation Dictionary PD]
[0066]
[0067] Table 1 above illustrates that the pronunciation of "ringo" (apple) is "ringo," the pronunciation of "hounashi" (pineapple), the pronunciation of "seisaku" (correct) is "seikaku," and the pronunciation of "shoujiki" (honest) is "shoujiki," all of which are registered in the pronunciation dictionary PD.
[0068] Alternatively, the pronunciation dictionary PD used by the extraction unit 411 in the fourth embodiment may store notations of character strings and Roman characters indicating the pronunciation of the character strings, as shown in Table 2 below. [Table 2: Example 2 of the Pronunciation Dictionary PD]
[0069]
[0070] Table 2 above illustrates that the pronunciation of "apple" is " / ringo / ", the pronunciation of "pineapple" is " / painappuru / ", the pronunciation of "seikaku" is " / seikaku / ", and the pronunciation of "shojiki" is " / shojiki / ", all of which are registered in the pronunciation dictionary PD.
[0071] For example, when a character string is pronounced in an uncommon reading, a relatively large weight may be assigned to the character string. For example, as shown in Tables 1 and 2 above, when the readings "pineapple" and " / painappuru / " of "pineapple" are uncommon readings, the extraction unit 411 may extract a feature quantity that is estimated to have a relatively large weight.
[0072] Also, there may be cases where the same character string has multiple pronunciations. For example, the notation "鳳梨" may be pronounced as "パイナップル / / painappuru / " or "ほうり / / hori / ". The extraction unit 411 may extract feature amounts corresponding to each pronunciation so that appropriate weights are assigned to each pronunciation of the same character string.
[0073] Note that the estimation unit 412 of the fourth embodiment may estimate the weight corresponding to the character string based on both the feature amounts extracted by the extraction unit 411 of the fourth embodiment and the feature amounts extracted by the extraction unit 311 of the third embodiment. By using the feature amounts extracted by the extraction unit 411 of the fourth embodiment in combination with the feature amounts extracted by the extraction unit 311 of the third embodiment, the estimation unit 412 can estimate the weight more accurately.
[0074] Also, although the feature amounts extracted by the extraction unit 411 of the present embodiment are feature amounts related to speech recognition, the estimation unit 412 may further estimate the weight by using feature amounts such as notation feature amounts and part-of-speech feature amounts that can be obtained from the character string regardless of speech recognition. The estimation unit 412 may use the feature amounts that can be obtained from the character string regardless of speech recognition in combination with the feature amounts related to speech recognition. [4 - 4: Technical Effects of Information Processing Apparatus 4]
[0075] The information processing apparatus 4 in the fourth embodiment can extract feature amounts corresponding to the pronunciation of a character string and can estimate a weight suitable for speech recognition. [5: Fifth Embodiment]
[0076] The fifth embodiment of the information processing apparatus, the information processing method, and the recording medium will be described. Hereinafter, the fifth embodiment of the information processing apparatus, the information processing method, and the recording medium will be described using the information processing apparatus 5 to which the fifth embodiment of the information processing apparatus, the information processing method, and the recording medium is applied. [5 - 1: Configuration of Information Processing Apparatus 5]
[0077] FIG. 7 is a block diagram showing the configuration of an information processing device 5 according to the fifth embodiment. As shown in FIG. 7 , the information processing device 5 according to the fifth embodiment includes a calculation device 21 and a storage device 22, similar to the information processing device 2 according to the second embodiment to the information processing device 4 according to the fourth embodiment. Furthermore, the information processing device 5 according to the fifth embodiment may include a communication device 23, an input device 24, and an output device 25, similar to the information processing device 2 according to the second embodiment to the information processing device 4 according to the fourth embodiment. However, the information processing device 5 according to the fifth embodiment does not necessarily include at least one of the communication device 23, the input device 24, and the output device 25. The information processing device 5 according to the fifth embodiment differs from the information processing device 2 according to the second embodiment to the information processing device 4 according to the fourth embodiment in that a weight determination unit 515 and a generation unit 516 are further implemented within the calculation device 21. Other features of the information processing device 5 may be the same as at least one other feature of the information processing device 2 according to the second embodiment to the information processing device 4 according to the fourth embodiment. Therefore, in the following, only the differences from the embodiments already described will be described in detail, and the description of other overlapping parts will be omitted as appropriate. [5-2: Weight Determination Unit 515]
[0078] The weight determination unit 515 determines a weight corresponding to a character string based on the recognition result when speech corresponding to the character string is recognized. The weight determination unit 515 causes the speech recognition engine SE to recognize the speech corresponding to the character string, and determines the weight corresponding to the character string based on the recognition result. In this embodiment, the weight determined by the weight determination unit 515 is a weight that enables the speech recognition engine SE to perform appropriate speech recognition. In other words, the weight determination unit 515 can be said to be a mechanism that calculates a weight necessary for successful speech recognition by the speech recognition engine SE.
[0079] Fig. 8(a) is a block diagram showing the configuration of the weight determination unit 515, and Fig. 8(b) is a flowchart showing the flow of the weight determination operation of the weight determination unit 515. As shown in Fig. 8(a), the weight determination unit 515 includes a speech synthesis unit 5151, a speech recognition unit 5152, and a weight determination unit 5153. [5-3: Generation of training data]
[0080] 8B, the speech synthesis unit 5151 synthesizes speech corresponding to the input character string (step S50). The speech recognition unit 5152 causes the speech recognition engine SE to recognize the speech synthesized in step S50 (step S51).
[0081] The weight determination unit 5153 determines a weight based on the recognition result, i.e., the character string recognized by the speech recognition engine SE (step S52). The weight determination unit 5153 may determine a weight based on a comparison between the character string recognized by the speech recognition engine SE and the character string input to the speech synthesis unit 5151. The weight determination unit 5153 may determine that the speech recognition is correct when the character string recognized by the speech recognition engine SE and the character string input to the speech synthesis unit 5151 are similar to each other by a predetermined amount or more. If the speech recognition is correct, the weight determination unit 5153 determines that it is not necessary to assign a weight, or that a small weight should be assigned. If the speech recognition is correct, the weight determination unit 5153 may determine that the smallest possible weight should be assigned.
[0082] The weight determination unit 5153 may determine that the speech recognition is incorrect when the character string recognized by the speech recognition engine SE and the character string input to the speech synthesis unit 5151 are not similar to each other by a predetermined amount or more. If the speech recognition is incorrect, the weight determination unit 5153 determines the weight required to recognize the correct character string.
[0083] The larger the weight corresponding to a character string, the easier it is to recognize the character string, and the smaller the weight corresponding to a character string, the harder it is to recognize the character string. In other words, if the weight corresponding to a character string is too small, the speech recognition engine SE cannot recognize the character string. On the other hand, if the weight corresponding to a character string is too large, the speech recognition engine SE can recognize the character string, but is more likely to mistakenly recognize other character strings as the character string.
[0084] For example, suppose that the voice recognition engine SE recognizes the character string "AA" as the character string "AE" included in the character string input to the voice synthesis unit 5151. In this case, the weight determination unit 515 may determine the weight corresponding to the character string "AA" based on the criterion of how large the weight corresponding to the character string "AA" should be so that the voice recognition engine SE does not erroneously recognize it as "AE" but recognizes it as "AA."
[0085] The weight determined by the weight determination unit 5153 is referred to as the "correct weight." The weight determination unit 515 generates training data, which is data that associates the character string input to the speech synthesis unit 5151 with the correct weight, which is the determination result by the weight determination unit 5153. The weight determination unit 515 may also be referred to as a mechanism that generates training data. [5-4: Generation of Weight Estimation Model WM]
[0086] The generation unit 516 generates a weight estimation model WM to be used by the estimation unit 512, using the training data generated by the weight determination unit 515. The generation unit 516 performs learning using the training data generated by the weight determination unit 515. The weight estimation model WM is generated based on the results of speech recognition by the speech recognition engine SE.
[0087] The operation of generating the weight estimation model WM performed by the information processing device 5 will be described with reference to Fig. 9. Fig. 9 is a flowchart showing the flow of the operation of generating the weight estimation model WM performed by the information processing device 5.
[0088] 9 , a character string included in the training data is input to the extraction unit 511 (step S53). The extraction unit 511 extracts features related to speech recognition from the character string included in the training data (step S54). The estimation unit 512 inputs the features extracted in step S53 to the weight estimation model WM, and outputs a weight corresponding to the character string output by the weight estimation model WM (step S55).
[0089] The generation unit 516 causes the weight estimation model WM to learn a weight estimation method based on the correct weights included in the training data and the weights output by the weight estimation model WM (step S56). The generation unit 516 may cause the weight estimation model WM to learn a weight estimation method based on a loss function in which the loss increases as the correct weights and the weights output by the weight estimation model WM become less similar. The generation unit 516 may generate the weight estimation model WM by adjusting the parameters of the weight estimation model WM so that this loss function becomes smaller (preferably, minimized).
[0090] The weight estimation model WM generated in this manner is used in the estimation operation by the estimation unit 212 in the second embodiment, the estimation operation by the estimation unit 312 in the third embodiment, and the estimation operation by the estimation unit 412 in the fourth embodiment. [5-5: Technical Effects of the Information Processing Device 5]
[0091] The information processing device 5 in the fifth embodiment generates a weight estimation model WM that outputs weights by using weights suited to the characteristics of a general-purpose speech recognition engine SE as correct weights. When the weight estimation model WM is used, it is possible to accurately estimate weights corresponding to character strings. [6: Sixth Embodiment]
[0092] An information processing device, an information processing method, and a recording medium according to a sixth embodiment will be described below. The sixth embodiment of the information processing device, the information processing method, and the recording medium will be described below using an information processing device 6 to which the sixth embodiment of the information processing device, the information processing method, and the recording medium is applied.
[0093] When the information processing device 6 wants to register a character string related to a specific field, purpose, etc., the information processing device 6 may re-learn the specific field, purpose, etc. In the sixth embodiment, the re-learning of the specific field, purpose, etc. by the information processing device 6 will be described. [6-1: Configuration of the information processing device 6]
[0094] FIG. 10 is a block diagram showing the configuration of an information processing device 6 according to the sixth embodiment. As shown in FIG. 10 , the information processing device 6 according to the sixth embodiment includes a calculation device 21 and a storage device 22, similar to the information processing device 2 according to the second embodiment to the information processing device 5 according to the fifth embodiment. Furthermore, the information processing device 6 according to the sixth embodiment may include a communication device 23, an input device 24, and an output device 25, similar to the information processing device 2 according to the second embodiment to the information processing device 5 according to the fifth embodiment. However, the information processing device 6 according to the sixth embodiment does not necessarily include at least one of the communication device 23, the input device 24, and the output device 25. The information processing device 6 according to the sixth embodiment differs from the information processing device 2 according to the second embodiment to the information processing device 5 according to the fifth embodiment in that a speech recognition model generation unit 617 is further implemented within the calculation device 21. Other features of the information processing device 6 may be the same as at least one other feature of the information processing device 2 according to the second embodiment to the information processing device 5 according to the fifth embodiment. Therefore, in the following, only the differences from the embodiments already described will be described in detail, and the description of other overlapping parts will be omitted as appropriate. [6-2: Information Processing Operation Performed by Information Processing Device 6]
[0095] The information processing operation performed by the information processing device 6 will be described with reference to Fig. 11. Fig. 11 is a flowchart showing the flow of the information processing operation performed by the information processing device 6.
[0096] First, document data containing many character strings related to a specific field, application, etc. is prepared as speech recognition training data (step S60). The speech recognition model generation unit 617 uses the speech recognition training data prepared in step S60 to train speech recognition related to the specific field, application, etc. (referred to as "specific speech recognition") and generates a specific speech recognition model (step S61). The specific speech recognition model may be a speech recognition model that excels in speech recognition for a specific field, application, etc.
[0097] The weight determination unit 615 determines a weight based on the input character string and the character string recognized by the speech recognition engine SE using the specific speech recognition model (step S62). The weight estimation model generation unit 616 generates a specific weight estimation model related to a specific field, application, etc. using the input character string and the weight determined by the weight determination unit 615 as training data (step S63). The weight estimation model generation unit 616 generates the specific weight estimation model by performing machine learning based on the weight estimated by the estimation unit 612 based on the feature extracted from the input character string and the weight determined by the weight determination unit 615. The specific weight estimation model is a model generated so as to output a weight corresponding to the character string when a feature related to specific speech recognition extracted from the character string is input. The specific weight estimation model may estimate a weight appropriate for a character string related to a specific field, application, etc.
[0098] The regenerated specific weight estimation model is used to generate a user dictionary UD to be used for speech recognition using the regenerated specific speech recognition model. The information processing device 6 also generates the user dictionary UD according to the operational flow shown in FIG. 4 referred to in the second embodiment.
[0099] If a character string is received (step S20: Yes), the extraction unit 611 extracts a feature quantity related to specific speech recognition from the received character string (step S21). The estimation unit 612 estimates a weight corresponding to the character string based on the feature quantity related to specific speech recognition using a specific weight estimation model. The registration unit 614 associates the character string with the weight estimated by the estimation unit 612 and registers the associated character string in the user dictionary UD (step S23). The registration unit 614 generates a user dictionary UD related to a specific field. [6-3: Technical Effects of Information Processing Device 6]
[0100] The information processing device 6 in the sixth embodiment can perform re-learning according to a specific field, application, etc. [7: Seventh Embodiment]
[0101] A seventh embodiment of an information processing device, an information processing method, and a recording medium will be described below. Hereinafter, the seventh embodiment of the information processing device, the information processing method, and the recording medium will be described using an information processing device 7 to which the seventh embodiment of the information processing device, the information processing method, and the recording medium is applied. [7-1: Information processing operation performed by the information processing device 7]
[0102] In the seventh embodiment, the user may specify a weight corresponding to a character string by operating the user dictionary interface UIF. In this case, the estimation unit 712 may estimate the weight by taking into account the weight specified by the user. The information processing device 7 updates the information registered in the user dictionary UD in response to the user's operation of the user dictionary interface UIF.
[0103] Alternatively, for example, if a character string registered in the user dictionary UD is recognized better than the user expected, the user may be able to edit the weight by operating the user dictionary interface UIF so that the weight corresponding to the character string in question becomes smaller. Also, if a character string registered in the user dictionary UD is not recognized as well as the user expected, the user may be able to edit the weight by operating the user dictionary interface UIF so that the weight corresponding to the character string in question becomes larger.
[0104] Furthermore, for example, if the purpose of use of a character string registered in the user dictionary UD changes and the user wants to change the weight, the user may be able to increase or decrease the weight of the corresponding character string by operating the user dictionary interface UIF. [7-2: Technical Effects of the Information Processing Device 7]
[0105] The information processing device 7 in the seventh embodiment can register weights according to the user's requests in the user dictionary UD, and can provide voice recognition according to the user's requests. [8: Eighth Embodiment]
[0106] An eighth embodiment of an information processing device, an information processing method, and a recording medium will be described below. Hereinafter, the eighth embodiment of the information processing device, the information processing method, and the recording medium will be described using an information processing device 8 to which the eighth embodiment of the information processing device, the information processing method, and the recording medium is applied. [8-1: Configuration of Information Processing Device 8]
[0107] FIG. 13 is a block diagram showing the configuration of an information processing device 8 according to the eighth embodiment. As shown in FIG. 13 , the information processing device 8 according to the eighth embodiment includes a calculation device 21 and a storage device 22, similar to the information processing device 2 according to the second embodiment to the information processing device 7 according to the seventh embodiment. Furthermore, the information processing device 8 according to the eighth embodiment may include a communication device 23, an input device 24, and an output device 25, similar to the information processing device 2 according to the second embodiment to the information processing device 7 according to the seventh embodiment. However, the information processing device 8 does not necessarily have to include at least one of the communication device 23, the input device 24, and the output device 25. The information processing device 8 according to the eighth embodiment differs from the information processing device 2 according to the second embodiment to the information processing device 7 according to the seventh embodiment in that a speech recognition unit 818 is further implemented in the calculation device 21, and a speech recognition model SM, a general-purpose dictionary GD, a user dictionary UD, and a weight estimation model WM are implemented in the storage device 22. Other features of the information processing device 8 may be the same as at least one other feature of the information processing device 2 in the second embodiment to the information processing device 7 in the seventh embodiment. Therefore, in the following, differences from the already described embodiments will be described in detail, and descriptions of other overlapping parts will be omitted as appropriate. [8-2: Information Processing Operation Performed by Information Processing Device 8]
[0108] The speech recognition unit 818 may be a mechanism that causes a speech recognition engine SE installed in the information processing device 8 to perform speech recognition. The speech recognition engine SE performs speech recognition using a user dictionary UD in addition to a speech recognition model SM and a general-purpose dictionary GD. The information processing device 8 has the speech recognition model SM, the general-purpose dictionary GD, and the user dictionary UD built in. The user dictionary UD registers weights corresponding to character strings estimated by at least one of the estimation operations described in the second embodiment to the fourth embodiment, and the character strings themselves.
[0109] The information processing device 8 may also have a built-in weight estimation model WM. The user can register character strings in the user dictionary UD at any time. The operation of registering characters in the user dictionary UD by the information processing device 8 may be the same as at least one of the operations described in the second embodiment to the fourth embodiment. [8-3: Technical Effects of the Information Processing Device 8]
[0110] The information processing device 8 in the eighth embodiment performs speech recognition using a user dictionary UD tailored to the user in addition to the general-purpose dictionary GD, and therefore can provide the user with speech recognition tailored to the user. Furthermore, since the information processing device 8 uses the user dictionary UD, it can appropriately recognize character strings with a large number of characters, including katakana words, etc. Furthermore, since the information processing device 8 uses the user dictionary UD, it can accurately recognize a collocation including multiple words as a collocation. [9: Note]
[0111] The following supplementary notes are further disclosed with respect to the above-described embodiments. [Supplementary Note 1] An information processing device comprising: extraction means for extracting features related to speech recognition from a character string; and estimation means for estimating a weight corresponding to the character string based on the features, wherein the weight is a quantity related to the speech recognition. [Supplementary Note 2] The information processing device according to Supplementary Note 1, wherein the estimation means, when receiving a character string to be registered in a user dictionary, estimates a weight corresponding to the character string; and further comprising registration means for registering the character string in the user dictionary in association with the weight estimated by the estimation means, wherein the user dictionary is a dictionary corresponding to a predetermined user who uses the speech recognition, and is used for the speech recognition. [Supplementary Note 3] The information processing device according to Supplementary Note 1 or 2, wherein the extraction means extracts the features depending on a degree of learning of the speech recognition for the character string. [Supplementary Note 4] The information processing device according to Supplementary Note 1 or 2, wherein the extraction means extracts the features depending on the pronunciation of the character string. [Supplementary Note 5] The information processing device according to Supplementary Note 1 or 2, further comprising: a determination means for determining a weight of a correct answer based on an estimated character string estimated from a correct answer speech by a speech recognition mechanism that performs speech recognition to estimate a character string corresponding to the speech from the speech, and the correct answer character string; and a generation means for generating an estimation model to be used by the estimation means by performing learning based on the weight estimated by the estimation means based on the feature extracted from the correct answer character string and the weight of the correct answer. [Supplementary Note 6] The information processing device according to Supplementary Note 5, wherein new speech recognition means for performing the predetermined speech recognition is generated by performing machine learning using speech recognition training data including a predetermined character string, wherein the determination means determines a new weight of the correct answer based on an estimated character string estimated by the new speech recognition means from a correct answer speech corresponding to a correct answer character string and the correct answer character string, wherein the generation means generates a new estimation model to be used by the estimation means by performing learning based on the weight estimated by the estimation means based on the feature extracted from the correct answer character string and the new correct answer weight, wherein the extraction means extracts feature values related to the predetermined speech recognition, and wherein the estimation means estimates the weight using the new estimation model.[Supplementary Note 7] The information processing device according to Supplementary Note 2, further comprising a user interface used by the predetermined user to edit the user dictionary. [Supplementary Note 8] The information processing device according to Supplementary Note 2, further comprising a voice recognition means for estimating a character string corresponding to a voice from the voice using a general-purpose dictionary for users who use the voice recognition, including the predetermined user, and the user dictionary. [Supplementary Note 9] An information processing method for extracting features related to voice recognition from a character string, and estimating a weight corresponding to the character string based on the features. [Supplementary Note 10] A recording medium having recorded thereon a computer program for causing a computer to execute an information processing method for extracting features related to voice recognition from a character string, and estimating a weight corresponding to the character string based on the features. [Supplementary Note 11] The information processing device according to Supplementary Note 7, wherein, when the predetermined user inputs a weight corresponding to the character string via the user interface, the estimation means estimates a weight corresponding to the character string based on the features and the weight input via the user interface. [Supplementary Note 12] The information processing device according to Supplementary Note 3, wherein the extraction means extracts the feature amounts using a mechanism generated using speech recognition training data used in training the speech recognition. [Supplementary Note 13] The information processing device according to Supplementary Note 4, wherein the extraction means extracts the feature amounts using a pronunciation dictionary that associates notations of the character strings with pronunciations of the character strings. [Supplementary Note 14] An information processing device comprising: means for extracting feature amounts related to speech recognition from a character string; and means for adjusting accuracy of the speech recognition for the character string based on the feature amounts.
[0112] At least some of the components of each of the above-described embodiments can be appropriately combined with at least some of the other components of each of the above-described embodiments. Some of the components of each of the above-described embodiments may not be used.
[0113] This disclosure is not limited to the above-described embodiments. This disclosure may be modified as appropriate within the scope of the claims and the technical idea that can be read from the entire specification. Information processing devices, information processing methods, and recording media incorporating such modifications are also included in the technical idea of this disclosure. Furthermore, to the extent permitted by law, all publications and papers described in this specification are incorporated herein by reference.
[0114] To the extent permitted by law, this application claims priority to Japanese Patent Application No. 2023-035697, filed March 8, 2023, the disclosure of which is incorporated herein in its entirety by reference.
[0115] 1, 2, 3, 4, 5, 6, 7, 8 Information processing device 11, 211, 311, 411, 511, 611, 811 Extraction unit 12, 212, 312, 412, 512, 612, 712, 812 Estimation unit SE Speech recognition engine SM Speech recognition model GD General-purpose dictionary UD User dictionary FE Feature extraction mechanism PD Pronunciation dictionary WM Weight estimation model 213, 813 Input acceptance unit 214, 614, 814 Registration unit UIF User dictionary interface 515, 615 Weight determination unit 5151 Speech synthesis unit 5152 Speech recognition unit 5153 Weight determination unit 516 Generation unit 616 Weight estimation model generation unit 617 Speech recognition model generation unit 818 Speech recognition unit
Claims
1. an extraction means for extracting features related to speech recognition from a character string; an estimation means for estimating a weight corresponding to the character string based on the feature amount, The weight is a quantity related to the speech recognition. Information processing device.
2. the estimation means, when receiving a character string to be registered in a user dictionary, estimates a weight corresponding to the character string; further comprising a registration means for registering the character string and the weight estimated by the estimation means in the user dictionary in association with each other, The user dictionary is a dictionary corresponding to a predetermined user who uses the voice recognition, and is used for the voice recognition. The information processing device according to claim 1 .
3. The extraction means extracts the feature amount according to a degree of learning of the speech recognition for the character string.
3. The information processing device according to claim 1 or 2.
4. The extraction means extracts the feature amount according to the pronunciation of the character string.
3. The information processing device according to claim 1 or 2.
5. a determination means for determining a weight of a correct answer based on an estimated character string estimated from a correct answer voice by a speech recognition mechanism that performs speech recognition to estimate a character string corresponding to the speech from the speech, and the correct answer character string; and generating means for generating an estimation model to be used by the estimation means by performing learning based on the weight estimated by the estimation means based on the feature amount extracted from the correct character string and the weight of the correct answer.
3. The information processing device according to claim 1 or 2.
6. generating a new speech recognition means for performing the predetermined speech recognition by performing machine learning using speech recognition training data including a predetermined character string; the determination means determines a weight of the new correct answer based on an estimated character string estimated from the correct speech corresponding to the correct answer character string by the new speech recognition means and the correct answer character string; the generation means generates a new estimation model to be used by the estimation means by performing learning based on the weights estimated by the estimation means based on the feature amounts extracted from the correct character string and the weights of the new correct answer; the extraction means extracts the predetermined feature amount related to speech recognition, The estimation means estimates the weights using the new estimation model. The information processing device according to claim 5 .
7. The system further includes a user interface that the predetermined user uses to edit the user dictionary. The information processing device according to claim 2 .
8. A means for extracting features related to speech recognition from a character string; and adjusting the accuracy of the speech recognition for the character string based on the feature amount. Information processing device.
9. Extract features related to speech recognition from the string, Estimating a weight corresponding to the character string based on the feature amount A computer-implemented information processing method.
10. On the computer, Extract features related to speech recognition from the string, Estimating a weight corresponding to the character string based on the feature amount A computer program for executing an information processing method.