Information processing apparatus, learning apparatus, information processing method, learning method, and recording medium

CN122826564APending Publication Date: 2026-09-25MITSUBISHI ELECTRIC CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202480088385.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-02-27
Filing Date
2024-04-25
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

因此,在端到端型构造的语音识别引擎中,存在难以通过部分变更来追加登记词汇这样的问题

Benefits of technology

[0022]根据本公开,提供一种能够将由用户输入的特定符号作为输出符号优先输出的技术。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122826564A_ABST
    Figure CN122826564A_ABST
Patent Text Reader

Abstract

The processor performs an increase process for increasing the likelihood value of the candidate symbol including the registered symbol among the at least one candidate symbol using the bias module. The processor decides a temporary output symbol based on the likelihood value of each of the at least one candidate symbol after the increase process is performed. In a case where the registered symbol is included in the temporary output symbol, the processor performs a conversion process for converting the registered symbol to a specific symbol corresponding to the registered symbol with reference to the combination table, and outputs the temporary output symbol after the conversion process is performed as an output symbol.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to information processing apparatus, learning apparatus, information processing method, learning method, and recording medium. Background Technology

[0002] In recent years, information processing devices that perform conversion processing from input data to output symbols have been applied in various fields. In particular, so-called end-to-end constructions that perform this conversion processing using only a single neural network are known.

[0003] For example, an information processing device employing a speech recognition engine accepts speech data as input data. Then, the information processing device outputs the text of that speech as output symbols. Furthermore, the symbols (output symbols) can be, for example, words or articles.

[0004] However, speech recognition engines have a priority output function that pre-outputs specific symbols as output symbols. In this priority output function, for example, the user registers specific symbols representing proper nouns such as names of people or places with the speech recognition engine. The speech recognition engine then prioritizes outputting these specific symbols during speech recognition. This priority output function is also known as a vocabulary registration function, etc.

[0005] In traditional speech recognition engines, vocabulary registration was a separate function, making it easy to implement. However, in end-to-end speech recognition engines, the conversion from input data to output symbol sequences is achieved through a single neural network. Therefore, end-to-end speech recognition engines face the challenge of adding registered vocabulary through partial modifications.

[0006] To address this problem, Japanese Patent Publication No. 2022-531615 (Patent Document 1) discloses a speech recognition engine that suppresses this issue. This speech recognition engine enables an end-to-end speech recognition model to collaborate with a weighted finite-state transducer (WFST). Furthermore, through this collaboration, the speech recognition engine makes it easier for specific words to appear in the speech recognition results, thus suppressing the aforementioned problem.

[0007] Existing technical documents

[0008] Patent documents

[0009] Patent Document 1: Japanese Patent Publication No. 2022-531615 Summary of the Invention

[0010] The problem the invention aims to solve

[0011] As in Patent Document 1, when a structure that easily outputs specific symbols by combining an end-to-end model and WFST is applied to a speech recognition engine, sometimes symbols not included in the learning data of the end-to-end model are not output as output symbols. In the technology described in Patent Document 1, a substitution process is performed during the learning process of the speech recognition model. The substitution process randomly substitutes a symbol with the same pronunciation as any symbol in the learning data but different from that symbol. Through this substitution process, new symbols are generated as new learning data. Thus, the technology described in Patent Document 1 discloses a method to suppress the aforementioned problem by diversifying the learning data.

[0012] However, in the aforementioned substitution process, it is difficult for users to pre-collect the words to be generated. Therefore, even if a specific symbol is registered, it is often not generated during the substitution process. In this case, the problem may arise that the specific symbol is difficult to output as an output symbol. That is, in existing information processing devices, it may be difficult to prioritize the output of specific symbols input by the user as output symbols.

[0013] This disclosure was made to solve the aforementioned problems, and its purpose is to provide a technique that enables a specific symbol input by the user to be output as the output symbol with priority.

[0014] Methods for solving problems

[0015] The information processing apparatus of this disclosure converts input data into output symbols and outputs them. The information processing apparatus includes an interface for accepting user input of input data, a memory for storing a learned model, and a processor. The processor extracts feature quantities from the input data. By applying the feature quantities to the learned model, the processor derives at least one candidate symbol and a likelihood value representing the likelihood of each of the at least one candidate symbol. The memory stores a combination table and a bias module. The combination table is a table representing combinations of a specific symbol specified by the user and a registered symbol corresponding to that specific symbol. The bias module is a module that increases the likelihood value of a candidate symbol containing a registered symbol. The processor also performs an increase process using the bias module to increase the likelihood value of the candidate symbol containing the registered symbol among the at least one candidate symbols. Based on the likelihood values ​​of each of the at least one candidate symbol after the increase process, the processor determines a temporary output symbol. Furthermore, if the temporary output symbol contains a registered symbol, the processor performs a conversion process by referring to the combination table to convert the registered symbol into a specific symbol corresponding to that registered symbol, and outputs the temporary output symbol after the conversion process as the output symbol.

[0016] The learning apparatus disclosed herein is a learning apparatus for updating a model. The learning apparatus includes an interface for acquiring learning data, which is a combination of learning input data and a first learning output symbol, and a computing unit. The computing unit extracts feature quantities from the learning input data. The computing unit obtains output symbols by applying the feature quantities to the model. The computing unit generates a second learning output symbol by performing preprocessing on the first learning output symbol. The computing unit updates the model in a manner that reduces the error between the first and second learning output symbols. The preprocessing is the process of generating a first symbol contained in the first learning output symbol and a second symbol representing that first symbol using a user-specified expression.

[0017] The information processing method disclosed herein is an information processing method that converts input data into output symbols and outputs them. A bias module is a module that increases the likelihood value of a candidate symbol containing a registered symbol. A combination table is a table representing combinations of a specific symbol specified by a user and a registered symbol corresponding to that specific symbol. The information processing method includes extracting feature quantities from the input data. The information processing method includes applying the feature quantities to a learned model to derive at least one candidate symbol and a likelihood value representing the likelihood of each of the at least one candidate symbol. The information processing method includes performing an increase process using the bias module to increase the likelihood value of at least one candidate symbol containing a registered symbol corresponding to a specific symbol specified by the user. The information processing method includes performing an increase process using the bias module to increase the likelihood value of at least one candidate symbol containing a registered symbol. The information processing method includes determining a temporary output symbol based on the likelihood values ​​of each of the at least one candidate symbol after the increase process. The information processing method includes performing a conversion process, referring to the combination table, to convert the registered symbol into a specific symbol corresponding to that registered symbol when the temporary output symbol contains a registered symbol, and outputting the temporary output symbol after the conversion process as the output symbol.

[0018] The learning method disclosed herein is a learning method for updating a model. The learning method includes acquiring learning data as a combination of learning input data and a first learning output symbol. The learning method includes extracting feature quantities from the learning input data. The learning method includes obtaining output symbols by applying the feature quantities to the model. The learning method includes generating a second learning output symbol by performing preprocessing on the first learning output symbol. The learning method includes updating the model in a manner that reduces the error between the first and second learning output symbols. The preprocessing is the process of generating a first symbol contained in the first learning output symbol and a second symbol representing that first symbol using a user-specified expression.

[0019] The recording medium disclosed herein is a non-transitory recording medium storing a program for instructing a computer to convert input data into output symbols and output them. A bias module is a module that increases the likelihood value of candidate symbols containing registered symbols. A combination table is a table representing combinations of a specific symbol specified by a user and the registered symbols corresponding to that specific symbol. The program instructs the computer to extract features from the input data. The program instructs the computer to derive at least one candidate symbol and a likelihood value representing the likelihood of each of the at least one candidate symbol by applying the features to a learned model. The program instructs the computer to use the bias module to increase the likelihood value of at least one candidate symbol that contains a registered symbol corresponding to a specific symbol specified by the user. The program instructs the computer to use the bias module to increase the likelihood value of at least one candidate symbol that contains a registered symbol. The program instructs the computer to determine a temporary output symbol based on the likelihood values ​​of each of the at least one candidate symbol after the increased processing. The program instructs the computer to perform a conversion process when the temporary output symbol contains a registered symbol, by referring to a combination table to convert the registered symbol into a specific symbol corresponding to the registered symbol, and then outputs the temporary output symbol after the conversion process as the output symbol.

[0020] The recording medium disclosed herein is a non-transitory recording medium storing a program for updating a model in a computer. The program instructs the computer to execute learning data, which is a combination of learning input data and a first learning output symbol. The program instructs the computer to execute feature extraction from the learning input data. The program instructs the computer to execute obtaining output symbols by applying the feature values ​​to the model. The program instructs the computer to execute generating a second learning output symbol by performing preprocessing on the first learning output symbol. The program instructs the computer to execute updating the model in a manner that reduces the error between the first and second learning output symbols. The preprocessing is the process of generating a first symbol contained in the first learning output symbol and a second symbol representing that first symbol in a user-specified expression.

[0021] Invention Effects

[0022] According to this disclosure, a technique is provided that prioritizes the output of specific symbols input by the user as output symbols. Attached Figure Description

[0023] Figure 1 This is a diagram illustrating the general outline of the information processing apparatus of this disclosure.

[0024] Figure 2 It is a block diagram representing the hardware structure of an information processing device.

[0025] Figure 3 This is a functional block diagram of an information processing system.

[0026] Figure 4 This is a functional block diagram of an information processing device.

[0027] Figure 5 This is a diagram representing an example of a specific symbol table.

[0028] Figure 6 This is a diagram representing an example of a combination table.

[0029] Figure 7 This is a diagram representing an example of a bias table.

[0030] Figure 8 This is a flowchart illustrating the main processes of an information processing device.

[0031] Figure 9 This is a functional block diagram of the learning device.

[0032] Figure 10 This is a diagram illustrating an example of word modification processing.

[0033] Figure 11 This is a diagram illustrating an example of preprocessing.

[0034] Figure 12 This is a flowchart of a learning device.

[0035] Figure 13 This is an example of a specific symbol table expressed in Japanese.

[0036] Figure 14 This is an example of a combination table expressed in Japanese.

[0037] Figure 15 This is an example of an offset table expressed in Japanese.

[0038] Figure 16 This is a diagram representing other application examples of input data.

[0039] Figure 17 This is a diagram representing an example of a probability symbol.

[0040] Figure 18 This is a flowchart of the method for determining the registration symbol. Detailed Implementation

[0041] Hereinafter, embodiments of the present invention will be described in detail with reference to the accompanying drawings. Several embodiments will be described below, but from the outset of this application, appropriate combinations of structures to be described in each embodiment were predetermined. Furthermore, identical or equivalent parts in the drawings will be labeled with the same reference numerals, and their descriptions will not be repeated.

[0042] Implementation method 1.

[0043] [Summary of this disclosure]

[0044] Hereinafter, the term "symbol" will be used. A symbol is a unit of information output from the information processing device 100. A symbol may represent at least one character. Characters may be, for example, hiragana, katakana, kanji, and letters. For example, a symbol may be a single character (e.g., "A") or multiple characters, such as "America". Furthermore, when a symbol is defined as a single character, multiple characters are also referred to as a "symbol sequence". Additionally, a symbol may also represent information indicating events such as the start and end of the symbol.

[0045] In the following disclosure, information representing a single character and information consisting of multiple characters will be collectively referred to as a "symbol". That is, "A" and "America" ​​are both "symbols".

[0046] Figure 1 This is a diagram illustrating the general outline of the information processing apparatus 100 of this disclosure. Figure 1 In the example shown, an information processing device 100, a microphone 2 (voice input device), and a display 3 (display device) are shown. The microphone 2 and the display 3 are connected to the information processing device 100.

[0047] exist Figure 1 In the example, user A utters the voice "i triple e". "i triple e" represents "IEEE". Furthermore, "IEEE" is a symbol (word) representing the Institute of Electrical and Electronics Engineers (IEEE), an example of a proper noun. The information processing device 100 receives the voice data input of the voice "i triple e", performs voice recognition processing, and outputs "IEEE" as text data. The display 3 displays "IEEE" as the text data.

[0048] In addition, user A registers "IEEE" as a specific symbol with information processing device 100. When user A issues the voice "ⅰtriple e", user A wants to prioritize outputting "IEEE" as the output symbol.

[0049] In addition, Figure 1 In the example, the user expects the output "IEEE" to be included in the output symbols, but in a conventional speech recognition device, text such as "i triple e" can be displayed directly as speech on the display 3. Furthermore, the information processing device 100 of this embodiment adopts an end-to-end configuration.

[0050] [Hardware Structure of Information Processing Device]

[0051] Figure 2This is a block diagram illustrating the hardware structure of the information processing apparatus 100 according to Embodiment 1. The information processing apparatus 100 can be implemented by, for example, a general-purpose computer or a special-purpose computer.

[0052] like Figure 2 As shown, the information processing device 100 includes an arithmetic unit 11, a storage unit 12, a voice interface 13, a communication unit 14, a display interface 15, an input device interface 16, and a reading unit 17 as its main hardware components.

[0053] The arithmetic unit 11 is an arithmetic entity (processor) that performs various processes by executing various programs, and is an example of a computer. The arithmetic unit 11 corresponds to the "processor" of this disclosure. The arithmetic unit 11 may be composed of, for example, a CPU (Central Processing Unit), an FPGA (Field-Programmable Gate Array), or a GPU (Graphics Processing Unit). Furthermore, the arithmetic unit 11 may be composed of at least one of a CPU, FPGA, and GPU, or may be composed of all of the following: a CPU and an FPGA, an FPGA and a GPU, a CPU and a GPU, or a CPU, an FPGA, and a GPU. The arithmetic unit 11 is also referred to as "at least one processor." Additionally, the arithmetic unit 11 may be composed of processing circuitry. Furthermore, the arithmetic unit 11 may be composed of a single chip or multiple chips. Moreover, all or part of the functions of the arithmetic unit 11 may be provided in a server device (e.g., a cloud-type server device) not shown.

[0054] The storage unit 12 includes volatile storage areas (e.g., working areas) for temporarily storing program code, working memory, etc., when the arithmetic unit 11 executes any program. For example, the storage unit 12 is composed of volatile storage devices such as DRAM (Dynamic Random Access Memory) or SRAM (Static Random Access Memory). Furthermore, the storage unit 12 includes non-volatile storage areas. For example, the storage unit 12 is composed of non-volatile storage devices such as ROM (Read Only Memory), hard disk, or SSD (Solid State Drive).

[0055] Furthermore, in this embodiment, an example is shown where volatile and non-volatile storage areas are contained in the same storage unit 12. However, volatile and non-volatile storage areas may also be contained in different storage units. For example, the arithmetic unit 11 may contain volatile storage areas, and the storage unit 12 may contain non-volatile storage areas. The information processing apparatus 100 may also include a microcomputer that includes the arithmetic unit 11 and the storage unit 12.

[0056] The storage unit 12 stores the processing program 301, the learned model D3, the decision model 305, the combination table D7, and the bias table D8. The processing program 301 includes a generation program and a learning program. The information processing program describes the information processing by which the arithmetic unit 11 generates output text based on the voice data (input data) obtained from the microphone 2 and the learned model D3. The learning program describes the learning process of updating (learning) the learned model D3.

[0057] The learned model D3 includes a neural network 303 and parameters 304 used by the neural network 303. The learned model D3 performs machine learning (update processing) using multiple learning data as a combination of learning input data and learning output symbols.

[0058] The learned model D3 includes well-known constructions such as DNN (deep neural network), CNN (convolutional neural network), LSTM (long short-term memory), Transformer, and Conformer. The determination of model 305, combination table D7, and bias table D8 will be discussed later.

[0059] Voice interface 13 is used to connect microphone 2. Communication unit 14 communicates with designated external devices. Display interface 15 is used to connect display 3, enabling data input and output between information processing device 100 and display 3.

[0060] The input device interface 16 is an interface for connecting the input device 4 (e.g., keyboard and mouse) to realize the input and output of data between the information processing device 100 and the input device.

[0061] The reading unit 17 reads various data stored in the removable disk 20, which serves as a storage medium. The removable disk 20 stores the processing program 301 of this disclosure (at least one of an information processing program and a learning program). Furthermore, the removable disk 20 storing the processing program 301 is sold, etc. Additionally, the reading unit 17 can also obtain the processing program 301 from the removable disk 20. The removable disk 20 corresponds to an example of the "non-transitory recording medium" of this disclosure.

[0062] [Functional block diagram of an information processing system]

[0063] Figure 3 This is a functional block diagram of the information processing system 30. The information processing system 30 includes an information processing device 100, a learning device 200, a learning data storage device 300, and a learned model storage device 400. The information processing device 100 outputs an output symbol D4 based on input data D1 and a specific symbol table D2, using the learned model D3.

[0064] The learning device 200 learns the learned model D3 used in the information processing device 100 (updates parameters 304). The learning device 200 uses the learning data E stored in the learning data storage device 300 to learn the model. The learned model E3, as the learned model, is output to the learned model storage device 400.

[0065] [Functional block diagram of an information processing device]

[0066] Figure 4 This is a functional block diagram of the information processing device 100. The information processing device 100 includes an extraction unit 110, a decision unit 120, a generation unit 130, a derivation unit 140, and a conversion unit 150.

[0067] The information processing device 100 of this embodiment functions as a speech recognition device. Input data D1 and a specific symbol table D2 are input to the information processing device 100. In this embodiment, the input data D1 is speech data. The input data D1 is, for example, a digital signal obtained by converting the speech input to the microphone 2 using an A / D (analog / digital) converter or similar method. Furthermore, the output symbol D4 is text data representing the speech data. Additionally, for example, before the input data D1 is input to the information processing device 100, the specific symbol table D2 is input by the user from the input device 4.

[0068] Input data D1 is input to extraction unit 110. Extraction unit 110 extracts feature quantity D5 from input data D1 in a form suitable for input to the learned model D3. The method for extracting feature quantity D5 includes, for example, a step of segmenting the speech signal as input data D1 at regular intervals. Furthermore, the extraction method includes processing to extract a known log-Mel filter bank from the segmented signal or processing to extract the Mel frequency cepstrum. Other methods may also be used for feature extraction.

[0069] Feature D5 is output to derivation unit 140. Derivation unit 140 derives at least one candidate symbol and its respective likelihood value by applying feature D5 to the learned model D3. The likelihood value is a parameter representing the likelihood of a candidate symbol, also known as a score. For example, negative log-likelihood (NLL) is used as the score. In the case of a negative log-likelihood score, a smaller score indicates a higher likelihood value. In this embodiment, an example where a smaller score indicates a higher likelihood value is applied. Furthermore, as a variation, a larger score could also be configured to indicate a higher likelihood value.

[0070] Then, the derivation unit 140 uses the bias table D8 described later to perform an increase process (a process that decreases the score) that increases the likelihood value. Then, the derivation unit 140 outputs the candidate deviation with the largest likelihood value after the increase process as a temporary output symbol D9.

[0071] Next, the specific symbol table D2 will be explained. The specific symbol table D2 is a table that stores at least one specific symbol. A specific symbol is, for example, a symbol that tends not to be included in the output symbol D4 when the information processing device 100 performs speech recognition processing on the input text without using the bias table D8 (described later). Based on this tendency, a specific symbol is typically a symbol that the user wishes to preferentially include in the output symbol D4. Specific symbols are proper nouns or characters with special meanings, etc. Furthermore, the specific symbol table D2 contains information pre-input to the information processing device 100 by the user A or others before the input data D1 is input to the information processing device 100.

[0072] The registration symbols described later are symbols that correspond to specific symbols. For example, a registration symbol is a symbol with the same meaning as a specific symbol but represented by a different method of expression (expression category). Regarding the pronunciation of a specific symbol as a word, registration symbols can include hiragana, katakana, romaji, and symbols that differ in spelling but have the same pronunciation. Furthermore, registration symbols can include symbols that express a specific symbol using the International Phonetic Alphabet (IPA), etc. Thus, the pronunciation of a registration symbol is the same as that of a specific symbol, but it can be expressed using different symbols. Moreover, registration symbols and specific symbols can also be the same.

[0073] Figure 5 This is an example of a specific symbol table D2. Specific symbols are, for example, proper nouns such as personal names, place names, or the organization name "IEEE" as mentioned above. In the specific symbol table, a specific symbol (IEEE) corresponds to the phonetic symbol of that specific symbol as a substitute symbol. There may also be no substitute symbol.

[0074] The decision unit 120 determines the registered symbols from the input specific symbol table D2. Registered symbols are those that are preferentially output by the derivation unit 140 in place of specific symbols. Furthermore, "preferred output symbol" refers to a symbol whose likelihood value is high compared to other candidate symbols (described later). Thus, in this embodiment, a high likelihood value for the registered symbol is preferred.

[0075] Next, the decision unit 120 extracts a specific symbol from the specific symbol table. Then, the decision unit 120 converts the specific symbol into a registered symbol.

[0076] Furthermore, the method for converting a specific symbol to a registered symbol can employ a method that determines the registered symbol based on supplementary information (not shown) contained in the specific symbol table D2. Alternatively, this conversion method can be implemented through a prescribed operation. Furthermore, this conversion method can also be implemented, for example, by having the user specify the expression category of the registered symbol. Additionally, alternative symbols can also be determined as registered symbols.

[0077] Then, the decision unit 120 generates a combination table D7 by combining a specific symbol and a registration symbol. The combination table D7 is a table that represents the combination of a specific symbol input by the user and the registration symbol corresponding to that specific symbol.

[0078] Figure 6 This is a diagram illustrating an example of combination table D7. Combination table D7 is a table that maps "IEEE" as a specific symbol to the phonetic symbol (registration symbol) of that specific symbol. Combination table D7 is stored in storage unit 12 ( Figure 2 Furthermore, combination table D7 is output to conversion unit 150, which will be described later.

[0079] Additionally, the decision unit 120 extracts the registration symbol D6 from the combination table D7 and outputs the registration symbol D6 to the generation unit 130. Registration symbol D6 is... Figure 6 The symbol enclosed in a dashed box.

[0080] The generation unit 130 generates a bias table D8 by determining the bias value corresponding to the registration symbol D6 contained in the combination table D7. The bias table D8 corresponds to an example of the "bias module" of this disclosure. The bias table D8 is a table that defines the correspondence between the registration symbol and the bias value that increases the likelihood value. The bias table D8 generated by the generation unit 130 is stored in the storage unit 12.

[0081] Figure 7 This is an example of a bias table. Figure 7 In the example, "-3.5" is used as an offset value and corresponds to the IEEE phoneme symbol (registration symbol D6). Additionally, the offset value also corresponds to other registration symbols.

[0082] The learned model D3 can be learned under the following premise: when the number of registered symbols in the learning data E of the learned model D3 is small, the likelihood value of the symbol containing the registered symbol is smaller than the likelihood value of the symbol containing other symbols different from the registered symbol.

[0083] Here, "i triple e" is segmented into words like "i", "triple", and "e", which are more frequently contained in the learning data E. Therefore, compared to IEEE phoneme symbols, the learning data E contains more "i triple e". Therefore, as... Figure 1 As shown, when “i triple e” is the input data D1, there is a tendency (premise) for “i triple e” to have a score smaller than that of the IEEE phoneme symbol (“i triple e” has a higher likelihood value than the phoneme symbol). In addition, “i triple e” is also referred to as “the general word”.

[0084] Under this premise, the bias value is set such that the phoneme symbols of the score derived by the derivation unit 140 are less than "i triple e". That is, as a guideline for determining the bias value, the generation unit 130 determines a bias value that significantly reduces the score (increases the likelihood value) for symbols that are few in number in the learning data E.

[0085] For example, the generation unit 130 uses the decision model 305 (see reference). Figure 2 The bias value is determined. Model 305 is, for example, a model made using the learning symbols DT2, which includes a language model. The language model is, for example, an N-gram language model. The language model outputs a large negative log-likelihood for symbols with small quantities in the learning symbols DT2, and a small negative log-likelihood for symbols with large quantities in the learning symbols DT2. Model 305 can be represented by equation (1).

[0086] Bias value = ρ - λ × (negative log-likelihood from the language model) (1)

[0087] Here, ρ is a constant, and λ is a positive constant. Thus, by using a decision model 305 based on equation (1), the generation unit 130 learns that the fewer the number of input registration symbols included in the learning registration symbols, the higher the likelihood value of the candidate symbol containing that registration symbol. For example, when the number of input registration symbols included in the learning registration symbols is small, the bias value is -5; when the number of input registration symbols included in the learning registration symbols is large, the bias value is -1. Furthermore, the input registration symbols are the registration symbols input to the decision model 305.

[0088] Furthermore, the bias table can also be generated in the WFST form described above. Alternatively, a machine learning model can be used instead of a language model.

[0089] Next, the processing of the derivation unit 140 using bias table D8 will be described. As described above, the derivation unit 140 derives at least one candidate symbol and its respective likelihood value by applying feature quantity D5 to the learned model D3. In the absence of bias table D8, the derivation unit 140 determines the candidate symbol with the smallest score (highest likelihood value) from the at least one candidate symbol and outputs this candidate symbol as a temporary output symbol D9.

[0090] On the other hand, in the presence of bias table D8, the derivation unit 140 adds the score of the candidate symbol containing the registered symbol to the bias value corresponding to the registered symbol in bias table D8.

[0091] For example, let's illustrate the case where the derivation unit 140 identifies the following first and second candidate symbols as candidate symbols. Let the first candidate symbol be "i triple e", and its score be "0.0". Let the second candidate symbol be "IEEE phoneme symbol", and its score be "1.0". In this example, the score of the first candidate symbol is smaller than the score of the second candidate symbol (higher likelihood value).

[0092] In this example, without bias table D8, the derivation unit determines the first candidate symbol as the temporary output symbol.

[0093] On the other hand, in this example, in the existence Figure 7 In the case of bias table D8, the second candidate symbol is registered as a registered symbol. Therefore, the derivation unit 140 performs an increase process to increase the likelihood value of the candidate symbol (the second candidate symbol) based on the bias value (-3.5) corresponding to the registered symbol. In this example, the increase process is performed by performing a process of adding the bias value to the score to calculate a score such as -2.5.

[0094] Thus, the derivation unit 140 performs an increase process based on the bias value corresponding to the registered symbol to increase the likelihood value of at least one candidate symbol that contains the registered symbol.

[0095] Then, based on the likelihood values ​​of at least one candidate symbol after the augmentation process, a temporary output symbol D9 is determined. In the example above, the score of the first candidate symbol after the augmentation process is "0.0", and the score of the second candidate symbol is "-2.5". Therefore, the derivation unit 140 determines the second candidate symbol as the temporary output symbol D9. This temporary output symbol D9 is input to the conversion unit 150.

[0096] When the temporary output symbol D9 contains the registration symbol, the conversion unit 150 refers to the combination table D7 ( Figure 6 Then, the conversion unit 150 performs a conversion process to convert the registered symbol into a specific symbol corresponding to the registered symbol. Then, the conversion unit 150 outputs the temporary output symbol D9 after the conversion process as the output symbol D4.

[0097] In the example above, the temporary output symbol D9 contains "IEEE phonemes". Therefore, the conversion unit 150 adds "IEEE phonemes" (registered symbols) to the combination table D7 ( Figure 6 The process is converted to "IEEE" (a specific symbol) corresponding to the phoneme symbol of "IEEE". Then, "IEEE" as the converted symbol is output as the output symbol D4.

[0098] Next, the significance of the decision unit 120 will be explained. Consider the scenario where the decision unit 120 is not set up, and the specific symbols in the specific symbol table D2 are directly treated as the registered symbols D6, generating the bias table D8. That is, in the example above, the generation unit 130 can also generate the bias table D8 using the specific symbol "IEEE" instead of the registered symbol "IEEE phonemes". This bias table D8 is based on... Figure 7 When illustrating the example, the bias value (-3.5) corresponds to "IEEE".

[0099] Here, "IEEE" is a proper noun; therefore, in the learning data E used for learning the already learned model D3, the amount of data containing "IEEE" is often zero or very small. In this case, the derivation unit 140 makes the score of the candidate symbol "IEEE" much larger than the scores of "i triple e" and "IEEE phoneme symbol".

[0100] In this case, it is assumed that even if the derivation unit 140 adds a bias value to the score of the candidate symbol "IEEE" using the bias table D8, it will not be less than the score of "i triple e" and the score of "IEEE". In this case, the user-expected "IEEE" will not be output as the output symbol D4.

[0101] The candidate symbol "IEEE phoneme symbol" is represented differently from the specific symbol "IEEE", but the individual phonemes of the phoneme symbol are more extensively included in the learning data. Therefore, the score of the candidate symbol "IEEE phoneme symbol" is lower than the score of the specific symbol "IEEE".

[0102] As described above, compared to generating bias table D8 using the specific symbol "IEEE", the registration of the specific symbol functions more effectively when generating bias table D8 using the registered symbol "IEEE phoneme symbols". That is, the decision unit 120 has the function of converting the specific symbol into a symbol with a higher likelihood value than the specific symbol (in this example, "IEEE phoneme symbols").

[0103] [flow chart]

[0104] Figure 8 This is a flowchart illustrating the main processes of the information processing device 100. First, in step S12, the information processing device 100 extracts feature quantity D5 from the input data D1. Next, in step S14, the information processing device 100 derives at least one candidate symbol and the likelihood value of each candidate symbol.

[0105] Next, in step S16, the information processing device 100 performs an increase process that increases the likelihood value of the candidate symbol containing the registration symbol D6 based on the bias value. Next, in step S18, the information processing device 100 determines a temporary output symbol D9 based on the likelihood value after the increase process. Next, in step S20, the information processing device 100 converts the registration symbol contained in the temporary output symbol D9 into a specific symbol corresponding to that registration symbol in the combination table D7, and outputs it as output symbol D4.

[0106] [Learning Device]

[0107] Next, the learning device 200 will be described (refer to...). Figure 3 When a predetermined start condition is met, the learning device 200 learns model E5. Model E5 corresponds to the previously learned model D3 before the start condition was met. Furthermore, the start condition is, for example, a condition that is met by a predetermined start operation performed by the user of the learning device 200.

[0108] Figure 9 This is a functional block diagram of the learning device 200. The learning device 200 includes an acquisition unit 210, an extraction unit 220, a processing unit 230, a model storage unit 235, a preprocessing unit 240, an update unit 250, and an output unit 260.

[0109] Multiple learning data E are stored in the learning data storage device 300. The acquisition unit 210 acquires the multiple learning data E from the learning data storage device 300. The learning data E (supervisory data) consists of learning input data E1 for learning model E5 and a first learning output symbol E2 corresponding to the learning input data E1. For example, the learning input data E1 is speech data, and the first learning output symbol E2 is text data written with the speech represented by the speech data. The learning input data E1 is output to the extraction unit 220, and the first learning output symbol E2 is output to the preprocessing unit 240. In addition, in the learned model D3, the more symbols included in the learning output symbol, the higher the likelihood value of the candidate symbol containing that symbol.

[0110] Extraction unit 220 extracts feature quantities E4 from the learning input data E1 in a form suitable for inputting into the model. The specific processing method for extracting feature quantities E4 from the learning input data E1 is the same as that in extraction unit 110 of information processing device 100. Figure 4 The process of extracting feature quantity D5 from input data D1 is the same as that in the previous process.

[0111] Model storage unit 235 stores model E5. After reading model E5 from model storage unit 235, processing unit 230 applies model E5 to feature quantity E4 extracted by extraction unit 220 to obtain output symbol E6. For example, model E5 outputs a score (likelihood value) indicating which symbol is reasonable to output in each frame of feature quantity E4. Alternatively, model E5 calculates the score when outputting a symbol candidate given feature quantity E4, already output symbol information, and candidate information for the next output symbol.

[0112] The preprocessing unit 240 generates the second learning output symbol E7 by applying prescribed preprocessing to the first learning output symbol E2. The preprocessing includes, for example, the following modification process: applying lexical analysis to the written text, and after segmenting it according to each word, randomly modifying a portion of the words (object words) into words represented by expressions different from those words.

[0113] For example, the modification process includes changing object words that are not in hiragana to words expressed in hiragana. Additionally, the modification process includes changing object words that are not in katakana to words expressed in katakana. The modification process includes changing object words that are not in romaji to words expressed in romaji. The modification process includes changing the object word to a word marked by lengthening its vowel. Figure 10 This is an example of how words are processed by changing to long vowel markers.

[0114] The preprocessing of the preprocessing unit 240 will be further explained. As described above, in the information processing apparatus 100, a high likelihood value for the registered symbol is preferred. Here, as described above, a registered symbol is a symbol that has the same meaning as a specific symbol but is expressed in a different way than the specific symbol (see [reference]). Figure 6 ).

[0115] Additionally, the first learning output symbol E2 is, for example, an English text that uses letters. An English text that serves as the first learning output symbol E2 is, for example, a text containing "i triple e" such as "I attend i triple e". Hereinafter, as mentioned above, "i triple e" and similar phrases contained in English texts are "general words".

[0116] That is, when the learning device 200 directly uses the first learning output symbol E2 to learn model E5, model E5 is learned by assigning high likelihood values ​​to general words such as "i triple e" and low likelihood values ​​to phoneme symbols.

[0117] When learning Model E5 in this way, the score of the output symbol for a general word (the first score) is sometimes excessively lower than the score of the output symbol for a phoneme (the second score). In this case, if the phoneme (e.g., IEEE phonemes) is a registered symbol, even if the score of the registered symbol (the second score) decreases due to the bias value, the first score cannot be lower than the second score, and the following first problem may still occur: the output symbol for a general word is output. Furthermore, if λ in the bias formula is set to a large value to compensate for the large difference in scores, the following second problem may occur: the registered symbol is incorrectly output in irrelevant scenarios.

[0118] Therefore, assuming the first and second problems, the preprocessing of the preprocessing unit 240 involves applying lexical analysis to the text, segmenting it according to each word, selecting a subset of words, and converting them into different tags. Lexical analysis is an example of the "prescribed processing" of this disclosure.

[0119] Figure 11 This is a diagram used to illustrate the preprocessing. Figure 11 In the example, the text is "I attend i triple e". Furthermore, the preprocessing unit 240 segments the text into "I", "attend", and "i triple e" by applying morpheme analysis. These segmented symbols are also referred to as unit symbols of this disclosure or simply words.

[0120] Then, the preprocessing unit 240 selects "i triple e" as the first symbol. The preprocessing unit 240 then converts "i triple e" into the phoneme symbol "IEEE" with a different expression. The phoneme symbol "IEEE" is an example of the second symbol. Thus, the learning device 200 can increase the learning data (second symbol) of the phoneme symbols. Therefore, the learning device 200 can increase the likelihood value of candidate symbols containing the learning data (second symbol) of the phoneme symbols. That is, it can appropriately reduce the score (second score) of the output symbol based on the phoneme symbols, reduce the difference between the first score and the second score, and more effectively assign bias values ​​to the registered symbols. The first learning output symbol E2 after preprocessing by the preprocessing unit 240 is called the "second learning output symbol E7".

[0121] Furthermore, the representation category of the second symbol is specified by the user. Additionally, the second learning output symbol E7 includes the first symbol and the second symbol corresponding to that first symbol. The second symbol can be represented as corresponding to the registration symbol. Furthermore, typically, the second symbol is the same as the registration symbol. Additionally, the second symbol can be contained within the registration symbol, and the registration symbol can be contained within the second symbol.

[0122] The update unit 250 calculates the error between the output symbol E6 (output symbol) output from the processing unit 230 and the second learning output symbol E7 output from the preprocessing unit 240. Then, the update unit 250 updates the model E5 in a manner that minimizes the error. The update unit 250 outputs the updated model E5 to the model storage unit 235. Here, the function used to calculate the error between the output symbol E6 and the second learning output symbol E7 is, for example, the cross-entropy error function or the CTC error function (CTC: connectionist temporal classification). Furthermore, in model updating, after calculating the gradients of each parameter (parameter 304) of the model using the known inverse error propagation method, the model E5 is updated using optimization methods such as SGD (Stochastic Gradient Descent) or the Adam method.

[0123] After learning model E5, output unit 260 outputs model E5 stored in model storage unit 235 as learned model E3 to learned model storage device 400.

[0124] [Flowchart of the learning device]

[0125] Figure 12This is a flowchart of the learning device 200. In step S52, learning data E, which is a combination of learning input data E1 and the first learning output symbol E2, is obtained. Next, in step S54, the learning device 200 extracts feature quantity E4 from the learning input data E1.

[0126] Next, in step S56, the learning device 200 obtains the output symbol E6 by applying feature quantity E4 to model E5. Then, in step S58, the learning device 200 generates the second learning output symbol E7 by performing preprocessing on the first learning output symbol E2 (see reference). Figure 11 ).

[0127] Next, in step 60, the learning device 200 updates model E5 in a manner that reduces the error between output symbol E6 and the first learning output symbol E2.

[0128] Next, the learning device 200 determines whether the termination condition is met. The termination condition may include, for example, processing a predetermined amount of learning data E. If the termination condition of processing a predetermined amount of learning data E is met (yes in step S62), the process proceeds to step S64. On the other hand, if the termination condition is not met (no in step S62), the process returns to step S52. In step S64, the learning device 200 outputs the model E5 updated in step S60 as the learned model E3.

[0129] Furthermore, at least one of the information processing device 100 and the learning device 200 can be implemented on a local computer or configured in the cloud. Alternatively, the information processing device 100 and the learning device 200 can be configured on the same computer. Alternatively, the information processing device 100 and the learning device 200 can be configured on different computers. In such a configuration, the learning device 200 can also output the learned model to the learned model storage device 400 via a network.

[0130] [summary]

[0131] (1) For example, sometimes a user wants to prioritize certain symbols as output symbols to an information processing device. In this case, in an information processing device that adopts an end-to-end model, since the conversion from input data to output symbol sequence is achieved through a single neural network, there is a problem that it is difficult to perform customization such as adding registered words through partial changes.

[0132] Furthermore, in the technology disclosed in Japanese Patent Application Publication No. 2022-531615, since random permutation processing is performed, there are many cases where a specific symbol is not generated during the permutation process. In this case, it may be difficult to output the specific symbol as an output symbol. That is, in existing information processing devices, it may be difficult to prioritize the output of a specific symbol input by the user as an output symbol.

[0133] Therefore, in this disclosure, the information processing apparatus 100 derives at least one candidate symbol and a likelihood value representing the likelihood of each of the at least one candidate symbol based on input data. Next, the information processing apparatus 100 performs an increase process based on a bias module to increase the likelihood value of the candidate symbol containing the registered symbol. This allows the registered symbol (or a candidate symbol containing the registered symbol) to be preferentially determined as a temporary output symbol. Then, if the temporary output symbol contains the registered symbol, the information processing apparatus 100 performs a conversion process, referring to a combination table, to convert the registered symbol into a specific symbol corresponding to the registered symbol, and outputs the temporary output symbol after the conversion process as an output symbol.

[0134] Therefore, the above customization can be easily performed by having the user input a specific symbol. Furthermore, by performing addition processing to prioritize the output of registered symbols and substitution processing to change registered symbols to specific symbols, the specific symbol can be prioritized as the output symbol.

[0135] Furthermore, if it is desired to prioritize the output of specific symbols as output symbols, it is considered to update the structure of the learned model D3 using learning data containing a large number of specific symbols. However, when adopting such a structure, the information processing device 100 may prioritize the output of less frequently used specific symbols such as proper nouns. In this embodiment, it is possible to avoid adopting such a structure and instead prioritize the output of specific symbols desired by the user as output symbols to the information processing device 100.

[0136] (2) In addition, as preprocessing, the learning device 200 performs preprocessing to generate a first symbol contained in the first learning output symbol and a second symbol representing the first symbol in an expression specified by the user. Figure 11 Therefore, the learning device 200 can increase the amount of learning data containing the second symbol. Therefore, the learning device 200 updates model E3 in a way that increases the likelihood value of the second symbol (the symbol corresponding to the registered symbol). Therefore, the learning device 200 can update the model in a way that increases the likelihood value of the second symbol, which represents the first symbol in an expression category specified by the user.

[0137] Implementation method 2.

[0138] In Embodiment 2, various other embodiments of the above structure are described.

[0139] [Symbols]

[0140] An example in which the symbols of the above-described Embodiment 1 are mainly in English is described. However, symbols may also be expressed in other languages. Examples of other languages include Japanese.

[0141] Figure 13 This is an example of a specific symbol table D2 expressed in Japanese. In Figure 13 the example, proper nouns in Chinese characters are shown as specific symbols. In addition, as alternative symbols, a first symbol expressed in katakana and a second symbol expressed in hiragana are shown.

[0142] Figure 14 This is an example of a combination table D7 expressed in Japanese. In Figure 14 the example, proper nouns in Chinese characters are shown as specific symbols. In addition, symbols expressed in katakana are shown as registered symbols.

[0143] Figure 15 This is an example of a bias table D8 expressed in Japanese. In Figure 15 the example, an example in which bias values are respectively defined for three registered symbols is shown.

[0144] The processing flow of the information processing apparatus 100 when employing such a structure will be described. Assume that the specific symbol input by a user is the proper noun "Choshi" that is a place name. Further, the structure in which the speech is "watashi wa choushi ni itta" will be described below. In the case of this structure, it is preferable to output a symbol such as "Watashi wa, Choshi ni itta" as an output symbol.

[0145] Assume that the derivation unit 140 determines the following first candidate symbol and second candidate symbol as candidate symbols. The first candidate symbol is "Watashi wa, Choushi e itta", and the score of this first candidate symbol is "0.0". "Choushi" is a common noun that means the state of linguistic expression.

[0146] In addition, assume that the second candidate symbol is "Watashi wa, Choushi e itta", and the score of this second candidate symbol is "1.0". In this example, the score of the first candidate symbol is smaller than the score of the second candidate symbol (the likelihood value is higher).

[0147] In this example, when there is no Figure 15 bias table D8, the derivation unit determines the first candidate symbol as a temporary output symbol.

[0148] On the other hand, in this example, when there is the Figure 15 bias table D8, as shown in Figure 14As shown, the second candidate symbol is registered as a registered symbol. Therefore, the derivation unit 140 executes an increasing process of increasing the likelihood value of the candidate symbol (the second candidate symbol) based on the bias value (-3.5) corresponding to the registered symbol. In this example, the increasing process is a process of calculating a score of -2.5 by performing a process of adding the bias value to the score.

[0149] In this way, the derivation unit 140 executes an increasing process of increasing the likelihood value of a candidate symbol including the registered symbol among at least one candidate symbol based on the bias value corresponding to the registered symbol.

[0150] Then, based on the likelihood value of each of the at least one candidate symbol after the increasing process is performed, a temporary output symbol D9 is determined. In the above example, the score of the first candidate symbol after the increasing process is "0.0", and the score of the second candidate symbol is "-2.5". Therefore, the derivation unit 140 determines the second candidate symbol as the temporary output symbol D9. The temporary output symbol D9 is input to the conversion unit 150.

[0151] When the temporary output symbol D9 includes a registered symbol, the conversion unit 150 refers to the combination table D7 ( Figure 6 ). Then, the conversion unit 150 executes a conversion process of converting the registered symbol into a specific symbol corresponding to the registered symbol. Then, the conversion unit 150 outputs the temporary output symbol D9 after the conversion process as an output symbol D4.

[0152] In the above example, the temporary output symbol D9 includes "私は、チョーシへ行った". Therefore, the conversion unit 150 refers to the combination table D7 ( Figure 6 ), and converts "私は、チョーシへ行った" into "私は、銚子へ行った" that includes the specific symbol. Then, "私は、銚子へ行った", which is the converted symbol, is output as the output symbol D4. As described above, even if the symbols are in another language, the information processing apparatus 100 can appropriately and preferentially output a specific symbol as an output symbol.

[0153] [Input data and output symbols]

[0154] In the above embodiment, an example in which the information processing apparatus 100 functions as a speech recognition apparatus is described. That is, an example where the input data is speech data and the output symbol is the text of the speech represented by the speech data is described. However, input data and output symbols may also take other forms. Figure 16 is a diagram showing this other example.

[0155] In Figure 16In the example, when the input data is speech data, the output symbols can also be summary data. Summary data is text that summarizes the speech represented by the speech data. In this structure, the information processing device functions as a speech summarizing device.

[0156] Alternatively, the input data can be any of the following: sound data (e.g., sound effects data), still image data, and moving image data. In this case, the output symbol is explanatory text. Explanatory text is text that describes an object (e.g., subtitles). The object is represented by sound based on sound data, still image based on still image data, and moving image based on moving image data. In this structure, the information processing device functions as a subtitle-generating device.

[0157] like Figure 16 As shown, in the information processing apparatus 100 of this disclosure, the input data and output symbols can be various.

[0158] [Method for determining bias value]

[0159] In the above embodiment, an example was described where the generation unit 130 uses an N-gram language model learned using the output symbols to determine the bias value. However, other methods can also be used to determine the bias value.

[0160] For example, such as Figure 15 The column of bias values ​​is shown in parentheses; multiple registration symbols can have the same bias value. Figure 15 The parenthesized examples show an instance where all bias values ​​are -2.0. Because of this structure, the same bias values ​​are used in all the incrementing processes described above, thus simplifying the incrementing process.

[0161] Alternatively, the generation unit 130 may determine that the longer the number of characters in the registration symbol, the greater the likelihood value will be as the bias value. In the above embodiment, the bias value that greatly increases the likelihood value is a negative number with a large absolute value.

[0162] Alternatively, model 305 can also be a machine learning model that takes the registration symbol as input and outputs a bias value.

[0163] Alternatively, a structure in which the user determines the bias value can be adopted. The information processing device 100 employing such a structure can also use the input device 4 (see reference 4). Figure 2 The bias value can be input using the input device 4. Additionally, with this structure, the bias value can be input (set) to a registration symbol already stored in the information processing device 100. Alternatively, the user can also use the input device 4 (see reference 4). Figure 2 Enter the information that maps the registration symbol to the bias value.

[0164] [Preprocessing]

[0165] exist Figure 11 The text describes a preprocessing step involving randomly selecting the first symbol from a set of unit symbols determined by prescribed operations (such as morpheme analysis).

[0166] However, preprocessing can be anything from simply increasing the likelihood value of candidate symbols that contain the registration symbol to other processing methods. For example, in preprocessing, unit symbols judged to be nouns or proper nouns can be selected as the first symbol from the segmented unit symbols. The reason for this is that it is assumed that the unit symbols (specific symbols) that the user wants to register usually contain many difficult-to-read nouns or proper nouns.

[0167] exist Figure 11 In the example, the unit symbol "i triple e" (a proper noun only) is selected from the first learning output symbol as the first symbol, and the proper noun is converted into the second learning output symbol.

[0168] Furthermore, the preprocessing determines the quantity contained in all the first learning output symbols E2 for each of the multiple unit symbols. Here, the first learning output symbols E2 are at least one of the symbols used in the learning device 200 and the symbols used before use in the learning device 200. The reason for this is that the unit symbols (specific symbols) that the user wants to register are usually those with a smaller quantity contained in the first learning output symbols E2.

[0169] exist Figure 11 In the example, for a unit symbol such as “i triple e” in the first learning output symbol, the quantity contained in the first learning output symbol E2 is less than (less than) the specified value. Therefore, this unit symbol is selected as the first symbol, and the first symbol is converted into the second learning output symbol.

[0170] [The information assigned to the symbol]

[0171] The decision-making department 120 may also assign prescribed information symbols to the symbols (registration symbols). An information symbol is a symbol with a prescribed meaning. An information symbol includes a start symbol indicating the beginning of the registration symbol and an end symbol indicating the end of the registration symbol. Start symbol' <bob>' is an abbreviation for "beginning of biasing," signifying the start of the registration symbol. Additionally, there is an end symbol ''. <eob>' is an abbreviation for end of biasing, signifying the end of the registration symbol. Thus, the decision-making unit 120 can clearly determine the registration symbol by assigning a start symbol and an end symbol to it.

[0172] Additionally, the preprocessing unit 240 may also assign a probability symbol to the first symbol. The probability symbol is a symbol that represents the probability (proportion) of replacing the first symbol with the second symbol.

[0173] Figure 17 This is a diagram representing an example of the probability symbol 80. In Figure 17 The text shows the assignment of the first symbol to... <r20>Examples of such probability symbols are 80. <r20>This means that there is a 20% probability of replacing the first symbol with the second symbol. That is, in... Figure 17 In the example, the first symbol is replaced with the second symbol at a rate of 1 in 5 times. Furthermore, the probability symbol 80 can be specified by the user or determined by the learning device 200 through a prescribed operation.

[0174] [Method for determining the registration symbol]

[0175] (1) In addition, there are cases where multiple specific symbols exist (e.g., Figure 13 (In the case of...). In this case, the decision unit 120 may also consider the performance of the registration process for the registered symbol, and make the method for determining the registered symbol different for each specific symbol. For example, the decision unit 120 may determine the first specific symbol among a plurality of specific symbols as the registered symbol itself. Alternatively, the decision unit 120 may determine a symbol different from the second specific symbol among a plurality of specific symbols as the registered symbol.

[0176] (2) In the case of Japanese speech recognition, words pronounced "chooshi" include "chooshi" and others besides "chooshi". In this case, when the decision unit 120 determines "chooshi" as the registration symbol for "chooshi", in the derivation unit 140, as a temporary output symbol D9, all symbols that should have been output as "chooshi" are also output as "chooshi". In this case, after the conversion unit 150 performs the conversion process, most of the parts that should have been output as "chooshi" become "chooshi". Even if "chooshi" is registered as the registration symbol, if the derivation unit 140 can prioritize outputting "chooshi", then "chooshi" can be determined as the registration symbol instead of "chooshi". The method for determining the registration symbol can also be based on the principle that there are fewer instances where the part that should be incorrectly set as "tone" becomes "mark".

[0177] (3) The method of determining the registration symbol D6 by the determination unit 120 (information processing device 100) is not limited to the above-described embodiment, and may be other methods. Figure 18 This is an example of a flowchart illustrating the first determination method for registration symbol D6. Figure 18 In the example, multiple learning output symbols are used (e.g., the first learning output symbol E2). The learning output symbol is at least one of the symbols used after being used in the learning device 200 and the symbols used before being used in the learning device 200.

[0178] In step S102, the determination unit 120 acquires a plurality of learning output symbols. Next, in step S104, the determination unit 120 determines the number of specific symbols included in the plurality of learning output symbols and judges whether the number of specific symbols is greater than or equal to a predetermined value. If the number of specific symbols is greater than or equal to the predetermined value, the process proceeds to step S106; if the number of specific symbols is less than the predetermined value, the process proceeds to step S108.

[0179] In step S106, the determination unit 120 determines a specific symbol as a registered symbol. For example, in Figure 5 In the example, "IEEE" as a specific symbol is itself stored as registration symbol D6. Then, Figure 18 The processing ends. On the other hand, in step S106, the determination unit 120 determines the symbol represented by an expression different from the specific symbol as the registered symbol. For example, in Figure 5 In the example, the "phonetic symbol of IEEE" corresponding to "IEEE" as a specific symbol is stored as registration symbol D6.

[0180] Next, explain the... Figure 18 The meaning of the flowchart. When a specific symbol included in multiple learning output symbols exceeds a specified value (in... Figure 18 In step S104 (if yes), there is a tendency for the likelihood value of candidate symbols containing a specific symbol to increase. In such a case, by determining the specific symbol as the registered symbol, it is possible to output that specific symbol as a temporary output symbol with priority. Furthermore, in this case, since the specific symbol is output as a temporary output symbol, the aforementioned substitution process is not performed.

[0181] On the other hand, when a specific symbol included in multiple learning output symbols is less than a specified value (in... Figure 18 (If step S104 is not specified), there is a tendency for the likelihood value of candidate symbols containing a specific symbol to decrease. In such cases, by determining a symbol represented by a different expression than the specific symbol as the registered symbol, the registered symbol can be preferentially output as a temporary output symbol.

[0182] (4) In addition, such as Figure 18 As indicated by the parentheses, the decision unit 120 may also use multiple sample output symbols instead of multiple (N: N is an integer greater than 2) learning output symbols. Regarding the sample output symbols, for example, a user inputs N input data (sample input data) to the information processing device 100, and N sample output symbols are output as output symbols. Figure 18 In the bracketed structure, the decision unit 120 uses N sample output symbols.

[0183] Next, the explanation Figure 18 The meaning of the bracketed structure. When a specific symbol included in multiple sample output symbols exceeds a specified value (in... Figure 18 In step S104 (if yes), there is a tendency for the likelihood value of a candidate symbol containing a specific symbol to increase. In such a case, by determining the specific symbol as the registered symbol, it is possible to output that specific symbol as a temporary output symbol with priority. Furthermore, in this case, since the specific symbol is output as a temporary output symbol, the above-described substitution process is not performed.

[0184] On the other hand, when a specific symbol in multiple sample output symbols is less than a specified value (in... Figure 18 (If step S104 is not specified), there is a tendency for the likelihood value of candidate symbols containing a specific symbol to decrease. In such cases, by determining a symbol represented by a different expression than the specific symbol as the registered symbol, the registered symbol can be preferentially output as a temporary output symbol.

[0185] In addition, through Figure 18 The bracketed structure allows for the determination of the registration symbol D6 for each specific symbol contained in a specific symbol table D2, in a manner that maximizes the frequency of output as a temporary output symbol D9. Furthermore, through... Figure 18 The bracketed structure, for each specific symbol contained in the specific symbol table D2, can reduce the number of specific symbols that are incorrectly output as output symbols.

[0186] [About output symbols]

[0187] For example, such as Figure 18 As shown, when the method for determining the registration symbol D6 is changed, the output symbol D4 may be changed accordingly. Therefore, the information processing device 100 can also determine multiple output symbols D4, select the best output symbol from the multiple output symbols D4, and output the best output symbol when the method for determining the registration symbol is changed.

[0188] [About combination tables and bias tables]

[0189] In the above embodiments, an example was described where the information processing apparatus 100 generates the combination table and the bias table itself. However, at least one of the combination table and the bias table may also be generated by other apparatus.

[0190] [Postscript]

[0191] The parentheses shown below are merely an example for reference and are not limited to the content containing parentheses.

[0192] (1) The information processing apparatus of this disclosure converts input data (input data D1) into output symbols (output symbols D4) for output. The information processing apparatus includes an interface for accepting user input of input data, a memory for storing a learned model, and a processor. The processor extracts feature quantities (feature quantity E5) from the input data. The processor derives at least one candidate symbol ("i triple e" and "IEEE phoneme symbol") and a likelihood value representing the likelihood of each of the at least one candidate symbol by applying the feature quantity to the learned model. The memory stores a combination table ( Figure 6 Combination table D7) and bias module ( Figure 7 The bias table (D8) is used. The combination table represents a combination of a specific symbol specified by the user and the corresponding registration symbol. The bias module is a module that increases the likelihood value of a candidate symbol containing the registration symbol. The processor also performs an increase process that uses the bias module to increase the likelihood value of at least one candidate symbol containing the registration symbol. Based on the likelihood values ​​of each of the at least one candidate symbol after the increase process, the processor determines a temporary output symbol (temporary output symbol D9). Then, if the temporary output symbol contains the registration symbol, the processor performs a conversion process that refers to the combination table to convert the registration symbol into the specific symbol corresponding to the registration symbol, and outputs the temporary output symbol after the conversion process as the output symbol (output symbol D4).

[0193] For example, there may be cases where a user wants a specific symbol (the aforementioned specific symbol) to be prioritized as the output symbol from the information processing device. In such cases, in information processing devices employing an end-to-end model, the conversion from input data to an output symbol sequence is achieved through a single neural network. Therefore, there is a difficulty in performing customizations such as adding registered vocabulary through partial changes.

[0194] Furthermore, in the technology disclosed in Japanese Patent Application Publication No. 2022-531615, there are many cases where the registration of the specific symbol is not generated during the aforementioned substitution process. In this case, the problem may arise that the specific symbol is difficult to output as an output symbol. That is, in existing information processing devices, the problem may arise that it is difficult to prioritize the output of a specific symbol input by the user as an output symbol.

[0195] Therefore, in this disclosure, the information processing apparatus derives at least one candidate symbol and a likelihood value representing the likelihood of each of the at least one candidate symbol based on input data. Next, the information processing apparatus performs an increase process based on a bias module, increasing the likelihood value of the candidate symbol containing the registered symbol. This allows the registered symbol (or a symbol containing the registered symbol) to be preferentially determined as a temporary output symbol. Furthermore, if the temporary output symbol contains the registered symbol, the information processing apparatus performs a conversion process that refers to a combination table to convert the registered symbol into a specific symbol corresponding to that registered symbol, and outputs the temporary output symbol after the conversion process as the output symbol.

[0196] Therefore, the above customization can be easily performed by inputting a specific symbol. Furthermore, by performing an addition process to prioritize the output of the registered symbol, and a substitution process to change the registered symbol to the specific symbol, the specific symbol can be prioritized as the output symbol.

[0197] (2) Preferably, the bias module is a bias table that specifies a bias value that increases the likelihood value of a candidate symbol containing the registration symbol and corresponds to the registration symbol. Figure 7 ).

[0198] Based on this structure, increment processing can be performed by reflecting the bias value corresponding to the registration symbol in the likelihood value.

[0199] (3) Preferably, the learned model is trained based on multiple learning data that are a combination of learning features and learning registration symbols. The memory stores a decision model (decision model 305) that outputs bias values ​​by inputting registration symbols. The decision model outputs a bias value that increases the likelihood of a candidate symbol containing a registration symbol if the number of input registration symbols in the learning registration symbols is smaller. The processor generates a bias table by mapping the bias values ​​output from the decision model to registration symbols that are the same as the learning registration symbols.

[0200] Based on this structure, a decision model is used that determines the likelihood value of candidate symbols containing registration symbols by using a bias value that increases the number of registration symbols output for learning. Furthermore, the processor generates a bias table by mapping the bias values ​​output from the decision model to registration symbols that are identical to the registration symbols used for learning. Therefore, even when the likelihood value of candidate symbols containing registration symbols is low, this likelihood value can be increased by increasing the processing.

[0201] (4) Preferably, the processor generates a bias table (input device 4) based on the user's input.

[0202] Based on this structure, a bias table can be generated that specifies bias values ​​corresponding to the user's expectations.

[0203] (5) Preferably, the bias table specifies the same bias value for each of the multiple registration symbols (see reference). Figure 15 (in parentheses).

[0204] With this structure, all applied bias values ​​are identical, thus simplifying the aforementioned augmentation process.

[0205] (6) Preferably, the processor generates a bias module ( Figure 4 (Processing of the generation unit 130).

[0206] Based on this structure, the information processing device itself can generate a bias module.

[0207] (7) Preferably, the interface accepts user input of a specific symbol. The processor generates a combination table representing the combination of the specific symbol input to the interface and the corresponding registered symbol (processing of the decision unit 120).

[0208] Based on this structure, the information processing device itself can generate combination tables.

[0209] (8) Preferably, the interface acquires multiple sample output symbols (by applying multiple sample input data to the information processing device) Figure 18 (The brackets). If the number of specific symbols contained in multiple sample output symbols is greater than or equal to a predetermined value (yes in step S104), the processor determines that specific symbol as a registered symbol. If the number of specific symbols contained in multiple sample output symbols is less than a predetermined value (no in step S104), the processor determines a symbol represented by an expression different from that specific symbol as a registered symbol.

[0210] Based on this structure, when the number of specific symbols included in multiple sample output symbols exceeds a predetermined value, there is a tendency for the likelihood value of candidate symbols containing the specific symbol to increase. In such cases, by determining the specific symbol as the registered symbol, it is possible to prioritize the output of that specific symbol as a temporary output symbol. Furthermore, in this case, since the specific symbol is output as a temporary output symbol, the aforementioned permutation process is not performed.

[0211] On the other hand, when the number of specific symbols included in multiple sample output symbols is less than a specified value, there is a tendency for the likelihood value of candidate symbols containing the specific symbol to decrease. In such cases, by determining the symbol represented by a different expression than the specific symbol as the registered symbol, the registered symbol can be output as a temporary output symbol with priority.

[0212] (9) Preferably, the learned model is a model learned based on multiple learning data, which are combinations of learning input data and learning output symbols. The interface obtains multiple learning output symbols ( Figure 18 If the number of specific symbols included in the multiple learning output symbols is greater than or equal to a predetermined value (yes in step S104), the processor determines that specific symbol as a registered symbol. If the number of specific symbols included in the multiple learning output symbols is less than the predetermined value (no in step S104), the processor determines a symbol represented by an expression different from that specific symbol as a registered symbol.

[0213] Based on this structure, when a specific symbol is included in multiple learning output symbols and exceeds a predetermined value, there is a tendency for the likelihood value of candidate symbols containing the specific symbol to increase. In such a case, by determining the specific symbol as a registered symbol, it is possible to prioritize outputting that specific symbol as a temporary output symbol. Furthermore, in this case, since the specific symbol is output as a temporary output symbol, the aforementioned substitution process is not performed.

[0214] On the other hand, when the number of specific symbols included in multiple learning output symbols is less than a specified value, there is a tendency for the likelihood value of candidate symbols containing the specific symbol to decrease. In such cases, by determining a symbol represented by an expression different from the specific symbol as the registered symbol, the registered symbol can be output as a temporary output symbol with priority.

[0215] (10) Preferably, the addition process is the process of not adding the likelihood value of candidate symbols that do not contain the registered symbol among at least one candidate symbol (see reference). Figure 15 (The "Other than this" column).

[0216] Based on this structure, it is possible to suppress the increase in the likelihood value of candidate symbols that do not contain the registration symbol, and as a result, it is possible to prioritize candidate symbols that contain the registration symbol as temporary output symbols.

[0217] (11) Preferably, the registration symbol is any one of the following: the same as the specific symbol, or the expression method is different from the specific symbol, or the reading method is the same as the specific symbol but the expression method is different (see reference). Figure 6 and Figure 14 ).

[0218] Based on this structure, it is possible to define a registration symbol that can be any of the above.

[0219] (12) Preferably, the specific symbol is expressed by any one of hiragana, katakana, romanization, English, and the International Phonetic Alphabet, and the registration symbol is expressed by another one of them (see reference). Figure 6 and Figure 14 ).

[0220] Based on this structure, specific symbols and registration symbols can be expressed using hiragana, katakana, romaji, English, and the International Phonetic Alphabet.

[0221] (13) Preferably, the input data is speech data. The output symbol is text representing the speech data, or text summarizing the content of the speech data (see reference). Figure 16 ).

[0222] Based on this structure, it is possible to output text that is speech or text that summarizes the content of speech.

[0223] (14) Preferably, the input data is any one of sound data, still image data, and moving image data. Furthermore, the output symbol is text representing a description of the object represented by the input data (see [reference]). Figure 16 ).

[0224] Based on this structure, it is possible to output text that represents a description of the object represented by the input data.

[0225] (15) The learning apparatus of this disclosure is a learning apparatus for updating a model. The learning apparatus includes an interface for acquiring learning data as a combination of learning input data (learning input data E1) and a first learning output symbol (first learning output symbol E2) and a computational device. The computational device extracts feature quantities (feature quantities E4) from the learning input data. The computational device obtains an output symbol (output symbol E6) by applying the feature quantities to the model. The computational device performs preprocessing on the first learning output symbol (…). Figure 11 The processing unit generates the second learning output symbol. The processing unit updates the model to reduce the error between the first learning output symbol and the second learning output symbol. Preprocessing involves generating the first symbol contained in the first learning output symbol and the second symbol representing that first symbol using a user-specified expression. Figure 11 ).

[0226] Based on this structure, as preprocessing, a process is performed to generate a first symbol included in the first learning output symbol and a second symbol representing the first symbol in a user-specified expression. Furthermore, the computing device can update the model in such a way that the likelihood value of the second symbol increases, thereby updating the model in a way that increases the likelihood value of the second symbol representing the first symbol in a user-specified expression.

[0227] (16) Preferably, the second symbol is any one of the following: the same as the first symbol, or the expression method is different from the specific symbol, or the reading method is the same as the specific symbol but the expression method is different ( Figure 11 ).

[0228] Based on this structure, a second symbol can be defined as any of the above.

[0229] (17) Preferably, the first learning output symbol is text. Preprocessing involves segmenting the text according to unit symbols by performing prescribed processing, and selecting the unit symbol as the first symbol. Figure 11 ).

[0230] Based on this structure, when the output symbol for the first learning is text, it is possible to select the word contained in the text as the first symbol.

[0231] (18) Preferably, the unit symbol is a noun or proper noun ( Figure 11 ).

[0232] Based on this structure, the model can be updated by increasing the likelihood value of the second symbol corresponding to the first symbol as a noun or proper noun.

[0233] (19) Preferably, in the preprocessing, when the number of unit symbols included in the first learning output symbol is less than a specified value, the unit symbol is selected as the first symbol. Figure 11 ).

[0234] Based on this structure, even when the number of symbols contained in the first learning output symbol is small, the symbol can be selected as the first symbol, thus improving the likelihood value of the second symbol corresponding to the first symbol.

[0235] (20) Preferably, the preprocessing assigns information representing the probability that the first symbol is converted into the second symbol to the first symbol (see reference). Figure 17 ).

[0236] Based on this structure, the probability of the first symbol being converted into the second symbol can be controlled.

[0237] The embodiments disclosed herein should be considered illustrative and not restrictive in all respects. The scope of the invention is defined not by the description of the above embodiments, but by the claims, and is intended to include all modifications within the meaning and scope equivalent to the claims.

[0238] Label Explanation

[0239] 2: Microphone; 3: Display; 4: Input device; 11: Arithmetic unit; 12: Storage unit; 13: Voice interface; 14: Communication unit; 15: Display interface; 16: Input device interface; 20: Removable disk; 30: Information processing system; 80: Probability symbol; 100: Information processing device; 110, 220: Extraction unit; 120: Decision unit; 130: Generation unit; 140: Derivation unit; 150: Conversion unit; 200: Learning device; 210: Acquisition unit; 230: Processing unit; 235: Model storage unit; 240: Preprocessing unit; 250: Update unit; 260: Output unit; 300: Learning data storage device; 301: Processing program; 303: Neural network; 304: Parameters; 305: Decision model; 400: Learned model storage device. < / eob> < / bob>

Claims

1. An information processing apparatus that converts input data into output symbols and outputs them, the information processing apparatus comprising: An interface that accepts user input of the input data; Memory, which stores the learned model; and processor, The processor Extract feature values ​​from the input data. By applying the feature quantities to the learned model, at least one candidate symbol and a likelihood value representing the likelihood of each of the at least one candidate symbol are derived. The memory stores the combination table and the bias module. The combination table is a table that represents combinations of a specific symbol specified by the user and the corresponding registered symbol. The bias module is a module that increases the likelihood value of candidate symbols containing the registered symbols. The processor also Perform the process of increasing the likelihood value of the candidate symbol containing the registered symbol among the at least one candidate symbol using the bias module. Based on the likelihood values ​​of each of the at least one candidate symbol after the augmentation process has been performed, a temporary output symbol is determined. If the temporary output symbol contains the registered symbol, a conversion process is performed to convert the registered symbol into the specific symbol corresponding to the registered symbol by referring to the combination table, and the temporary output symbol after the conversion process is performed is output as the output symbol.

2. The information processing apparatus according to claim 1, wherein, The bias module is a bias table that defines a bias value that increases the likelihood value of a candidate symbol containing the registered symbol and corresponds to the registered symbol.

3. The information processing apparatus according to claim 2, wherein, The learned model is trained based on multiple learning data sets, which are combinations of learning features and learning registration symbols. The memory stores a decision model that outputs bias values ​​by inputting the registration symbol. The decision model outputs a bias value that increases the likelihood of candidate symbols containing a given registration symbol if the number of such registration symbols included in the input registration symbols is smaller. The processor generates the bias table by mapping the bias values ​​output from the decision model to the registration symbols that are the same as the learning registration symbols.

4. The information processing apparatus according to claim 2, wherein, The processor generates the bias table based on user input.

5. The information processing apparatus according to claim 2, wherein, The bias table specifies the same bias value for each of the registered symbols.

6. The information processing apparatus according to claim 1, wherein, The processor generates the bias module.

7. The information processing apparatus according to claim 1, wherein, The interface accepts user input for the specific symbol. The processor generates a table as the combination table, which represents the combination of the specific symbol input to the interface and the registration symbol corresponding to the specific symbol.

8. The information processing apparatus according to claim 1, wherein, The interface acquires multiple sample output symbols by applying multiple sample input data to the information processing device. The processor If the number of the specific symbol included in the plurality of sample output symbols exceeds a predetermined value, then the specific symbol is determined to be the registered symbol. If the number of the specific symbols included in the plurality of sample output symbols is less than the specified value, the symbol represented by an expression different from that specific symbol will be determined as the registered symbol.

9. The information processing apparatus according to claim 1, wherein, The learned model is a model learned based on multiple learning data, which are combinations of input data and output symbols used for learning. The interface acquires multiple output symbols used for learning. The processor If the number of the specific symbols included in the plurality of learning output symbols exceeds a predetermined value, then the specific symbol is determined to be the registered symbol. If the number of the specific symbols included in the plurality of learning output symbols is less than the specified value, the symbol represented by an expression different from that specific symbol will be determined as the registered symbol.

10. The information processing apparatus according to claim 1, wherein, The increment process is the process of not increasing the likelihood value of the candidate symbols that do not contain the registered symbol among the at least one candidate symbols.

11. The information processing apparatus according to claim 1, wherein, The registration symbol is any one of the following: it is the same as the specific symbol, or its expression method is different from the specific symbol, or its pronunciation is the same as the specific symbol but its expression method is different.

12. The information processing apparatus according to claim 10, wherein, The specific symbol is expressed by any one of hiragana, katakana, romanization, English, and the International Phonetic Alphabet, and the registration symbol is expressed by another one of them.

13. The information processing apparatus according to claim 1, wherein, The input data is voice data. The output symbol is text representing speech data, or text summarizing the content of speech data.

14. The information processing apparatus according to claim 1, wherein, The input data can be any one of sound data, still image data, and moving image data. The output symbol is a textual description of the object represented by the input data.

15. A learning device for updating a model, the learning device comprising: An interface that acquires learning data as a combination of input data for learning and the first output symbol for learning; and processor, The processor Extract features from the input data used for learning. The output sign is obtained by applying the feature quantity to the model. The second learning output symbol is generated by performing preprocessing on the first learning output symbol. The model is updated in a manner that reduces the error between the output symbol and the second learned output symbol. The preprocessing is as follows: generating a first symbol contained in the first learning output symbol and a second symbol representing the first symbol in an expression specified by the user.

16. The learning device according to claim 15, wherein, The second symbol is any one of the following: it is the same as the first symbol, or its expression method is different from that of the specific symbol, or its pronunciation is the same as that of the specific symbol but its expression method is different.

17. The learning device according to claim 15, wherein, The first learning method uses text as the output symbol. The preprocessing involves performing prescribed processing on the text to segment it according to a unit symbol, and selecting that unit symbol as the first symbol.

18. The learning device according to claim 17, wherein, The unit symbols mentioned are nouns or proper nouns.

19. The learning device according to claim 15, wherein, In the preprocessing, if the number of unit symbols included in the first learning output symbol is less than a predetermined value, the unit symbol is selected as the first symbol.

20. The learning device according to claim 15, wherein, The preprocessing assigns information representing the probability that the first symbol is converted into the second symbol to the first symbol.

21. An information processing method that converts input data into output symbols and outputs them, wherein, The bias module is a module that increases the likelihood value of candidate symbols containing registered symbols. A combination table is a table that represents combinations of a specific symbol specified by the user and the corresponding registered symbol. The information processing method includes: Extract feature quantities from the input data; By applying the feature quantities to the learned model, at least one candidate symbol and a likelihood value representing the likelihood of each of the at least one candidate symbol are derived; Perform the process of increasing the likelihood value of the candidate symbol among the at least one candidate symbol that includes the registered symbol corresponding to a specific symbol specified by the user, using the bias module; Perform the process of increasing the likelihood value of the candidate symbol containing the registered symbol among the at least one candidate symbol using the bias module; The temporary output symbol is determined based on the likelihood value of each of the at least one candidate symbol after the augmentation process has been performed; as well as If the temporary output symbol contains the registered symbol, a conversion process is performed to convert the registered symbol into the specific symbol corresponding to the registered symbol by referring to the combination table, and the temporary output symbol after the conversion process is performed is output as the output symbol.

22. A learning method for updating a model, the learning method having: Acquire learning data as a combination of input data for learning and the first output symbol for learning; Extract features from the input data used for learning; The output sign is obtained by applying the feature quantity to the model; The second learning output symbol is generated by performing preprocessing on the first learning output symbol; and The model is updated in a manner that reduces the error between the output symbol and the second learned output symbol. The preprocessing is as follows: generating a first symbol contained in the first learning output symbol and a second symbol representing the first symbol in an expression specified by the user.

23. A non-transitory recording medium storing a program for enabling a computer to convert input data into output symbols and output them, wherein, The bias module is a module that increases the likelihood value of candidate symbols containing registered symbols. A combination table is a table that represents combinations of a specific symbol specified by the user and the corresponding registered symbol. The program causes the computer to execute: Extract feature quantities from the input data; By applying the feature quantities to the learned model, at least one candidate symbol and a likelihood value representing the likelihood of each of the at least one candidate symbol are derived; Perform the process of increasing the likelihood value of the candidate symbol among the at least one candidate symbol that includes the registered symbol corresponding to a specific symbol specified by the user, using the bias module; Perform the process of increasing the likelihood value of the candidate symbol containing the registered symbol among the at least one candidate symbol using the bias module; The temporary output symbol is determined based on the likelihood value of each of the at least one candidate symbol after the augmentation process has been performed; as well as If the temporary output symbol contains the registered symbol, a conversion process is performed to convert the registered symbol into the specific symbol corresponding to the registered symbol by referring to the combination table, and the temporary output symbol after the conversion process is performed is output as the output symbol.

24. A non-transitory recording medium storing a program for updating a model in a computer, wherein, The program causes the computer to execute: Acquire learning data as a combination of input data for learning and the first output symbol for learning; Extract feature quantities from the input data used for learning; The output sign is obtained by applying the feature quantity to the model; The second learning output symbol is generated by performing preprocessing on the first learning output symbol; and The model is updated in a manner that reduces the error between the output symbol and the second learned output symbol. The preprocessing is the process of generating a first symbol contained in the first learning output symbol and a second symbol representing the first symbol in an expression specified by the user.

Citation Information

Patent Citations

  • Using context information in an end-to-end model for speech recognition

    JP2022531615A