Processing apparatus, processing method, and program
By determining and processing important pairs of audio and text data with speech recognition technologies, the method enhances learning relevance and accuracy by generating additional data through conversions, addressing the uniform treatment of all data in existing systems.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- NEC CORP
- Filing Date
- 2024-05-10
- Publication Date
- 2026-07-22
AI Technical Summary
Existing speech recognition technologies do not adequately account for the varying importance of learning data, leading to uniform treatment that fails to reflect the content of important data effectively.
A processing device and method that determines the importance of pairs of input audio data and text data, processing important pairs through conversions like speech rate, voice quality, and noise overlay to generate additional learning data, while ignoring unimportant pairs, thereby enhancing the relevance of important data in learning.
This approach increases the number of learning data that reflects important content, improving the accuracy and efficiency of speech recognition by focusing on significant data points and reducing user workload.
Smart Images

Figure 0007893375000001 
Figure 0007893375000002 
Figure 0007893375000003
Abstract
Description
Technical Field
[0004] ,
[0006]
[0001] The present invention relates to a processing device, a processing method, and a program.
Background Art
[0002] The technology related to the present invention is disclosed in Patent Document 1. Patent Document 1 discloses a technology for generating learning data of a speech recognition engine. Specifically, Patent Document 1 discloses a technology that uses learning data (speech data) collected in the past and conversion data (speech data) converted using the collected learning data.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] The importance of the collected learning data may vary for each learning data. Therefore, it is desired to perform learning that reflects the content of important learning data more. If learning is performed by uniformly treating all learning data in the same manner, such learning cannot be realized. Patent Document 1 does not disclose the problem and its solution means.
[0005] An example of the object of the present invention is to provide a processing device, a processing method, and a program that can perform learning that reflects the content of important learning data in the learning of a speech recognition engine in view of the above-described problem.
Means for Solving the Problems
[0006] According to one aspect of the present invention, A speech recognition means that performs speech recognition processing on input speech data and outputs text data indicating the content of the speech in the input speech data, A determination means for determining the importance of each pair of input audio data and text data as learning data, A generation means that processes the input audio data of the important pair, which is the pair whose importance satisfies the predetermined conditions, to generate at least one processed audio data, and generates additional learning data by combining the processed audio data and the text data, A processing device having the following is provided.
[0007] According to one aspect of the present invention, One or more computers, Speech recognition processing is performed on the input audio data, and text data representing the content of the utterance from the input audio data is output. For each pair of input audio data and text data, the importance of the data as training data is determined. A processing method is provided which involves processing the input audio data of the important pairs, which are pairs whose importance satisfies predetermined conditions, to generate at least one processed audio data, and generating additional learning data by combining the processed audio data and the text data.
[0008] According to one aspect of the present invention, Computers, A speech recognition means that performs speech recognition processing on input speech data and outputs text data indicating the content of the speech in the input speech data. For each pair of input audio data and text data, a determination means for determining the importance of the data as learning data. A generation means that processes the input audio data of the important pair, which is the pair whose importance satisfies the predetermined conditions, to generate at least one processed audio data, and generates additional learning data by combining the processed audio data and the text data. A program is provided to enable it to function as such. [Effects of the Invention]
[0009] According to one aspect of the present invention, in the learning of a speech recognition engine, a processing device, a processing method, and a program are realized that enable learning that more reflects the content of important learning data among the collected learning data.
Brief Description of the Drawings
[0010] The above-described objects and other objects, features, and advantages will become even more apparent from the preferred embodiments described below and the accompanying drawings.
[0011] [Figure 1] It is a diagram showing an example of a functional block diagram of a processing device. [Figure 2] It is a diagram for explaining an example of the processing executed by the processing device. [Figure 3] It is a diagram for explaining another example of the processing executed by the processing device. [Figure 4] It is a diagram showing an example of the hardware configuration of the processing device. [Figure 5] It is a diagram showing an example of the screen output by the processing device. [Figure 6] It is a flowchart showing an example of the processing flow of the processing device. [Figure 7] It is a flowchart showing an example of the processing flow of the processing device. [Figure 8] It is a diagram showing an example of a functional block diagram of the processing device.
Embodiments for Carrying Out the Invention
[0012] Hereinafter, embodiments of the present invention will be described with reference to the drawings. In all the drawings, the same components are denoted by the same reference numerals, and the description will be omitted as appropriate.
[0013] <First Embodiment> FIG. 1 is a functional block diagram showing an overview of a processing device 10 according to the first embodiment. The processing device 10 includes a speech recognition unit 11, a determination unit 12, and a generation unit 13.
[0014] The voice recognition unit 11 executes voice recognition processing on the input voice data and outputs text data indicating the utterance content of the input voice data.
[0015] The determination unit 12 determines the importance (hereinafter sometimes simply referred to as "importance") as learning data for each pair of the input voice data and the text data (hereinafter sometimes simply referred to as "pair").
[0016] The generation unit 13 processes the input voice data of the important pair, which is a pair whose importance satisfies a predetermined condition, to generate at least one processed voice data, and generates additional learning data combining the processed voice data and the text data.
[0017] As described above, the processing device 10 according to the first embodiment generates processed voice data from the input voice data for an important pair that is important as learning data, and generates additional learning data including each of the generated processed voice data. As a result, the learning data regarding the important pair increases by the amount of the additional learning data. Thus, the processing device 10 according to the first embodiment enables learning that more reflects the content of important learning data by increasing the number of learning data regarding important pairs that are important as learning data.
[0018] <Second Embodiment> "Overview"
[0019] The processing device 10 of the second embodiment is a具体化 of the processing device 10 of the first embodiment.
[0020] That is, the processing device 10 executes voice recognition processing on the input voice data and outputs text data indicating the utterance content of the input voice data. Next, the processing device 10 determines the importance as learning data for each pair of the input voice data and the text data.
[0021] The processing unit 10 then performs training data generation processing for each pair according to its importance. For important pairs whose importance meets predetermined conditions (conditions indicating importance), the processing unit 10 performs the processing shown in Figure 2. On the other hand, for unimportant pairs whose importance does not meet predetermined conditions (conditions indicating importance), the processing unit 10 performs the processing shown in Figure 3.
[0022] Figure 2 shows key pairs. A key pair includes input audio data and text data generated by performing speech recognition processing on that input audio data.
[0023] The processing unit 10 processes the input audio data of important pairs to generate at least one processed audio data. The generation of the processed audio data is achieved by at least one of the following processes: speech rate conversion, voice quality conversion, noise superposition, and sound quality conversion. Therefore, the content of the processed audio data is the same as the content of the input audio data.
[0024] As shown in the diagram, the processing unit 10 generates training data by combining the input audio data and text data of important pairs. The processing unit 10 also generates additional training data by combining the text data of important pairs with at least one processed audio data for each pair.
[0025] Figure 3 shows non-essential pairs. Non-essential pairs include input audio data and text data generated by performing speech recognition processing on that input audio data.
[0026] The processing unit 10 generates training data by combining the input audio data and text data of non-important pairs. For non-important pairs, the processing unit 10 does not generate processed audio data or additional training data that includes processed audio data.
[0027] In this way, for important pairs, the processing unit 10 generates "training data combining the input audio data and text data of the important pair" and "additional training data combining the text data of the important pair with at least one processed audio data for each pair." Then, for non-important pairs, the processing unit 10 generates "training data combining the input audio data and text data of the non-important pair."
[0028] With this type of processing device 10, the amount of training data for important pairs can be increased compared to the amount of training data for unimportant pairs by the amount of additional training data. By training the speech recognition engine with this training data, learning that better reflects the content of the training data for important pairs can be achieved.
[0029] The configuration of the processing unit 10 will be described in detail below.
[0030] "Hardware configuration" An example of the hardware configuration of the processing unit 10 is described below. Each functional unit of the processing unit 10 is realized by any combination of hardware and software. It will be understood by those skilled in the art that there are various modifications to the implementation method and the device. The software includes programs that are pre-installed at the time of shipment of the device, as well as programs downloaded from recording media such as CDs (Compact Discs) or from servers on the Internet.
[0031] Figure 4 is a block diagram illustrating the hardware configuration of the processing unit 10. As shown in Figure 4, the processing unit 10 includes a processor 1A, memory 2A, input / output interface 3A, peripheral circuitry 4A, and bus 5A. Peripheral circuitry 4A includes various modules. The processing unit 10 does not necessarily have peripheral circuitry 4A. The processing unit 10 may also be composed of multiple physically and / or logically separated devices. In this case, each of the multiple devices may have the above hardware configuration.
[0032] Bus 5A is a data transmission path for the processor 1A, memory 2A, peripheral circuits 4A, and input / output interface 3A to send and receive data to and from each other. Processor 1A is a processing unit such as a CPU or GPU (Graphics Processing Unit). Memory 2A is a memory such as RAM (Random Access Memory) or ROM (Read Only Memory). Input / output interface 3A includes interfaces for acquiring information from input devices, external devices, external servers, external sensors, cameras, etc., and interfaces for outputting information to output devices, external devices, external servers, etc. Input / output interface 3A also includes interfaces for connecting to communication networks such as the Internet. Input devices include, for example, keyboards, mice, microphones, physical buttons, and touch panels. Output devices include, for example, displays, speakers, printers, and mailers. Processor 1A can issue commands to each module and perform calculations based on their calculation results.
[0033] "Functional Configuration" Next, the functional configuration of the processing device 10 of the second embodiment will be described in detail. Figure 1 shows an example of a functional block diagram of the processing device 10 of the second embodiment. As shown in the figure, the processing device 10 of the second embodiment has a speech recognition unit 11, a determination unit 12, and a generation unit 13.
[0034] The speech recognition unit 11 performs speech recognition processing on the input speech data and outputs speech content text data based on the results of the speech recognition processing.
[0035] "Input audio data" is audio data that represents the content of a person's speech. The speech recognition unit 11 acquires multiple input audio data. The length of one input audio data is, for example, a few seconds, but is not limited to this. The lengths of multiple input audio data may all be the same, or they may be different from one another.
[0036] The speech recognition unit 11 can acquire input speech data by any means. "Acquisition" includes at least one of the following: the device retrieving data or information stored in another device or storage medium (active acquisition), and the device inputting data or information output from another device (passive acquisition). Examples of active acquisition include making a request to another device and receiving a reply, and accessing and reading data from another device or storage medium. Examples of passive acquisition include receiving information that is distributed (or transmitted, push notification, etc.). Furthermore, "acquisition" may also mean selecting and acquiring data or information from among the received data or information, or selecting and receiving distributed data or information.
[0037] The speech recognition unit 11 performs speech recognition processing on input speech data using a pre-prepared speech recognition engine. The speech recognition engine is generated by learning based on training data that combines speech data and text data representing the content of the speech utterance. When speech data is input to the speech recognition engine, text data representing the content of the speech utterance is output. The configuration of the speech recognition engine is not particularly limited, and any technology can be adopted. The training data generated by the processing unit 10 can be used to retrain this speech recognition engine.
[0038] "Speech content text data" is text data that shows the content of the speech input. The speech recognition unit 11 may output the recognition result (text data) output from the speech recognition engine as is, as speech content text data.
[0039] In addition, the speech recognition unit 11 may perform correction processing on the recognition result (text data) output from the speech recognition engine. The speech recognition unit 11 may then output the text data after the correction has been applied to the recognition result (text data) output from the speech recognition engine as the utterance content text data.
[0040] Here, we will explain an example of the correction process. The speech recognition unit 11 outputs the recognition result (text data) output from the speech recognition engine to the user. This output is realized via an output device such as a display or projection device. The speech recognition unit 11 then accepts user input to correct the outputted recognition result (text data). The user specifies the part of the outputted text data that needs correction and inputs the correct content for the specified part. Receiving such user feedback on the recognition result is realized using any widely known technology.
[0041] The determination unit 12 determines the importance of each pair of input audio data and spoken content text data as training data. One pair consists of one input audio data and spoken content text data that shows the content of the input audio data.
[0042] "Importance" can be expressed using any indicator. For example, importance may be expressed using two values: "important" and "unimportant." Alternatively, importance may be expressed using four values from 0 to 3. For example, a higher value indicates higher importance. Note that the examples given here are merely examples, and the indicators used to show importance are not limited to these.
[0043] Next, the process for determining importance will be explained. The determination unit 12 accepts user input specifying the importance of each pair as training data. The determination unit 12 then determines that the importance specified by the user input is the importance of each pair.
[0044] Here, we will describe an example of a user interface (UI) screen that accepts user input to specify the importance level. Note that this UI screen is merely an example and is not limited to this example.
[0045] The determination unit 12 outputs a UI screen to the user, for example, as shown in Figure 5. This output is realized via an output device such as a display or projection device. The determination unit 12 can then receive user input specifying the importance level of each pair via this UI screen.
[0046] The UI screen shown in Figure 5 includes fields for "before editing," "after editing," and "importance."
[0047] The "Edited" column displays the speech content text data output from the speech recognition unit 11. That is, if user input has been used to modify the recognition result (text data) output from the speech recognition engine, the modified text data is displayed. Conversely, if no user input has been used to modify the recognition result (text data) output from the speech recognition engine, the original recognition result is displayed. Each text data is linked to the date and time of utterance. The date and time of utterance are identified, for example, based on the time code or timestamp of the input audio data. In the diagram, one pair of speech content text data is displayed per row.
[0048] The "Before Editing" column displays the recognition results (text data) output by the speech recognition engine. Each text data entry is linked to the date and time of speaking. The date and time of speaking are identified, for example, based on the time code or timestamp of the input audio data.
[0049] The "Importance" column displays the importance level specified by the user. The number of filled-in black stars indicates the importance level. In this example, importance is shown with four values from 0 to 3. A higher value indicates higher importance. If there are 3 filled-in black stars, the importance level is 3. If there are 0 filled-in black stars, the importance level is 0. Also, if no stars are displayed, the importance level is 0.
[0050] The determination unit 12 displays three stars only for pairs of text that have been modified based on user input in the importance column of the illustrated UI screen. The determination unit 12 accepts user input to either fill in these stars with black or leave them white. The determination unit 12 does not display three stars for text data that has not been modified based on user input in the importance column of the illustrated UI screen. Furthermore, the determination unit 12 does not accept user input to specify the importance of such text data. The determination unit 12 determines the importance of such text data to be 0.
[0051] The UI screen shown in the illustration displays data for multiple pairs in a list, and allows users to specify the importance level for multiple pairs at once. Alternatively, the UI screen could be configured to display data for each pair and accept input to specify the importance level.
[0052] In this way, the determination unit 12 can display the results of the speech recognition process and the corrections based on user input, and can accept user input specifying the importance level via a UI screen that accepts user input specifying the importance level. The following embodiment will describe other processes for determining importance.
[0053] Returning to Figure 1, the generation unit 13 generates training data for important pairs, which are pairs whose importance satisfies predetermined conditions, using the first method. The first method is the method described above using Figure 2. Then, the generation unit 13 generates training data for unimportant pairs, which are pairs whose importance does not satisfy predetermined conditions, using the second method. The second method is the method described above using Figure 3.
[0054] The "predetermined conditions" are the criteria for determining each pair as important, and are defined in advance using importance levels. For example, if importance is represented by two values, "important" and "unimportant," the predetermined condition is "importance is 'important'." Alternatively, if importance is represented by four values from 0 to 3 (higher values indicate greater importance), the predetermined condition is "importance is M or greater," where M is one of 1, 2, or 3.
[0055] Next, we will describe the first and second methods for generating training data.
[0056] -First Method- As shown in Figure 2, the generation unit 13 processes the input audio data of important pairs to generate at least one processed audio data. The generation unit 13 then generates training data by combining the input audio data of important pairs with the utterance content text data (labeled "text data" in the figure). The generation unit 13 also generates additional training data by combining the utterance content text data of important pairs with at least one processed audio data for each pair.
[0057] The generation of processed audio data is achieved by at least one of the following processes: speech rate conversion, voice quality conversion, noise overlay, and sound quality conversion.
[0058] "Speech speed conversion processing" is a process that speeds up or slows down the speaking speed of an audio file.
[0059] "Voice quality conversion processing" is a process that transforms the voice quality of a speaker. For example, it can convert the voice quality of input audio data to that of a different gender, age, etc.
[0060] "Noise overlay processing" is a process that overlays arbitrary noise onto input audio data.
[0061] "Audio quality conversion processing" is the process of converting the audio quality. For example, it can convert the audio quality of input audio data to telephone audio, microphone audio, etc.
[0062] The content of the processed audio data generated by this process is the same as the content of the input audio data.
[0063] -Second Method- As shown in Figure 3, the generation unit 13 generates training data by combining the input audio data and utterance content text data (labeled "text data" in the figure) of non-important pairs. Note that the generation unit 13 does not generate processed audio data or additional training data including processed audio data for non-important pairs.
[0064] The generation unit 13 can send the generated training data to the training server. The training server is a server that trains the speech recognition engine described above. The generation unit 13 may send the training data generated by real-time processing to the training server. Alternatively, the generation unit 13 may send the training data generated by batch processing to the training server.
[0065] In addition, the generation unit 13 may store the generated learning data in a predetermined storage device. This storage device may be located within the processing unit 10, or it may be located within an external device configured to communicate with the processing unit 10.
[0066] Next, an example of the processing flow of the processing device 10 will be explained using the flowcharts in Figures 6 and 7. The processing shown in Figure 6 is performed by the speech recognition unit 11. The processing shown in Figure 7 is performed by the determination unit 12 and the generation unit 13.
[0067] First, let's explain the processing flow shown in Figure 6.
[0068] The processing unit 10 is in a state of waiting to acquire input audio data (S10). When the processing unit 10 acquires input audio data (Yes in S10), it performs speech recognition processing on the input audio data (S11) and outputs the recognition result to the user (S12). The user can either input a correction to the outputted recognition result or input an input to confirm it as is without correction.
[0069] If the processor receives input to correct the output recognition result (Yes in S13), the processor 10 corrects the recognition result based on that input (S14). Then, the processor 10 outputs the corrected recognition result as speech content text data that shows the content of the input voice data (S15).
[0070] On the other hand, if no input is accepted to correct the output recognition result (No. in S13), the processing unit 10 outputs the output recognition result as speech content text data indicating the content of the input voice data (S15).
[0071] Next, we will explain the processing flow shown in Figure 7.
[0072] The processing unit 10 is in a state of waiting to acquire pairs of input speech data and spoken content text data (S20). When the processing unit 10 acquires a pair of input speech data and spoken content text data (Yes in S20), it determines the importance of that pair as training data (S21).
[0073] For example, the processing unit 10 may accept user input specifying the importance of the pair. The processing unit 10 may then determine the importance of the pair to be the importance specified by the user input. The input specifying the importance of each pair may be performed in advance, and the results may be stored in the processing unit 10. Then, in S21, the processing unit 10 may read the importance of the pair to be processed from the importance of each pair that has been stored in advance. In addition, the processing unit 10 may accept input from the user specifying the importance of the pair to be processed in S21.
[0074] If the importance level meets a predetermined condition (Yes in S22), the processing unit 10 processes the input audio data to generate at least one processed audio data (S23). Then, the processing unit 10 sends the learning data, which is a combination of the input audio data and the at least one processed audio data, and the utterance content text data, to the learning server (S24).
[0075] On the other hand, if the importance does not meet the predetermined conditions (No. in S22), the processing unit 10 sends learning data, which is a combination of the input voice data and the spoken content text data, to the learning server (S25).
[0076] In addition, in steps S24 and S25, the processing unit 10 may store the learning data in a predetermined storage device instead of sending the learning data to the learning server.
[0077] "Effects and Effects" In the second embodiment, the processing unit 10 determines the importance of each pair of input speech data and utterance content text data as training data. For important pairs that are important as training data, the processing unit 10 generates processed speech data from the input speech data and generates additional training data that includes each of the generated processed speech data. On the other hand, for unimportant pairs that are not important as training data, the processing unit 10 does not generate processed speech data or additional training data that includes the generated processed speech data. As a result, the training data for important pairs is greater than the training data for unimportant pairs by the amount of additional training data.
[0078] Thus, the processing device 10 according to the second embodiment increases the number of training data related to important pairs that are important as training data, thereby enabling learning that better reflects the content of important training data. Furthermore, since the processing device 10 according to the second embodiment increases the number of training data related to important pairs using audio data processing, the workload on the user to increase the training data can be reduced.
[0079] Furthermore, the processing unit 10 of the second embodiment can generate processed speech data using at least one of the following processes: speech rate conversion, voice quality conversion, noise superposition, and sound quality conversion. In this case, the utterance content of the processed speech data is the same as the utterance content of the input speech data. Therefore, the utterance content text data representing the utterance content of the input speech data can be used as is as text data to be combined with each processed speech data in the generation of training data. Thus, the processing unit 10 of the second embodiment can generate multiple additional training data using the utterance content text data as is, thereby efficiently increasing the amount of training data.
[0080] <Third Embodiment> The processing unit 10 of the third embodiment determines the importance of each pair of input speech data and utterance content text data as training data using a method not described in the first and second embodiments. Specifically, the processing unit 10 of the third embodiment determines the importance based on the content of the corrections to the recognition results output from the speech recognition engine. This will be explained in detail below.
[0081] The determination unit 12 determines the importance based on the content of the modifications to the recognition results output from the speech recognition engine. "Content of the modifications to the recognition results output from the speech recognition engine" means whether or not modifications based on user input have been made, and if so, the specific content of those modifications. As described in the second embodiment, the speech recognition unit 11 can make modifications to the recognition results (text data) output from the speech recognition engine based on user input.
[0082] The determination unit 12 determines the importance of each pair based on the value of at least one of the following items 1 to 9.
[0083] (Item 1) Whether or not user input was accepted to correct the recognition result output from the speech recognition engine.
[0084] In this case, the item values are, for example, P1, which corresponds to "User input to correct the recognition result output from the speech recognition engine was received," and P2, which corresponds to "User input to correct the recognition result output from the speech recognition engine was not received." P1 and P2 are, for example, predetermined numerical values.
[0085] The determination unit 12 assigns a higher importance to pairs that have received user input than to pairs that have not received user input. In other words, the importance of pairs that have received user input to correct the recognition result output from the speech recognition engine will be higher than the importance of pairs that have not received user input.
[0086] (Item 2) The number of characters modified by user input.
[0087] In this case, the item value is, for example, the number of characters. The determination unit 12 assigns a higher importance the more characters that have been modified by user input.
[0088] (Item 3) Whether or not the words corrected by user input are included in the already trained data.
[0089] In this case, the item values are, for example, Q1, which corresponds to "The word modified by user input is included in the previously trained data," and Q2, which corresponds to "The word modified by user input is not included in the previously trained data." Alternatively, the item values may also be, for example, Q1, which corresponds to "All of the words modified by user input are included in the previously trained data," and Q2, which corresponds to "At least one of the words modified by user input is not included in the previously trained data." Q1 and Q2 are, for example, predetermined numerical values.
[0090] The judgment unit 12 increases the importance of a word if it is not included in the previously learned training data.
[0091] If this item is selected, the determination unit 12 performs morphological analysis on the speech content text data and extracts the words contained in the speech content text data. Then, the determination unit 12 determines whether or not each word has been modified by user input and extracts the modified words.
[0092] Subsequently, the determination unit 12 determines whether the extracted modified word is included in the previously learned training data. For example, the previously learned training data is stored in the processing unit 10 or in the storage device of an external device configured to communicate with the processing unit 10. The determination unit 12 determines whether the extracted modified word is included in the spoken content text data of the previously learned training data stored in this storage device.
[0093] As another example, a database of words included in the spoken content text data of pre-trained learning data may be generated in advance. The determination unit 12 may then determine whether or not the extracted modified words are included in this database.
[0094] (Item 4) The number of trained data points that include words modified by user input.
[0095] In this case, the item value is, for example, the number of trained data points that include words modified by user input.
[0096] The determination unit 12 assigns a higher importance to words that have been modified by user input, as the number of trained data points containing those words decreases.
[0097] The determination unit 12 can perform the determination using the same method as described in item 3. For example, a database of words included in the spoken content text data of the learned training data is generated in advance. In this database, the number of learned training data containing each word is registered in association with each word. The determination unit 12 determines whether or not the extracted modified word is included in this database. If the extracted modified word is included in the database, the determination unit 12 reads the "number of learned training data" registered in association with that word.
[0098] (Item 5) The number of words that have been modified by user input but are not included in the previously trained data.
[0099] In this case, the item value is, for example, the number of words that have been modified by user input and are not included in the previously trained data.
[0100] When correcting the recognition results output from the speech recognition engine, multiple words may be modified. Therefore, the determination unit 12 determines the importance of the words that have been modified by user input but are not included in the previously learned training data.
[0101] The determination unit 12 assigns a higher importance to words that have been modified by user input but are not included in the previously learned training data, based on the number of such words.
[0102] The determination unit 12 can determine whether each word modified by user input is included in the previously learned training data, using a method similar to the method described in item 3. Based on the determination result, the determination unit 12 can then count the number of words modified by user input that are not included in the previously learned training data.
[0103] (Item 6) Whether the words corrected by user input are included in the previously trained additional training data.
[0104] In this case, the item values are, for example, R1 corresponding to "the words corrected by user input are included in the previously trained additional training data" and R2 corresponding to "the words corrected by user input are not included in the previously trained additional training data." Alternatively, the item values could be, for example, R1 corresponding to "all of the words corrected by user input are included in the previously trained additional training data" and R2 corresponding to "at least one of the words corrected by user input is not included in the previously trained additional training data." R1 and R2 are, for example, predetermined numerical values.
[0105] Item 3, mentioned above, concerns whether the words modified by user input are included in the "pre-trained training data." In contrast, item 6 concerns whether the words modified by user input are included in the "pre-trained additional training data." The "pre-trained training data" is a concept that partially includes the "pre-trained additional training data."
[0106] The additional training data is training data generated in relation to important pairs generated by the generation unit 13, as described in the second embodiment.
[0107] For example, the learned additional training data is stored in the processing unit 10, or in the storage device of an external device configured to communicate with the processing unit 10. The determination unit 12 determines whether or not the spoken content text data of the learned additional training data stored in this storage device contains words that have been modified by user input.
[0108] As another example, a database of words included in the speech content text data of pre-trained additional training data may be generated in advance. The determination unit 12 may then determine whether or not the words modified by user input are included in this database.
[0109] (Item 7) The number of additional trained data points that include words corrected by user input.
[0110] In this case, the item value is, for example, the number of additional trained data points that include words modified by user input.
[0111] The determination unit 12 assigns a higher importance to words that have been modified by user input, the fewer the number of additional training data points that include the modified word.
[0112] The determination unit 12 can perform the determination using the same method as described in item 6. For example, a database of words included in the spoken content text data of the previously learned additional training data is generated in advance. In this database, the number of previously learned additional training data containing each word is registered in association with that word. The determination unit 12 determines whether or not the word modified by user input is included in this database. If the word modified by user input is included in the database, the determination unit 12 reads the "number of previously learned additional training data" registered in association with that word.
[0113] (Item 8) The number of words that have been modified by user input but are not included in the previously trained additional training data.
[0114] In this case, the item value is, for example, the number of words that have been modified by user input and are not included in the previously trained additional training data.
[0115] When correcting the recognition results output from the speech recognition engine, multiple words may be modified. Therefore, the determination unit 12 determines the importance of the words that have been modified by user input but are not included in the previously learned additional training data.
[0116] The determination unit 12 assigns a higher importance to words that have been modified by user input and are not included in the previously learned additional training data, based on the number of such words.
[0117] The determination unit 12 can determine, using a method similar to the one described in item 6, whether or not each word modified by user input is included in the previously learned additional training data. Based on the determination result, the determination unit 12 can then count the number of words modified by user input that are not included in the previously learned additional training data.
[0118] (Item 9) Whether the word modified by user input is a designated noun.
[0119] In this case, the item values are, for example, S1 corresponding to "the word modified by user input is a specified noun" and S2 corresponding to "the word modified by user input is not a specified noun". Alternatively, the item values may also be, for example, S1 corresponding to "at least one of the words modified by user input is a specified noun" and S2 corresponding to "none of the words modified by user input are specified nouns". S1 and S2 are, for example, predetermined numerical values.
[0120] A "designated noun" is an important noun specified by the user. Users can set proper nouns such as place names and personal names, or other arbitrary nouns such as technical terms, as designated nouns.
[0121] The determination unit 12 assigns a higher importance to the word modified by user input if it is a predetermined noun, compared to when it is not.
[0122] The determination unit 12 can determine whether a word modified by user input is a proper noun, for example, using well-known techniques such as morphological analysis. Alternatively, a dictionary data containing predetermined nouns may be generated in advance and stored in the processing unit 10. The determination unit 12 may then determine whether the word modified by user input is registered in this dictionary data.
[0123] The determination unit 12 can determine the importance of each pair based on at least one item value from items 1 to 9. An importance calculation model is generated in advance and stored in the processing unit 10, taking at least one item value from items 1 to 9 as input and outputting importance. The determination unit 12 uses this importance calculation model to determine the importance. The importance calculation model consists of functions, tables, etc. As an example, a model can be considered in which the sum of the products of the item value of each of the multiple items and the weight coefficient of each item is calculated as a score, and the system determines which of the score ranges corresponding to each of the multiple importance levels this score falls into. However, this is merely an example and is not limited to this.
[0124] Furthermore, the user may set the weight for each of the above items 1 through 9. The importance calculation model may then be modified based on the weights set by the user for each item. In the example model above, the weight coefficients for each item are modified based on the weights set by the user for each item.
[0125] The other configurations of the processing apparatus 10 in the third embodiment are the same as those of the processing apparatus 10 in the first and second embodiments.
[0126] The processing device 10 of the third embodiment achieves the same effects as the processing device 10 of the first and second embodiments. Furthermore, the processing device 10 of the third embodiment can automatically determine the importance of pairs of input voice data and spoken content text data as learning data. As a result, the user does not need to specify the importance for each pair, reducing the user's workload.
[0127] Furthermore, according to the processing device 10 of the third embodiment, the importance can be determined based on the content of the correction to the recognition result output from the speech recognition engine. Specifically, the importance can be determined based on at least one of the above items 1 to 9. As a result, the importance of the training data can be determined with high accuracy.
[0128] <Fourth Embodiment> The processing unit 10 of the fourth embodiment determines the importance of each pair of input speech data and utterance content text data as training data using a method not described in the first to third embodiments. Specifically, the processing unit 10 of the fourth embodiment determines the importance based on at least one of whether or not it has been trained in the past, and the number of times it has been trained in the past. This will be explained in detail below.
[0129] The determination unit 12 determines the importance based on at least one of whether or not it has been learned in the past, and the number of times it has been learned in the past.
[0130] The determination unit 12 determines the importance of each pair based on the value of at least one of the following items 10 and 11.
[0131] (Item 10) Whether or not it has been learned in the past.
[0132] In this case, the item values are, for example, T1, which corresponds to "previously learned," and T2, which corresponds to "not previously learned." T1 and T2 are predetermined numerical values.
[0133] The determination unit 12 can determine whether a pair has been previously trained by comparing the utterance content text data of the pair to be determined with the utterance content text data of previously trained data. There are various conditions for determining whether a pair has been trained, but for example, it may be any of the following.
[0134] • There is pre-trained training data that contains speech content text data that perfectly matches the speech content text data of the pair to be judged. • There is pre-trained training data that contains speech content text data containing a predetermined percentage or more of the words included in the speech content text data of the pair to be judged. • There is pre-trained training data that includes utterance text data containing words modified by the user within the utterance text data of the pair to be judged. • There is additional trained data containing speech content text data that perfectly matches the speech content text data of the pair being judged. • There is additional trained training data that contains speech content text data with a predetermined percentage or more of the words included in the speech content text data of the pair to be judged. • There is additional trained training data that includes utterance text data containing words modified by the user within the utterance text data of the pair being judged.
[0135] (Item 11) The number of times it has been learned in the past.
[0136] In this case, the item value is, for example, the number of past learning sessions. The determination unit 12, for example, compares the utterance content text data of the pair to be judged with the utterance content text data of the previously learned learning data to extract the previously learned learning data that satisfies predetermined conditions between the utterance content text data of the pair to be judged and the previously learned learning data. The determination unit 12 can then set the number of extracted previously learned learning data as the item value.
[0137] The determination unit 12 may, for example, set one of the following numbers as the item value.
[0138] The number of trained data points that have utterance text data that perfectly match the utterance text data of the pair to be judged. The number of trained data sets that contain a predetermined percentage or more of the words included in the speech content text data of the pair to be judged. The number of trained data sets containing speech content text data that includes words modified by the user within the speech content text data of the pair being judged. The number of additional trained data sets that have speech content text data that perfectly match the speech content text data of the pair being judged. The number of additional trained data sets that contain a predetermined percentage or more of the words included in the speech content text data of the pair being judged. The number of additional trained data sets that contain speech content text data with words modified by the user within the speech content text data of the pair being judged.
[0139] The determination unit 12 can determine the importance of each pair based on at least one item value from items 10 and 11. An importance calculation model is generated in advance and stored in the processing unit 10, taking at least one item value from items 10 and 11 as input and outputting importance. The determination unit 12 uses this importance calculation model to determine the importance. The importance calculation model consists of functions, tables, etc. As an example, a model can be considered in which the sum of the products of each item value and the weight coefficient of each item is calculated as a score, and the system determines which of the score ranges corresponding to each of the multiple importance levels this score falls into. However, this is merely an example and is not limited to this.
[0140] Furthermore, the user may set the weights for each of the above items 10 and 11. The importance calculation model may then be modified based on the weights of each item set by the user. In the example model above, the weight coefficients for each item are modified based on the weights of each item set by the user.
[0141] Furthermore, the determination unit 12 may determine the importance of each pair based on the value of at least one of the items 1 to 11 above, using a method similar to the method described above. Items 1 to 9 are the items described in the third embodiment.
[0142] The other configurations of the processing apparatus 10 in the fourth embodiment are the same as those of the processing apparatus 10 in the first to third embodiments.
[0143] The processing device 10 of the fourth embodiment achieves the same effects as the processing device 10 of the first to third embodiments. Furthermore, the processing device 10 of the fourth embodiment can automatically determine the importance of pairs of input voice data and spoken content text data as learning data. As a result, the user does not need to specify the importance for each pair, thereby reducing the user's workload.
[0144] Furthermore, according to the processing device 10 of the fourth embodiment, importance can be determined based on whether or not it has been previously learned and the number of times it has been learned. Specifically, importance can be determined based on at least one of the above items 10 and 11. As a result, the importance of the learning data can be determined with high accuracy.
[0145] <Fifth Embodiment> The processing device 10 of the fifth embodiment determines the importance of each pair of input speech data and utterance content text data as training data using a method not described in the first to fourth embodiments. Specifically, the processing device 10 of the fifth embodiment determines the importance based on at least one of the following: whether or not the utterance content text data contains a predetermined noun, and the number of predetermined nouns it contains. This will be explained in detail below.
[0146] The determination unit 12 determines the importance based on at least one of the following: whether or not the utterance content text data contains a predetermined noun, and the number of predetermined nouns contained in the utterance content text data.
[0147] The determination unit 12 determines the importance of each pair based on the value of at least one of the following items 12 and 13.
[0148] (Item 12) Whether the utterance text data contains the specified noun.
[0149] In this case, the item values are, for example, U1, which corresponds to "the utterance text data contains the specified noun," and U2, which corresponds to "the utterance text data does not contain the specified noun." U1 and U2 are predetermined numerical values.
[0150] A "designated noun" is an important noun specified by the user. Users can set proper nouns such as place names and personal names, or other arbitrary nouns such as technical terms, as designated nouns.
[0151] The determination unit 12 extracts nouns from the speech content text data using, for example, well-known techniques such as morphological analysis, and determines whether the extracted nouns are proper nouns. Alternatively, a dictionary data containing predetermined nouns may be generated in advance and stored in the processing unit 10. The determination unit 12 may then determine whether the noun extracted from the speech content text data is registered in the dictionary data.
[0152] (Item 13) The number of specified nouns contained in the utterance text data.
[0153] In this case, the item value is, for example, the number of predetermined nouns contained in the utterance content text data. The determination unit 12 can count the number of predetermined nouns contained in the utterance content text data using the method described in item 12.
[0154] The determination unit 12 can determine the importance of each pair based on at least one item value from items 12 and 13. An importance calculation model is generated in advance, taking at least one item value from items 12 and 13 as input and outputting importance, and is stored in the processing unit 10. The determination unit 12 uses this importance calculation model to determine the importance. The importance calculation model consists of functions, tables, etc. As an example, a model can be considered in which the sum of the products of the item value of each of the multiple items and the weight coefficient of each item is calculated as a score, and the system determines which of the score ranges corresponding to each of the multiple importance levels this score falls into. However, this is merely an example and is not limited to this.
[0155] Furthermore, the user may set the weights for each of the above items 12 and 13. The importance calculation model may then be modified based on the weights of each item set by the user. In the example model above, the weight coefficients for each item are modified based on the weights of each item set by the user.
[0156] Furthermore, the determination unit 12 may determine the importance of each pair based on the value of at least one of the items 1 to 13 above, using a method similar to the method described above. Items 1 to 9 are the items described in the third embodiment. Items 10 and 11 are the items described in the fourth embodiment.
[0157] The other configurations of the processing apparatus 10 in the fifth embodiment are the same as those of the processing apparatus 10 in the first to fourth embodiments.
[0158] The processing device 10 of the fifth embodiment achieves the same effects as the processing device 10 of the first to fourth embodiments. Furthermore, the processing device 10 of the fifth embodiment can automatically determine the importance of pairs of input voice data and spoken content text data as learning data. As a result, the user does not need to specify the importance for each pair, reducing the user's workload.
[0159] Furthermore, according to the processing device 10 of the fifth embodiment, importance can be determined based on at least one of the following: whether or not the utterance content text data contains a predetermined noun, and the number of predetermined nouns it contains. Specifically, importance can be determined based on at least one of the above items 12 and 13. As a result, the importance of the data as training data can be determined with high accuracy.
[0160] <Sixth Embodiment> The processing device 10 of the sixth embodiment differs from the first to fifth embodiments in that it includes means for training the speech recognition engine based on the training data generated by the generation unit 13. In the first to fifth embodiments, an external device (learning server) different from the processing device 10 is provided for training the speech recognition engine based on the training data generated by the generation unit 13.
[0161] Figure 8 shows an example of a functional block diagram of the processing device 10 of the sixth embodiment. As shown in the figure, the processing device 10 of the sixth embodiment includes a speech recognition unit 11, a determination unit 12, a generation unit 13, and a learning unit 14.
[0162] The learning unit 14 learns the speech recognition engine based on the training data generated by the generation unit 13. The speech recognition engine to be learned is, for example, the speech recognition engine used in speech recognition processing by the speech recognition unit 11. The learning unit 14 can learn the speech recognition engine based on any widely known technology.
[0163] The other configurations of the processing apparatus 10 in the sixth embodiment are the same as those of the processing apparatus 10 in the first to fifth embodiments.
[0164] According to the processing apparatus 10 of the sixth embodiment, the same effects and advantages as those of the processing apparatus 10 of the first to fifth embodiments are achieved.
[0165] <Variation> Modified examples applicable to the processing apparatus 10 of the first to sixth embodiments will be described.
[0166] "Variation 1" The generation unit 13 can generate a number of additional training data points corresponding to their importance. The higher the importance, the more additional training data points the generation unit 13 generates. The number of additional training data points to be generated for each level of importance is predetermined. The generation unit 13 then generates a number of additional training data points according to the importance level, following these rules.
[0167] According to this modified version, the same effects and advantages as the processing device 10 of the first to sixth embodiments are achieved. Furthermore, according to this modified version, it is possible to generate more training data for important pairs with higher importance. As a result, it becomes possible to perform training that better reflects the content of important pairs with higher importance.
[0168] "Variation 2" In the first to sixth embodiments, the generation unit 13 generated training data for non-important pairs by combining the input audio data and spoken content text data of the non-important pairs using the method described with reference to Figure 3.
[0169] As a variation of this process, the generation unit 13 does not need to perform the process of generating training data for non-important pairs. In other words, the generation unit 13 does not need to generate training data that combines the input audio data and spoken content text data of non-important pairs.
[0170] According to this modified example, the same effects and advantages as those of the processing apparatus 10 in the first to sixth embodiments are achieved.
[0171] <Usage scene> The processing unit 10 of the first to sixth embodiments can be used, for example, in a call center or by a police department that receives reports by telephone. For example, audio data recorded from a phone call is input to the processing unit 10 as input audio data. The operator receiving the call checks the results of the speech recognition processing output from the processing unit 10 and makes corrections as necessary. The operator can also save the corrected text data of the spoken content as a record of the call.
[0172] According to the processing unit 10, training data for training a speech recognition engine used in situations such as the one described above can be generated.
[0173] The embodiments of the present invention have been described above with reference to the drawings, but these are illustrative examples of the present invention, and various other configurations can be adopted. The configurations of the embodiments described above may be combined with each other, or some configurations may be replaced with other configurations. Furthermore, the configurations of the embodiments described above may be modified in various ways without departing from the spirit of the invention. In addition, the configurations and processes disclosed in each of the embodiments and modifications described above may be combined with each other.
[0174] Furthermore, the flowcharts used in the above description show multiple steps (processes) in sequence. However, the execution order of the steps performed in each embodiment is not limited to the order in which they are described. In each embodiment, the order of the illustrated steps can be changed to the extent that it does not impede the content. Also, the above embodiments can be combined to the extent that their content is not contradictory.
[0175] Some or all of the above embodiments may also be described as follows, but are not limited to the following. 1. A speech recognition means that performs speech recognition processing on input audio data and outputs text data indicating the content of the utterance of the input audio data, A determination means for determining the importance of each pair of input audio data and text data as learning data, A generation means that processes the input audio data of the important pair, which is the pair whose importance satisfies the predetermined conditions, to generate at least one processed audio data, and generates additional learning data by combining the processed audio data and the text data, A processing device. 2. The generating means is The processing apparatus according to claim 1, which does not generate the processed audio data and the additional learning data including the processed audio data for non-important pairs whose importance does not satisfy the predetermined conditions. 3. The speech recognition means is: The system accepts user input to correct the results of the aforementioned speech recognition processing. The determination means is, Whether or not the aforementioned user input was received, The number of characters modified by the user input mentioned above, Whether or not the word modified by the user input is included in the previously learned training data, The number of trained data sets that include the word corrected by the user input, The number of words corrected by the user input that are not included in the previously learned training data, Whether or not the word modified by the user input is included in the previously learned additional training data, The number of the additional training data that have been trained and include the word modified by the user input, The number of words modified by the user input that are not included in the previously learned additional training data, and Whether the word modified by the user input is a specified noun or not, The processing apparatus according to 1 or 2, which determines the importance based on at least one of the following. 4. The determination means is, A processing device according to any one of 1 to 3 that determines the importance based on whether or not it has been learned in the past, and at least one of the number of times it has been learned in the past. 5. The determination means is, A processing apparatus according to any one of 1 to 4 that determines the importance based on whether the text data contains a predetermined noun and the number of predetermined nouns contained in the text data. 6. The determination means is, A processing device according to any one of 1 to 5, which accepts user input specifying the importance of each pair as training data. 7. The voice recognition means is The system accepts user input to correct the results of the aforementioned speech recognition processing. The determination means is, The processing device according to 6, which displays the results of the speech recognition processing and the corrections, and accepts user input to specify the importance level via a UI screen that accepts user input to specify the importance level. 8. The generating means is An apparatus according to any one of 1 to 7, which performs at least one of the following processes on the input audio data: speech rate conversion process, voice quality conversion process, noise superposition process, and sound quality conversion process, to generate the processed audio data. 9. One or more computers, Speech recognition processing is performed on the input audio data, and text data representing the content of the utterance from the input audio data is output. For each pair of input audio data and text data, the importance of the data as training data is determined. A processing method that processes the input audio data of the important pair, which is the pair whose importance satisfies a predetermined condition, to generate at least one processed audio data, and generates additional learning data by combining the processed audio data and the text data. 10. The one or more computers: The processing method according to 9, wherein for unimportant pairs whose importance does not satisfy the predetermined conditions, the processed audio data is not generated, and the additional learning data including the processed audio data is not generated. 11. The one or more computers mentioned above, The system accepts user input to correct the results of the aforementioned speech recognition processing. Whether or not the aforementioned user input was received, The number of characters modified by the user input mentioned above, Whether or not the word modified by the user input is included in the previously learned training data, The number of trained data sets that include the word corrected by the user input, The number of words corrected by the user input that are not included in the previously learned training data, Whether or not the word modified by the user input is included in the previously learned additional training data, The number of the additional training data that have been trained and include the word modified by the user input, The number of words modified by the user input that are not included in the previously learned additional training data, and Whether the word modified by the user input is a specified noun or not, The processing method according to 9 or 10, which determines the importance based on at least one of the following. 12. The one or more computers mentioned above, A processing method according to any one of 9 to 11, which determines the importance based on whether or not it has been learned in the past, and at least one of the number of times it has been learned in the past. 13. The one or more computers mentioned above, A processing method according to any one of 9 to 12, which determines the importance based on whether the text data contains a predetermined noun and the number of predetermined nouns contained in the text data. 14. The one or more computers mentioned above, A processing method according to any one of 9 to 13, which accepts user input specifying the importance of each pair as training data. 15. The one or more computers: The system accepts user input to correct the results of the aforementioned speech recognition processing. The processing method according to 14, wherein the results of the speech recognition processing and the corrections are displayed, and the user input specifying the importance is received via a UI screen that accepts user input specifying the importance. 16. The one or more computers: A processing method according to any one of 9 to 15, which involves performing at least one of the following processes on the input audio data: speech rate conversion, voice quality conversion, noise overlay, and sound quality conversion, in order to generate the processed audio data. 17. Computers, A speech recognition means that performs speech recognition processing on input speech data and outputs text data indicating the content of the speech in the input speech data. For each pair of input audio data and text data, a determination means for determining the importance of the data as learning data. A generation means that processes the input audio data of the important pair, which is the pair whose importance satisfies the predetermined conditions, to generate at least one processed audio data, and generates additional learning data by combining the processed audio data and the text data. A program that makes it function as such. 18. The generating means is The program according to 17, which does not generate the processed audio data or the additional learning data including the processed audio data for non-important pairs whose importance does not satisfy the predetermined conditions. 19. The speech recognition means is The system accepts user input to correct the results of the aforementioned speech recognition processing. The determination means is, Whether or not the aforementioned user input was received, The number of characters modified by the user input mentioned above, Whether or not the word modified by the user input is included in the previously learned training data, The number of trained data sets that include the word corrected by the user input, The number of words corrected by the user input that are not included in the previously learned training data, Whether or not the word modified by the user input is included in the previously learned additional training data, The number of the additional training data that have been trained and include the word modified by the user input, The number of words modified by the user input that are not included in the previously learned additional training data, and Whether the word modified by the user input is a specified noun or not, A program according to 17 or 18 that determines the importance based on at least one of the following. 20. The determination means is, A program according to any one of 17 to 19 that determines the importance based on whether or not it has been learned in the past, and at least one of the number of times it has been learned in the past. 21. The determination means is, A program according to any one of 17 to 20 that determines the importance based on whether the text data contains a predetermined noun and the number of predetermined nouns contained in the text data. 22. The determination means is, A program according to any one of 17 to 21 that accepts user input specifying the importance of each pair as training data. 23. The aforementioned speech recognition means is The system accepts user input to correct the results of the aforementioned speech recognition processing. The determination means is, The program according to 22, which displays the results of the speech recognition processing and the corrections, and accepts user input to specify the importance level via a UI screen that accepts user input to specify the importance level. 24. The generating means is A program according to any one of 17 to 23, which performs at least one of the following processes on the input audio data: speech rate conversion, voice quality conversion, noise overlay, and sound quality conversion, to generate the processed audio data.
[0176] This application claims priority based on Japanese Patent Application No. 2023-078497, filed on 11 May 2023, and incorporates all of its disclosures herein. [Explanation of Symbols]
[0177] 10 Processing Unit 11. Voice Recognition Unit 12 Judgment section 13 Generation part 14. Learning Department 1A Processor 2A Memory 3A input / output I / F 4A Peripheral Circuits 5A bus
Claims
1. A speech recognition means that performs speech recognition processing on input speech data and outputs text data indicating the content of the speech in the input speech data, A determination means for determining the importance of each pair of input audio data and text data as learning data, A generation means that processes the input audio data of the important pair, which is the pair whose importance satisfies the predetermined conditions, to generate at least one processed audio data, and generates additional learning data by combining the processed audio data and the text data, A processing device.
2. The generating means is The processing apparatus according to claim 1, wherein for unimportant pairs whose importance does not satisfy the predetermined conditions, the processing audio data is not generated, and the additional learning data including the processing audio data is not generated.
3. The aforementioned speech recognition means is The system accepts user input to correct the results of the aforementioned speech recognition processing. The determination means is, Whether or not the aforementioned user input was received, The number of characters modified by the user input mentioned above, Whether or not the word modified by the user input is included in the previously learned training data, The number of trained data sets that include the word corrected by the user input, The number of words corrected by the user input that are not included in the previously learned training data, Whether or not the word modified by the user input is included in the previously learned additional training data, The number of the additional training data that have been trained and include the word modified by the user input, The number of words modified by the user input that are not included in the previously learned additional training data, and Whether the word modified by the user input is a specified noun or not, The apparatus according to claim 1 or 2, which determines the importance based on at least one of the following.
4. The determination means is, The processing apparatus according to claim 1 or 2, which determines the importance based on whether or not it has been learned in the past, and at least one of the number of times it has been learned in the past.
5. The determination means is, The apparatus for determining importance according to claim 1 or 2, based on whether the text data contains a predetermined noun and the number of predetermined nouns contained in the text data.
6. The determination means is, The processing apparatus according to claim 1 or 2, which accepts user input specifying the importance of each pair as learning data.
7. The aforementioned speech recognition means is The system accepts user input to correct the results of the aforementioned speech recognition processing. The determination means is, The processing apparatus according to claim 6, which displays the results of the speech recognition processing and the corrections, and accepts user input to specify the importance level via a UI (user interface) screen that accepts user input to specify the importance level.
8. The generating means is The processing apparatus according to claim 1 or 2, which performs at least one of the following processes on the input audio data: speech rate conversion process, voice quality conversion process, noise superposition process, and sound quality conversion process, in order to generate the processed audio data.
9. One or more computers, Speech recognition processing is performed on the input audio data, and text data representing the content of the utterance from the input audio data is output. For each pair of input audio data and text data, the importance of the data as training data is determined. A processing method that processes the input audio data of the important pair, which is the pair whose importance satisfies the aforementioned conditions, to generate at least one processed audio data, and generates additional learning data by combining the processed audio data and the text data.
10. Computers, A speech recognition means that performs speech recognition processing on input speech data and outputs text data indicating the content of the speech in the input speech data. For each pair of input audio data and text data, a determination means for determining the importance of the data as learning data. A generation means that processes the input audio data of the important pair, which is the pair whose importance satisfies the predetermined conditions, to generate at least one processed audio data, and generates additional learning data by combining the processed audio data and the text data. A program that makes it function as such.