Information processing system, information processing method, and computer program
The information processing system enhances speech recognition training by converting text data to generate expanded and corrected voice data, addressing limitations in existing systems and improving accuracy.
Patent Information
- Application Number
- JP2023555999
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-10-28
- Publication Date
- 2025-11-26
- Estimated Expiration
- 2041-10-28
AI Technical Summary
Existing speech recognition systems face challenges in accurately training on speech data without corresponding text data, and existing methods for generating pseudo training data or altering text expressions are limited in effectiveness.
An information processing system that acquires first text data, converts it to generate converted text data, and uses this data along with generated voice data to train a voice recognition system, incorporating methods to correct mispronunciations and expand training data.
The system enables more accurate training of speech recognizers by expanding the training data set with converted text data, allowing for automatic correction of mispronunciations and improving recognition accuracy.
Smart Images

Figure 0007775890000001 
Figure 0007775890000002 
Figure 0007775890000003
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to the technical fields of an information processing system, an information processing device, an information processing method, and a recording medium. [Background technology]
[0002] As this type of system, a system that performs training on a speech recognizer is known. For example, Patent Document 1 discloses that when training a speech recognizer using speech data and text data, for text data that does not have corresponding speech data, pseudo training data is generated without speech recognition and training is performed.
[0003] As other related techniques, Patent Document 2 discloses generating a converted utterance sentence by obscuring at least a part of an original utterance sentence. Patent Document 3 discloses replacing a part of text with an alternative expression that is least likely to cause a change in voice quality from a set of alternative expressions. [Prior art documents] [Patent documents]
[0004] [Patent Document 1] Patent Publication No. 2014-074732 [Patent Document 2] Japanese Patent Application Publication No. 2017-208003 [Patent Document 3] International Publication No. 2007 / 010680 Summary of the Invention [Problem to be solved by the invention]
[0005] This disclosure aims to improve upon the techniques disclosed in the prior art documents. [Means for solving the problem]
[0006] One aspect of the information processing system disclosed herein comprises a first text data acquisition means for acquiring first text data, a text data conversion means for converting the first text data to generate converted text data, a converted voice data generation means for generating converted voice data corresponding to the converted text data, and a learning means for training a voice recognition means that uses the first text data and the converted voice data as input and generates text data corresponding to the voice data from the voice data.
[0007] One aspect of the information processing device disclosed herein comprises a first text data acquisition means for acquiring first text data, a text data conversion means for converting the first text data to generate converted text data, a converted voice data generation means for generating converted voice data corresponding to the converted text data, and a learning means for training a voice recognition means that uses the first text data and the converted voice data as input and generates text data corresponding to the voice data from the voice data.
[0008] One aspect of the information processing method disclosed herein is an information processing method executed by at least one computer, which acquires first text data, converts the first text data to generate converted text data, generates converted voice data corresponding to the converted text data, and trains a voice recognition means that uses the first text data and the converted voice data as inputs and generates text data corresponding to the voice data from voice data.
[0009] One aspect of the recording medium of this disclosure has recorded thereon a computer program that causes at least one computer to execute an information processing method, which includes acquiring first text data, converting the first text data to generate converted text data, generating converted voice data corresponding to the converted text data, and training a voice recognition means that uses the first text data and the converted voice data as inputs to generate text data corresponding to the voice data from the voice data. [Brief explanation of the drawings]
[0010] [Figure 1] 1 is a block diagram showing a hardware configuration of an information processing system according to a first embodiment. [Figure 2] 1 is a block diagram showing a functional configuration of an information processing system according to a first embodiment. [Figure 3] 10 is a table showing an example of first text data and converted text data. [Figure 4] 4 is a flowchart showing the flow of operations performed by the information processing system according to the first embodiment. [Figure 5] FIG. 10 is a block diagram showing the functional configuration of an information processing system according to a second embodiment. [Figure 6] 10 is a flowchart showing the flow of operations performed by the information processing system according to the second embodiment. [Figure 7] FIG. 10 is a block diagram showing the functional configuration of an information processing system according to a third embodiment. [Figure 8] FIG. 10 is a block diagram showing the functional configuration of an information processing system according to a fourth embodiment. [Figure 9] FIG. 11 is a block diagram showing the functional configuration of an information processing system according to a fifth embodiment. [Figure 10] 13 is a flowchart showing the flow of a conversion unit learning operation by the information processing system according to the fifth embodiment. [Figure 11] FIG. 13 is a block diagram showing the functional configuration of an information processing system according to a sixth embodiment. [Figure 12] 13 is a flowchart showing the flow of a conversion unit learning operation by the information processing system according to the sixth embodiment. [Figure 13] FIG. 13 is a plan view showing an example of presentation of second text data by the information processing system according to the sixth embodiment. [Figure 14] FIG. 13 is a block diagram showing the functional configuration of an information processing system according to a seventh embodiment. [Figure 15] 13 is a flowchart showing the flow of a conversion unit learning operation by the information processing system according to the seventh embodiment. [Figure 16] FIG. 13 is a block diagram showing the functional configuration of an information processing system according to an eighth embodiment. [Figure 17] FIG. 13 is a block diagram showing the functional configuration of an information processing system according to a ninth embodiment. [Figure 18] 13 is a flowchart showing the flow of a voice recognition operation by the information processing system according to the ninth embodiment. [Figure 19] FIG. 20 is a block diagram showing the functional configuration of an information processing system according to a tenth embodiment. [Figure 20] 13 is a flowchart showing the flow of a voice recognition operation by the information processing system according to the tenth embodiment.
[0011] Hereinafter, embodiments of an information processing system, an information processing device, an information processing method, and a recording medium will be described with reference to the drawings.
[0012] First Embodiment An information processing system according to a first embodiment will be described with reference to FIGS.
[0013] (Hardware configuration) First, the hardware configuration of the information processing system according to the first embodiment will be described with reference to Fig. 1. Fig. 1 is a block diagram showing the hardware configuration of the information processing system according to the first embodiment.
[0014] 1, an information processing system 10 according to the first embodiment includes a processor 11, a RAM (Random Access Memory) 12, a ROM (Read Only Memory) 13, and a storage device 14. The information processing system 10 may further include an input device 15 and an output device 16. The processor 11, RAM 12, ROM 13, storage device 14, input device 15, and output device 16 are connected via a data bus 17.
[0015] The processor 11 loads a computer program. For example, the processor 11 is configured to load a computer program stored in at least one of the RAM 12, the ROM 13, and the storage device 14. Alternatively, the processor 11 may load a computer program stored in a computer-readable storage medium using a storage medium reading device (not shown). The processor 11 may acquire (i.e., load) the computer program from a device (not shown) located outside the information processing system 10 via a network interface. The processor 11 controls the RAM 12, the storage device 14, the input device 15, and the output device 16 by executing the loaded computer program. In particular, in this embodiment, when the processor 11 executes the loaded computer program, a functional block for training a speech recognizer is realized within the processor 11. That is, the processor 11 may function as a controller that executes each control of the information processing system 10.
[0016] The processor 11 may be configured as, for example, a central processing unit (CPU), a graphics processing unit (GPU), a field-programmable gate array (FPGA), a demand-side platform (DSP), or an application-specific integrated circuit (ASIC). The processor 11 may be configured as one of these, or may be configured to use multiple processors in parallel.
[0017] The RAM 12 temporarily stores computer programs executed by the processor 11. The RAM 12 temporarily stores data that is temporarily used by the processor 11 while the processor 11 is executing the computer programs. The RAM 12 may be, for example, a D-RAM (Dynamic RAM).
[0018] The ROM 13 stores computer programs executed by the processor 11. The ROM 13 may also store fixed data. The ROM 13 may be, for example, a programmable ROM (P-ROM).
[0019] The storage device 14 stores data that is to be saved long-term by the information processing system 10. The storage device 14 may operate as a temporary storage device for the processor 11. The storage device 14 may include, for example, at least one of a hard disk device, a magneto-optical disk device, an SSD (Solid State Drive), and a disk array device.
[0020] The input device 15 is a device that receives input instructions from a user of the information processing system 10. The input device 15 may include, for example, at least one of a keyboard, a mouse, and a touch panel. The input device 15 may be configured as a mobile terminal such as a smartphone or a tablet.
[0021] The output device 16 is a device that outputs information related to the information processing system 10 to the outside. For example, the output device 16 may be a display device (e.g., a display) that can display information related to the information processing system 10. The output device 16 may also be a speaker or the like that can output information related to the information processing system 10 as audio. The output device 16 may be configured as a mobile terminal such as a smartphone or a tablet.
[0022] 1 shows an example of information processing system 10 including a plurality of devices, but all or some of the functions may be realized by a single device (information processing device). This information processing device may be configured to include only processor 11, RAM 12, and ROM 13 described above, and the other components (i.e., storage device 14, input device 15, output device 16) may be provided by an external device connected to the information processing device. Furthermore, some of the calculation functions of the information processing device may be realized by an external device (e.g., an external server, a cloud, etc.).
[0023] (Functional configuration) Next, the functional configuration of the information processing system 10 according to the first embodiment will be described with reference to Fig. 2. Fig. 2 is a block diagram showing the functional configuration of the information processing system according to the first embodiment.
[0024] As shown in FIG. 2, the information processing system 10 according to the first embodiment is configured to execute training of a speech recognizer 50. The speech recognizer 50 is a device that generates text data from speech data. The training of the speech recognizer 50 is executed, for example, to generate text data with higher accuracy. The speech recognizer 50 according to this embodiment may also have a function of correcting slip-ups and converting them into text. The training of the speech recognizer 50 may also be to train a conversion model used by the speech recognizer 50 (i.e., a model that converts speech data into text data). Note that the information processing system 10 according to the first embodiment does not include the speech recognizer 50 itself as a component, but may be configured as a system including the speech recognizer 50.
[0025] The information processing system 10 according to the first embodiment is configured to include, as components for realizing its functions, a first text data acquisition unit 110, a text data conversion unit 120, a converted voice data generation unit 130, and a learning unit 140. Each of the first text data acquisition unit 110, the text data conversion unit 120, the converted voice data generation unit 130, and the learning unit 140 may be a processing block realized by, for example, the above-mentioned processor 11 (see FIG. 1).
[0026] The first text data acquisition unit 110 is configured to be able to acquire first text data. The first text data is text data acquired for training a speech recognizer. The first text data may be, for example, data consisting of words only, or text data in sentence format. The first text data acquisition unit 110 may acquire multiple pieces of first text data. Note that the first text data acquisition unit 110 may acquire the first text data by voice input. That is, voice data may be converted into text data and acquired as the first text data.
[0027] The text data conversion unit 120 is configured to be able to convert the first text data acquired by the first text data acquisition unit 110 to generate converted text data. The converted text data is text data in which at least a portion of the first text data has been converted into different characters. The text data conversion unit 120 may generate one converted text data from one piece of first text data, or may generate multiple converted text data from one piece of first text data. A specific method for generating the converted text data will be described in detail in another embodiment described later.
[0028] The converted voice data generation unit 130 is configured to be able to generate converted voice data from the converted text data generated by the text data conversion unit 120. That is, the converted voice data generation unit 130 has a function of converting text data into voice data. Note that, as the method for converting text data into voice data can be appropriately adopted from existing technologies, a detailed description thereof will be omitted here.
[0029] The learning unit 140 is configured to be able to train the speech recognizer 50 using the first text data acquired by the first text data acquisition unit 110 and the converted voice data generated by the converted voice data generation unit 130. That is, the learning unit 140 is configured to perform training using a set of corresponding first text data and converted voice data. The learning unit 140 may perform training using a plurality of first text data and a plurality of converted voice data.
[0030] (Example of converted text data) Next, a specific example of converted text data will be described with reference to Fig. 3. Fig. 3 is a table showing an example of first text data and converted text data.
[0031] As shown in FIG. 3, assume that the first text data acquisition unit 110 acquires the first text data "innovation." In this case, the text data conversion unit 120 may generate converted text data such as "invation," "innoinnovation," and "inoeshow." In this way, the text data conversion unit 120 may generate converted text data as expected mistakes in the first text data. Note that, while an example of generating three converted text data from one piece of first text data is given here, one or two converted text data may be generated, or four or more converted text data may be generated. Furthermore, while the above example illustrates a mistake in speech caused by a person being at a loss for words, converted text data may be generated assuming other mistakes, etc. For example, converted text data may be generated assuming a mistake in speech caused by a misuse of expressions such as "redeem one's honor" or "redeem one's name."
[0032] When the first text data is in the form of a sentence, the text data conversion unit 120 may convert some of the words contained in the sentence to generate converted text data. In other words, the converted text data may be generated by converting only some of the words contained in the sentence, without converting the remaining parts. For example, the text data conversion unit 120 may convert only long words or katakana words among the multiple words contained in the first text data.
[0033] More specifically, for example, if first text data "collect various data to bring about innovation" is acquired, the text data conversion unit 120 may convert only the word "innovation" therein to generate converted text data "collect various data to bring about innovation." The text data conversion unit 120 may also generate converted text by converting multiple words contained in a sentence. For example, the text data conversion unit 120 may convert the words "innovation" and "data" respectively in the first text data "collect various data to bring about innovation" described above to generate converted text data "collect various data to bring about innovation."
[0034] If a word included in the converted text data becomes an existing word, the text data conversion unit 120 may exclude that word (i.e., may not output the word as converted text data). For example, if converted text data "invention" is generated as a result of converting first text data "innovation," that word may not be output as converted text data.
[0035] (Operation flow) Next, the flow of operations performed by the information processing system 10 according to the first embodiment (i.e., operations performed when training the speech recognizer 50) will be described with reference to Fig. 4. Fig. 4 is a flowchart showing the flow of operations performed by the information processing system according to the first embodiment.
[0036] 4, when the information processing system 10 according to the first embodiment operates, the first text data acquisition unit 110 first acquires first text data (step S101). The first text data acquired by the first text data acquisition unit 110 is output to each of the text data conversion unit 120 and the learning unit 140.
[0037] Next, text data conversion unit 120 converts the first text data acquired by first text data acquisition unit 110 to generate converted text data (step S102). The converted text data generated by text data conversion unit 120 is output to converted voice data generation unit 130.
[0038] Next, the converted voice data generation unit 130 generates converted voice data from the converted text data generated by the text data conversion unit 120 (step S103). The converted voice data generated by the converted voice data generation unit 130 is output to the learning unit 140.
[0039] Next, the learning unit 140 performs learning of the speech recognizer 50 using the first text data acquired by the first text data acquisition unit 110 and the converted voice data generated by the converted voice data generation unit 130 (step S104). Note that the above-described series of processes may be repeatedly executed every time the first text data is acquired.
[0040] (Technical Effects) Next, the technical effects obtained by the information processing system 10 according to the first embodiment will be described.
[0041] As described with reference to FIGS. 1 to 4, in the information processing system 10 according to the first embodiment, the speech recognizer 50 is trained using the first text data and the converted speech text data as inputs. In this way, the data used for training can be expanded by converting the text data, thereby enabling more appropriate training. For example, if the converted text data is generated assuming a mispronunciation in the first text data, the speech recognizer 50 can recognize the mispronunciation in the speech data and generate text data. This makes it possible for the speech recognizer 50 to generate text data in which the mispronunciation has been automatically corrected.
[0042] Second Embodiment An information processing system 10 according to the second embodiment will be described with reference to Figures 5 and 6. The second embodiment differs from the first embodiment described above only in some configurations and operations, and other parts may be the same as the first embodiment. Therefore, the following will describe in detail the parts that differ from the first embodiment already described, and will omit explanations of other overlapping parts as appropriate.
[0043] (Functional configuration) First, the functional configuration of the information processing system 10 according to the second embodiment will be described with reference to Fig. 5. Fig. 5 is a block diagram showing the functional configuration of the information processing system according to the second embodiment. Note that in Fig. 5, the same elements as those shown in Fig. 2 are denoted by the same reference numerals.
[0044] As shown in Fig. 5, the information processing system 10 according to the second embodiment is configured to include, as components for realizing its functions, a first text data acquisition unit 110, a text data conversion unit 120, a converted voice data generation unit 130, a learning unit 140, and a first voice data generation unit 150. That is, the information processing system 10 according to the second embodiment further includes a first voice data generation unit 150 in addition to the configuration of the first embodiment already described (see Fig. 2). The first voice data generation unit 150 may be a processing block realized by, for example, the above-mentioned processor 11 (see Fig. 1).
[0045] The first voice data generation unit 150 is configured to be able to generate first voice data from the first text data acquired by the first text data acquisition unit 110. That is, the first voice data generation unit 150 has a function of converting text data into voice data. The first voice data generation unit 150 has the same function as the converted voice data generation unit 130 already described. Therefore, the converted voice data generation unit 130 and the first voice data generation unit 150 may be configured as a single common voice data generation unit. In this case, the voice data generation unit may generate and output converted voice data when converted text data is input, and may generate and output first voice data when first text data is input.
[0046] (Operation flow) Next, the flow of operations performed by the information processing system 10 according to the second embodiment will be described. Fig. 6 is a flowchart showing the flow of operations performed by the information processing system according to the second embodiment. In Fig. 6, the same processes as those shown in Fig. 4 are denoted by the same reference numerals.
[0047] 6, when the information processing system 10 according to the second embodiment operates, the first text data acquisition unit 110 first acquires first text data (step S101). The first text data acquired by the first text data acquisition unit 110 is output to each of the text data conversion unit 120 and the learning unit 140.
[0048] Next, the first voice data generation unit 150 generates first voice data from the first text data acquired by the first text data acquisition unit 110 (step S201). The first voice data generated by the first voice data generation unit 150 is output to the learning unit 140. Note that, although an example is given here in which the first voice data is generated immediately after acquiring the first text data, the first voice data generation unit 150 may generate the first voice data at another timing. For example, the first voice data generation unit 150 may generate the first voice data after the converted text data is generated, or may generate the first voice data after the converted voice data is generated.
[0049] Next, text data conversion unit 120 converts the first text data acquired by first text data acquisition unit 110 to generate converted text data (step S102). The converted text data generated by text data conversion unit 120 is output to converted voice data generation unit 130.
[0050] Next, the converted voice data generation unit 130 generates converted voice data from the converted text data generated by the text data conversion unit 120 (step S103). The converted voice data generated by the converted voice data generation unit 130 is output to the learning unit 140.
[0051] Next, the learning unit 140 performs learning of the speech recognizer 50 using the first text data acquired by the first text data acquisition unit 110, the converted speech data generated by the converted speech data generation unit 130, and the first speech data generated by the first speech data generation unit 150 (step S202). That is, in the second embodiment, in addition to the first text data and the converted speech data, the first speech data (i.e., speech data corresponding to the first text data before conversion) is used for learning the speech recognizer 50.
[0052] (Technical Effects) Next, the technical effects obtained by the information processing system 10 according to the second embodiment will be described.
[0053] 5 and 6, in the information processing system 10 according to the second embodiment, the speech recognizer 50 is trained using the first text data, converted speech data, and the first speech data as inputs. In this way, the speech recognizer 50 can be trained more appropriately than when the first speech data is not used for training (i.e., when training is performed using only the first text data and the converted speech data). Specifically, training can be performed taking into consideration the specific type of speech that the first text data contains, and therefore a more accurate speech recognizer 50 can be realized.
[0054] <Third embodiment> An information processing system 10 according to the third embodiment will be described with reference to Fig. 7. The third embodiment differs from the first and second embodiments only in some configurations and operations, and other parts may be the same as the first and second embodiments. Therefore, the following will describe in detail the parts that differ from the embodiments already described, and will omit a description of other overlapping parts as appropriate.
[0055] (Functional configuration) First, the functional configuration of the information processing system 10 according to the third embodiment will be described with reference to Fig. 7. Fig. 7 is a block diagram showing the functional configuration of the information processing system according to the third embodiment. In Fig. 7, the same elements as those shown in Fig. 2 are denoted by the same reference numerals.
[0056] 7, the information processing system 10 according to the third embodiment is configured to include, as components for realizing its functions, a first text data acquisition unit 110, a text data conversion unit 120, a converted voice data generation unit 130, and a learning unit 140. In particular, the text data conversion unit 120 according to the third embodiment includes a conversion rule storage unit 121. The conversion rule storage unit 121 may be realized by, for example, the above-mentioned storage device 14 (see FIG. 1).
[0057] The conversion rule storage unit 121 is configured to store conversion rules for converting first text data into converted text data. The text data conversion unit 120 according to this embodiment reads out conversion rules stored in the conversion rule storage unit 121 and converts the first text data into converted text data. The conversion rule storage unit 121 may store only one conversion rule, or may store multiple conversion rules. If the conversion rule storage unit 121 stores multiple conversion rules, the text data conversion unit 120 may select one conversion rule from the multiple conversion rules to generate converted text data. In this case, the text data conversion unit 120 may select a conversion rule suitable for the input first text data. Alternatively, the text data conversion unit 120 may generate converted text data using each of the multiple conversion rules. For example, after conversion using a first conversion rule, the text data may be further converted using a second conversion rule.
[0058] The conversion rules stored in the conversion rule storage unit 121 may be configured to be updateable (for example, by adding, modifying, deleting, etc.) as needed. The conversion rules may be updated manually. Alternatively, the conversion rules may be updated mechanically (for example, by machine learning). The conversion rule storage unit 121 may also be configured as a database external to the system. In this case, the text data conversion unit 120 itself does not have the conversion rule storage unit 121, and it is sufficient to read the conversion rules from a database external to the system and generate converted text data.
[0059] (Example of conversion rules) The conversion rules stored in the conversion rule storage unit 121 will be described below with some specific examples.
[0060] The conversion rule may be "remove some characters." In this case, the first text data "innovation" may be converted into converted text data "ivashon," for example. The conversion rule may be "add some characters." In this case, the first text data "innovation" may be converted into converted text data "innonovation," for example. The conversion rule may be "change some characters (for example, replace them with similar sounds)." In this case, the first text data "innovation" may be converted into converted text data "innoshon," for example. The conversion rule may be "repeat the first few characters." In this case, the first text data "innovation" is converted into converted text data "innoinnovation," for example.
[0061] Alternatively, the conversion rules may be rules that assume actual mistakes in speech. For example, suppose that the word "patent permission (tokkyokyoka)" is frequently mistakenly pronounced as "tokkyokyokya." Based on such examples, a conversion rule may be set that "in words with many consonant 'k' after 'patent,' change the vowel or consonant." Such example-based conversion rules can also be trained using, for example, actual speech data.
[0062] The above-described conversion rules are merely examples, and the conversion rules stored in the conversion rule storage unit 121 are not limited to the above-described rules.
[0063] (Technical Effects) Next, the technical effects obtained by the information processing system 10 according to the third embodiment will be described.
[0064] As explained in Fig. 7, in the information processing system 10 according to the third embodiment, converted text data is generated based on the conversion rules. In this way, it becomes possible to generate converted text data more easily and appropriately. Furthermore, by updating the conversion rules as appropriate, it becomes possible to generate more appropriate converted text data compared to the case where the same conversion rules are continuously used.
[0065] <Fourth embodiment> An information processing system 10 according to the fourth embodiment will be described with reference to Fig. 8. The fourth embodiment differs only in part of the configuration and operation from the first to third embodiments described above, and other parts may be the same as the first to third embodiments. Therefore, the following will describe in detail the parts that differ from the embodiments already described, and will omit a description of other overlapping parts as appropriate.
[0066] (Functional configuration) First, the functional configuration of the information processing system 10 according to the fourth embodiment will be described with reference to Fig. 8. Fig. 8 is a block diagram showing the functional configuration of the information processing system according to the fourth embodiment. Note that in Fig. 8, the same elements as those shown in Fig. 2 are denoted by the same reference numerals.
[0067] 8, the information processing system 10 according to the fourth embodiment is configured to include, as components for realizing its functions, a first text data acquisition unit 110, a text data conversion unit 120, a converted voice data generation unit 130, a learning unit 140, a second text data acquisition unit 200, and a conversion learning unit 210. That is, the information processing system 10 according to the fourth embodiment further includes, in addition to the configuration of the first embodiment already described (see FIG. 2), a second text data acquisition unit 200 and a conversion learning unit 210. Each of the second text data acquisition unit 200 and the conversion learning unit 210 may be a processing block realized by, for example, the above-mentioned processor 11 (see FIG. 1).
[0068] The second text data acquisition unit 200 is configured to be able to acquire second text data for training the text data conversion unit 120. The second text data may include, for example, phrases that anticipate slip-ups. The second text data acquisition unit 200 may acquire a plurality of pieces of second text data. Note that the second text data acquisition unit 200 may acquire the second text data by voice input. That is, voice data may be converted into text data and acquired as the second text data.
[0069] The conversion learning unit 210 is configured to be able to train the text data conversion unit 120 using the second text data acquired by the second text data acquisition unit 200. The learning of the text data conversion unit 120 here is performed so that the text data conversion unit 120 can generate more appropriate converted text data from the first text data. The learning of the text data conversion unit 120 may be, for example, learning the conversion rules described in the third embodiment (see FIG. 7). Alternatively, the learning of the text data conversion unit 120 may be machine learning of a generative model that generates converted text data. Specific learning techniques used by the conversion learning unit 210 will be described in detail in other embodiments described later.
[0070] (Technical Effects) Next, the technical effects obtained by the information processing system 10 according to the fourth embodiment will be described.
[0071] As described in Fig. 8, in the information processing system 10 according to the fourth embodiment, the text data conversion unit 120 is trained using the second text data. In this way, it becomes possible to train the text data conversion unit 120 easily and appropriately. Furthermore, by training the text data conversion unit 120, it becomes possible to generate more appropriate converted text data from the first text data.
[0072] Fifth Embodiment An information processing system 10 according to the fifth embodiment will be described with reference to Figures 9 and 10. The fifth embodiment differs from the fourth embodiment described above only in some configurations and operations, and other parts may be the same as the first to fourth embodiments. Therefore, the following will describe in detail the parts that differ from the embodiments already described, and will omit a description of other overlapping parts as appropriate.
[0073] (Functional configuration) First, the functional configuration of the information processing system 10 according to the fifth embodiment will be described with reference to Fig. 9. Fig. 9 is a block diagram showing the functional configuration of the information processing system according to the fifth embodiment. Note that in Fig. 9, the same elements as those shown in Fig. 8 are denoted by the same reference numerals.
[0074] 9, the information processing system 10 according to the fifth embodiment is configured to include, as components for realizing its functions, a first text data acquisition unit 110, a text data conversion unit 120, a converted voice data generation unit 130, a learning unit 140, a second text data acquisition unit 200, and a conversion learning unit 210. In particular, the conversion learning unit 210 according to the fifth embodiment includes a similar word detection unit 211.
[0075] The similar word detection unit 211 is configured to detect whether similar words are included in the second text data. More specifically, the similar word detection unit 211 is configured to detect whether a first word and a second word that are similar to each other are included within a predetermined range of the second text data. The "predetermined range" here corresponds to the period of time it takes for a user who has made a mistake to correct the mistake (specifically, to rephrase it with the correct word), and may be set to an appropriate value in advance. The predetermined range may be set, for example, based on the number of characters in the text data. For example, the similar word detection unit 211 may determine whether similar words are included within a range of 20 characters. The predetermined range may be changeable by the user. For example, if too many similar words are detected, the predetermined range may be narrowed (for example, 20 characters may be changed to 15 characters). Conversely, if it is difficult to detect similar words, the predetermined range may be widened (for example, 20 characters may be changed to 30 characters). Here, similar words refer to, for example, words that differ from each other by only one or several letters, or words that have at least one consonant letter in common but different vowels.
[0076] The similar word detection unit 211 may calculate the similarity of each word included in the second text data and detect first and second words that are similar to each other. For example, the similar word detection unit 211 extracts words included in the second text data and calculates the similarity of each extracted word. It is possible to appropriately adopt existing technology as a similarity calculation method. If the similar word detection unit 211 determines that a pair of words exists whose similarity is higher than a predetermined threshold, it detects these words as a first word and a second word. The predetermined threshold is a threshold that is set in advance to determine whether words are similar or not. The predetermined threshold may be changeable by the user. For example, if too many similar words are detected, the predetermined threshold may be increased. Conversely, if it is difficult to detect similar words, the predetermined threshold may be decreased. It is also possible for the similar word detection unit 211 to detect similar words (i.e., a first word and a second word) using a method other than the above-described method.
[0077] (Conversion learning operation) Next, the flow of operations (hereinafter referred to as "conversion learning operations") performed when training the text data conversion unit 120 in the information processing system 10 according to the fifth embodiment will be described with reference to Fig. 10. Fig. 10 is a flowchart showing the flow of the conversion learning operations performed by the information processing system according to the fifth embodiment.
[0078] 10, when the conversion learning operation of the information processing system 10 according to the fifth embodiment is started, first, the second text data acquisition unit 200 acquires second text data (step S501). The second text data acquired by the second text data acquisition unit 200 is output to the conversion learning unit 210.
[0079] Next, the similar word detection unit 211 in the conversion learning unit 210 determines whether or not similar words exist within a predetermined range of the second text data (step S502). If similar words exist within the predetermined range (step S502: YES), the similar word detection unit 211 detects these words as first and second words (step S503).
[0080] For example, if the second text data includes the sentence "We are trying to innovate, to innovate...", the similar word detection unit 211 may detect "invention" and "innovation" as the first word and the second word, respectively. In this way, when a speaker makes a mistake in speech, there is a possibility that the speaker will correct the mistake immediately after realizing the mistake. The similar word detection unit 211 may detect the mistaken word and the corrected word as the first word and the second word, respectively.
[0081] Furthermore, the similar word detection unit 211 may detect multiple pairs of first words and second words from the second text data. For example, if the second text data includes a sentence such as "We are collecting various dates and data to innovate and innovate," the similar word detection unit 211 may detect "invention" and "innovation" as the first word and the second word, respectively, and may also detect "date" and "data" as the first word and the second word, respectively.
[0082] Furthermore, the similar word detection unit 211 may detect a third word similar to the first word and the second word in addition to them. For example, if the second text data includes the sentence "We are trying to innovate, to innovate, to innovate...", the similar word detection unit 211 may detect "invention", "innovation", and "innovation" as the first word, the second word, and the third word, respectively. In this way, if there are three or more similar words, all of them may be detected as similar words. In other words, the words detected by the similar word detection unit 211 are not limited to the first word and the second word.
[0083] If there are no similar words within the predetermined range (step S502: YES), the similar word detection unit 211 does not need to detect the first word and the second word (that is, the process of step S503 may be omitted).
[0084] Next, the conversion learning unit 210 performs training of the text data conversion unit 120 using the second text data (step S504). Particularly here, if the first and second words are detected in step S503 described above, the conversion learning unit 210 trains the text data conversion unit 120 by regarding one of the first and second words as a misspelling of the other. For example, if "invasion" and "innovation" are detected as the first and second words, the conversion learning unit 210 trains the text data conversion unit 120 by regarding "invasion" as a misspelling of "innovation." Furthermore, if three or more similar words are detected, training may be performed taking all of these words into consideration. For example, if the first, second, and third words are detected, the text data conversion unit 120 may train by regarding the first and second words as misspelled words and the third word as a corrected word. If the first word and the second word are not detected, the conversion learning unit 210 may perform learning of the text data conversion unit 120 without taking into consideration the presence of the first word and the second word.
[0085] (Technical Effects) Next, the technical effects obtained by the information processing system 10 according to the fifth embodiment will be described.
[0086] 9 and 10, in the information processing system 10 according to the fifth embodiment, first and second words that are similar to each other are detected to train the text data conversion unit 120. In this way, the misspoken word and the corrected word can be taken into consideration, and therefore the text data conversion unit 120 can be trained more appropriately.
[0087] Sixth Embodiment An information processing system 10 according to the sixth embodiment will be described with reference to Figures 11 to 13. The sixth embodiment differs only in part of the configuration and operation from the fourth and fifth embodiments described above, and other parts may be the same as the first to fifth embodiments. Therefore, the following will describe in detail the parts that differ from the embodiments already described, and will omit a description of other overlapping parts as appropriate.
[0088] (Functional configuration) First, the functional configuration of the information processing system 10 according to the sixth embodiment will be described with reference to Fig. 11. Fig. 11 is a block diagram showing the functional configuration of the information processing system according to the sixth embodiment. In Fig. 11, the same elements as those shown in Fig. 8 are denoted by the same reference numerals.
[0089] As shown in FIG. 11 , the information processing system 10 according to the sixth embodiment includes, as components for realizing its functions, a first text data acquisition unit 110, a text data conversion unit 120, a converted voice data generation unit 130, a learning unit 140, a second text data acquisition unit 200, a conversion learning unit 210, a second text data presentation unit 220, and a third text data acquisition unit 230. That is, the information processing system 10 according to the sixth embodiment further includes, in addition to the configuration of the fourth embodiment already described (see FIG. 8 ), a second text data presentation unit 220 and a third text data acquisition unit 230. Each of the second text data presentation unit 220 and the third text data acquisition unit 230 may be a processing block realized by, for example, the above-described processor 11 (see FIG. 1 ). Furthermore, the second text data presentation unit 220 may be realized by including the above-described output device 16 (see FIG. 1 ).
[0090] The second text data presentation unit 220 is configured to be able to present the second text data acquired by the second text data acquisition unit to the user. The method of presenting the second text data by the second text data presentation unit 220 is not particularly limited. For example, the second text data presentation unit 220 may display the second text data to the user via a display. Alternatively, the second text data presentation unit 220 may output the second text data as sound via a speaker (i.e., may convert the text data into sound data and output it). The specific presentation method by the second text data presentation unit 220 will be described in detail later.
[0091] The third text data acquisition unit 230 is configured to be able to acquire third text data in response to a user input presented by the second text data presentation unit 220. The third text data acquisition unit 230 may acquire the third text data, for example, via the above-mentioned input device 15 (see FIG. 1). The third text data is text data used for training the text data conversion unit 120, and is acquired as data corresponding to the second text data. For example, the third text data may be acquired as text data showing examples of misspellings of the second text data.
[0092] (Conversion learning operation) Next, the flow of the conversion learning operation in the information processing system 10 according to the sixth embodiment will be described with reference to Fig. 12. Fig. 12 is a flowchart showing the flow of the conversion learning operation by the information processing system according to the sixth embodiment.
[0093] 12, when the conversion learning operation of the information processing system 10 according to the sixth embodiment is started, first, the second text data acquisition unit 200 acquires second text data (step S601). The second text data acquired by the second text data acquisition unit 200 is output to the conversion learning unit 210 and the second text data presentation unit, respectively.
[0094] Next, the second text data presentation unit 220 presents the second text data acquired by the second text data acquisition unit 200 to the user (step S602). After that, the third text data acquisition unit 230 accepts an input from the user and acquires third text data (step S603). The third text data acquired by the third text data acquisition unit 230 is output to the conversion learning unit 210.
[0095] Next, the conversion learning unit 210 performs learning of the text data conversion unit 120 using the second text data acquired by the second text data acquisition unit 200 and the third text data acquired by the third text data acquisition unit 230 (step S604). Note that if the third text data has not been acquired (for example, if no input has been made by the user), the conversion learning unit 210 may perform learning of the text data conversion unit 120 using only the second text data.
[0096] (Example of second text data presentation) Next, a specific presentation example will be described with reference to Fig. 13 regarding a method for presenting second text data by the second text data presenting unit 220. Fig. 13 is a plan view showing an example of presenting second text data by the information processing system according to the sixth embodiment.
[0097] In the example shown in FIG. 13, second text data is presented using a display. Here, the second text data is displayed in a character string field. A conversion example field is displayed as a space for the user to input third text data. Specifically, the character string field displays the second text data "innovation." The conversion example field also displays a message prompting the user to input, saying "Enter a new character string here." This message may be configured to disappear when the user starts inputting.
[0098] When the above-described presentation is made, the user who receives the presentation inputs third text data corresponding to the second text data, "innovation." The user may input multiple pieces of third text data. For example, the user may input "i-bashion," "ino-innovation," "inoesho," and the like, which are examples of misspellings of "innovation," as the third text data.
[0099] Although an example in which only one piece of second text data is displayed has been given here, if a plurality of pieces of second text data have been acquired, the acquired plurality of pieces of second text data may be displayed in a list format, and third text data corresponding to each of the plurality of second text data may be input. Also, if one piece of second text data contains a plurality of words, a plurality of words contained in the second text data may be extracted, and each word may be displayed in a list format, and third text data corresponding to each word may be input.
[0100] (Technical Effects) Next, the technical effects obtained by the information processing system 10 according to the sixth embodiment will be described.
[0101] As described with reference to FIGS. 11 to 13, in the information processing system 10 according to the sixth embodiment, second text data is presented, and third text data is acquired in response to a user's input. When training the text data conversion unit 120, the third text data is used in addition to the second text data. This allows for more appropriate training than when training is performed using only the second text data. For example, by using the third text data, which is an example of a misspelling of the second text data, for training, the text data conversion unit 120 can generate appropriate converted text data.
[0102] Seventh Embodiment An information processing system 10 according to the seventh embodiment will be described with reference to Figures 14 and 15. The seventh embodiment differs only in part of the configuration and operation from the fourth to sixth embodiments described above, and other parts may be the same as the first to sixth embodiments. Therefore, the following will describe in detail the parts that differ from the embodiments already described, and will omit a description of other overlapping parts as appropriate.
[0103] (Functional configuration) First, the functional configuration of the information processing system 10 according to the seventh embodiment will be described with reference to Fig. 14. Fig. 14 is a block diagram showing the functional configuration of the information processing system according to the seventh embodiment. In Fig. 14, the same elements as those shown in Fig. 8 are denoted by the same reference numerals.
[0104] 14, the information processing system 10 according to the seventh embodiment is configured to include, as components for realizing its functions, a first text data acquisition unit 110, a text data conversion unit 120, a converted voice data generation unit 130, a learning unit 140, a second text data acquisition unit 200, a conversion learning unit 210, a minutes text data acquisition unit 240, and a tension level acquisition unit 250. That is, the information processing system 10 according to the seventh embodiment further includes, in addition to the configuration of the fourth embodiment already described (see FIG. 8), the minutes text data acquisition unit 240 and the tension level acquisition unit 250. Each of the minutes text data acquisition unit 240 and the tension level acquisition unit 250 may be a processing block realized by, for example, the above-mentioned processor 11 (see FIG. 1).
[0105] The minutes text data acquisition unit 240 is configured to be able to acquire multiple minutes text data. The minutes text data is data in which the contents of speeches made in a meeting have been converted into text. The minutes text data acquisition unit 240 may acquire minutes text data that has been converted into text outside the system, or may acquire the contents of speeches (audio data) and then convert it into text to acquire the minutes text data. The minutes text data may include information about the meeting and information about the participants in the meeting. The minutes text data may include information that identifies the speaker. For example, each sentence included in the minutes text data may be linked to information that identifies the speaker.
[0106] The tension level acquisition unit 250 is configured to acquire the tension level of a meeting that serves as the source of the minutes text data. The tension level acquisition unit 250 may acquire the tension level based on the minutes text data. Alternatively, the tension level acquisition unit 250 may acquire information about the meeting separately from the minutes text data and acquire the tension level from that information. The tension level may be acquired based on, for example, the participants in the meeting. For example, a high tension level may be acquired for a meeting attended by company executives or a meeting including participants from other companies. Furthermore, a low tension level may be acquired for a meeting attended only by employees from the same department or only by junior employees. Alternatively, the tension level may be acquired according to the size of the meeting. For example, a high tension level may be acquired for a meeting with 1,000 or more participants. Furthermore, a low tension level may be acquired for a meeting with only two or three participants. The tension level may be expressed in three levels, for example, "low," "medium," and "high," or may be expressed in more detailed values (for example, a value from "1 to 100").
[0107] (Conversion learning operation) Next, the flow of the conversion learning operation in the information processing system 10 according to the seventh embodiment will be described with reference to Fig. 15. Fig. 15 is a flowchart showing the flow of the conversion learning operation by the information processing system according to the seventh embodiment.
[0108] 15, when the conversion learning operation of the information processing system 10 according to the seventh embodiment is started, first the minutes text data acquisition unit 240 acquires a plurality of minutes text data (step S701). The plurality of minutes text data acquired by the minutes text data acquisition unit 240 is output to the tension level acquisition unit 250. The minutes text data acquisition unit 240 may be configured to output only information related to the meeting corresponding to the plurality of minutes text data (i.e., only information used to acquire the tension level) to the tension level acquisition unit 250.
[0109] Next, the tension level acquisition unit 250 acquires the tension level of the meeting (step S702). Information relating to the tension level acquired by the tension level acquisition unit 250 is output to the second text data.
[0110] Next, the second text data acquisition unit 200 acquires second text data based on the tension level acquired by the tension level acquisition unit 250 (step S703). Specifically, the second text data acquisition unit 200 acquires, as second text data, those minutes of meeting data acquired by the minutes text data acquisition unit 240 that have a tension level higher than a predetermined value. The "predetermined value" here is a threshold value for determining whether the tension level is high enough to determine that a slip of the tongue is likely to occur, and is set in advance. The predetermined value may be configured to be appropriately changeable by, for example, a user. For example, if the user wants to increase the amount of minutes of meeting text data acquired as second text data (i.e., increase the number of text data used for learning), the predetermined value may be changed to a lower value. On the other hand, if the user wants to decrease the amount of minutes of meeting text data acquired as second text data (i.e., decrease the number of text data used for learning), the predetermined value may be changed to a higher value. The second text data acquired by the second text data acquisition unit 200 is output to the conversion learning unit 210.
[0111] Next, the conversion learning unit 210 uses the second text data to train the text data conversion unit 120 (step S704). That is, the conversion learning unit 210 trains the text data conversion unit 120 using the minutes text data in which the degree of tension is higher than a predetermined value.
[0112] (Technical Effects) Next, the technical effects obtained by the information processing system 10 according to the seventh embodiment will be described.
[0113] 14 and 15, in the information processing system 10 according to the seventh embodiment, minutes text data in which the degree of tension of the meeting is higher than a predetermined value is acquired as the second text data. In this way, learning is performed using data that is likely to contain slip-ups, allowing the text data conversion unit 120 to learn more appropriately.
[0114] In the fourth to seventh embodiments, the configurations for executing learning of the text data conversion unit 120 using the second text data have been described, but the configurations of these embodiments may be combined. That is, the configurations of the fourth to seventh embodiments may be combined to perform learning of the text data conversion unit 120.
[0115] Eighth Embodiment An information processing system 10 according to the eighth embodiment will be described with reference to Fig. 16. The eighth embodiment differs only in part of the configuration and operation from the first to seventh embodiments described above, and other parts may be the same as the first to seventh embodiments. Therefore, the following will describe in detail the parts that differ from the embodiments already described, and will omit a description of other overlapping parts as appropriate.
[0116] (Functional configuration) First, the functional configuration of the information processing system 10 according to the eighth embodiment will be described with reference to Fig. 16. Fig. 16 is a block diagram showing the functional configuration of the information processing system according to the eighth embodiment. In Fig. 16, the same elements as those shown in Fig. 2 are denoted by the same reference numerals.
[0117] 16, the information processing system 10 according to the eighth embodiment is configured to include, as components for realizing its functions, a first text data acquisition unit 110, a text data conversion unit 120, a converted voice data generation unit 130, a learning unit 140, and a voice recognition unit 300. That is, the information processing system 10 according to the eighth embodiment further includes the voice recognition unit 300 in addition to the configuration of the first embodiment already described (see FIG. 2). The voice recognition unit 300 may be a processing block realized by, for example, the above-mentioned processor 11 (see FIG. 1).
[0118] The speech recognition unit 300 is configured to be able to convert input speech data into text data and output the text data. That is, the speech recognition unit 300 has the same function as the speech recognizer 50 described in the first to seventh embodiments. Furthermore, the speech recognition unit 300 is configured to be trained by the training unit 140, similar to the speech recognizer 50. That is, the speech recognition unit 300 is trained using first text data and converted speech data. Note that the speech recognizer 50 described in the first to seventh embodiments is not included in the components of the information processing system 10, whereas the speech recognition unit 300 is included in the components of the information processing system 10. Furthermore, the speech recognition unit 300 includes a slip-up correction unit 301.
[0119] The slip-up correction unit 301 is configured to be able to correct slip-ups included in the voice data. Therefore, when voice data including a slip-up is input to the voice recognition unit 300, text data in which the slip-up has been corrected is output. The slip-up correction unit 301 may correct the slip-up, for example, after the voice data has been converted into text. That is, the voice data may first be converted into text with the slip-up included, and then the slip-up is corrected. Alternatively, the slip-up correction unit 301 may correct the slip-up during the process of converting the voice data into text. That is, when voice data including a slip-up is input, text data in which the slip-up has been corrected may be generated.
[0120] When the input speech data includes multiple mistakes in speech, the mistaken speech corrector 301 may correct all or some of the mistakes in speech. The configuration for correcting some of the mistakes in speech will be described in detail in another embodiment below.
[0121] (Technical Effects) Next, the technical effects obtained by the information processing system 10 according to the eighth embodiment will be described.
[0122] 16, in the information processing system 10 according to the eighth embodiment, a process of correcting a slip of the tongue (or a process of generating text data in which the slip of the tongue has been corrected) is executed in the speech recognition unit 300. In this way, even when speech data containing a slip of the tongue is input, the slip of the tongue can be corrected and appropriate text data (text data that does not contain the slip of the tongue) can be output.
[0123] Ninth Embodiment An information processing system 10 according to the ninth embodiment will be described with reference to Figures 17 and 18. The ninth embodiment differs from the eighth embodiment described above only in some configurations and operations, and other parts may be the same as the first to eighth embodiments. Therefore, the following will describe in detail the parts that differ from the embodiments already described, and will omit a description of other overlapping parts as appropriate.
[0124] (Functional configuration) First, the functional configuration of the information processing system 10 according to the ninth embodiment will be described with reference to Fig. 17. Fig. 17 is a block diagram showing the functional configuration of the information processing system according to the ninth embodiment. In Fig. 17, the same elements as those shown in Fig. 16 are denoted by the same reference numerals.
[0125] 17, the information processing system 10 according to the ninth embodiment is configured to include, as components for realizing its functions, a first text data acquisition unit 110, a text data conversion unit 120, a converted voice data generation unit 130, a learning unit 140, and a voice recognition unit 300. In particular, the voice recognition unit 300 according to the ninth embodiment includes a score calculation unit 302 in addition to the misspoken speech correction unit 301 described in the eighth embodiment (see FIG. 16).
[0126] The score calculation unit 302 is configured to calculate a score indicating the possibility that the speech data contains a misspelling. This score may be calculated based on words included in the speech data. For example, if "innovation" is misspelled as "ivasion," "innovation" is a word found in a general dictionary, but "ivasion" is not. In this case, it may be determined that "ivasion" is likely to be a misspelling of "innovation," and a relatively high score may be calculated. On the other hand, if "data" is misspelled as "date," both "data" and "date" are words found in a general dictionary. In this case, it may be determined that "date" is unlikely to be a misspelling of "data," and a relatively low score may be calculated. Furthermore, if similar words frequently appear before and after a specific word in the speech data, or if there is a large difference between the frequency of the specific word and the frequency of the similar word, it may be determined that the specific word is likely to be a misspelling of a similar word. In this case, both the specific word and the similar word are words registered in the dictionary. For example, if "data" appears frequently before and after "date," or if "date" appears once and "data" appears 20 times, it is determined that "date" is likely a misspelling of "data."
[0127] The slip-up correction unit 301 according to this embodiment is configured to determine whether to correct a slip-up based on the score calculated by the score calculation unit 302. For example, the slip-up correction unit 301 may compare the calculated score with a predetermined reference score to determine whether to correct a slip-up. Specifically, the slip-up correction unit 301 may correct the slip-up if the calculated score is higher than the reference score, and may not correct the slip-up if the calculated score is lower than the reference score. Alternatively, the slip-up correction unit 301 may correct the slip-up if the score is high, insert a caution (a display warning that a slip-up may have occurred) if the score is medium, and not correct the slip-up if the score is low. The degree of correction may also be changed depending on the score. For example, the degree of correction may be increased if the score is high, so that a relatively large number of words are corrected, and the degree of correction may be decreased if the score is low, so that a relatively small number of words are corrected.
[0128] (Voice recognition operation) Next, the flow of the operation of converting voice data into text data in the information processing system 10 according to the ninth embodiment (hereinafter referred to as "voice recognition operation") will be described with reference to Fig. 18. Fig. 18 is a flowchart showing the flow of the voice recognition operation by the information processing system according to the ninth embodiment.
[0129] 18, when the speech recognition operation of the information processing system 10 according to the ninth embodiment is started, the speech recognition unit 300 first acquires speech data (step S901). Then, the score calculation unit 302 calculates a score indicating the possibility that the speech data contains a slip of the tongue (step S902).
[0130] Next, the slip-up correction unit 301 determines whether the score calculated by the score calculation unit 302 is higher than the standard score (step S903). If the calculated score is higher than the standard score (step S903: YES), the slip-up correction unit 301 corrects the slip-up. As a result, text data in which the slip-up has been corrected is output (step S904). On the other hand, if the calculated score is lower than the standard score (step S903: NO), the slip-up correction unit 301 does not correct the slip-up. As a result, text data in which the slip-up has not been corrected is output (step S905).
[0131] While the example given here is one in which it is determined whether or not to correct a slip of the tongue based on a standard score, as already explained, it is also possible to insert a caution, change the degree of correction, etc. Furthermore, whether or not to correct a slip of the tongue may be determined on a word-by-word basis, a sentence-by-sentence basis, or a data-by-data basis.
[0132] (Technical Effects) Next, the technical effects obtained by the information processing system 10 according to the ninth embodiment will be described.
[0133] 17 and 18, in the information processing system 10 according to the ninth embodiment, it is determined whether or not to correct a slip of the tongue included in the speech data based on the calculated score. In this way, it is possible to appropriately correct a slip of the tongue while preventing a non-sleep error from being erroneously corrected.
[0134] Tenth Embodiment An information processing system 10 according to a tenth embodiment will be described with reference to Figures 19 and 20. The tenth embodiment differs only in part of the configuration and operation from the eighth and ninth embodiments described above, and other parts may be the same as the first to eighth embodiments. Therefore, the following will describe in detail the parts that differ from the embodiments already described, and will omit a description of other overlapping parts as appropriate.
[0135] (Functional configuration) First, the functional configuration of the information processing system 10 according to the tenth embodiment will be described with reference to Fig. 19. Fig. 19 is a block diagram showing the functional configuration of the information processing system according to the tenth embodiment. In Fig. 19, the same elements as those shown in Fig. 16 are denoted by the same reference numerals.
[0136] 19, the information processing system 10 according to the tenth embodiment includes, as components for realizing its functions, a first text data acquisition unit 110, a text data conversion unit 120, a converted voice data generation unit 130, a learning unit 140, and a voice recognition unit 300. In particular, the voice recognition unit 300 according to the tenth embodiment includes a tension level determination unit 303 in addition to the slip of the tongue correction unit 301 described in the eighth embodiment (see FIG. 16). Note that the voice recognition unit 300 according to the tenth embodiment receives input of recorded voice data of the proceedings including the content of speech in the conference.
[0137] The tension level determination unit 303 is configured to be able to determine the tension level of a conference from which the minutes recording voice data has been recorded. The tension level determination unit 303 may determine the tension level, for example, in the same manner as the above-mentioned tension level acquisition unit 250 (see FIG. 14). The tension level determination unit 303 may acquire the tension level based on pseudo-theoretical voice data. Alternatively, the tension level determination unit 303 may acquire information about the conference separately from the minutes recording voice data and acquire the tension level from that information. The tension level may be acquired according to, for example, the participants in the conference, the scale of the conference, etc.
[0138] The slip-up correction unit 301 according to this embodiment is configured to determine whether to correct a slip-up based on the level of tension determined by the tension level determination unit 303. For example, the slip-up correction unit 301 may compare the determined level of tension with a predetermined reference value to determine whether to correct a slip-up. Specifically, the slip-up correction unit 301 may correct the slip-up if the determined level of tension is higher than the reference value, and may not correct the slip-up if the determined level of tension is lower than the reference value. Alternatively, the slip-up correction unit 301 may correct the slip-up if the level of tension is high, insert a caution (a display warning that a slip-up may occur) if the level of tension is medium, and not correct the slip-up if the level of tension is low. The degree of correction may also be changed depending on the level of tension. For example, the degree of correction may be increased when the level of tension is high, thereby correcting a relatively large number of words, and the degree of correction may be decreased when the level of tension is low, thereby correcting a relatively small number of words.
[0139] (Voice recognition operation) Next, the flow of the operation of converting voice data into text data in the information processing system 10 according to the tenth embodiment (hereinafter referred to as "voice recognition operation") will be described with reference to Fig. 20. Fig. 20 is a flowchart showing the flow of the voice recognition operation by the information processing system according to the tenth embodiment.
[0140] 20, when the speech recognition operation of the information processing system 10 according to the tenth embodiment is started, the speech recognition unit 300 first acquires speech data (minutes recording voice data) (step S1001). Then, the tension determination unit 303 determines the tension level of the conference from which the minutes recording voice data was recorded (step S1002).
[0141] Next, the slip-up correction unit 301 determines whether the level of tension determined by the tension level determination unit 303 is higher than a reference value (step S1003). If the determined level of tension is higher than the reference value (step S1003: YES), the slip-up correction unit 301 corrects the slip-up. As a result, text data in which the slip-up has been corrected is output (step S1004). On the other hand, if the determined level of tension is lower than the reference value (step S1003: NO), the slip-up correction unit 301 does not correct the slip-up. As a result, text data in which the slip-up has not been corrected is output (step S1005).
[0142] While the example given here is one in which it is determined whether or not to correct a slip of the tongue based on a reference value, as already explained, it is also possible to insert a caution, change the degree of correction, etc. Furthermore, whether or not to correct may be determined on a word-by-word basis, a sentence-by-sentence basis, or a data-by-data basis.
[0143] (Technical Effects) Next, the technical effects obtained by the information processing system 10 according to the tenth embodiment will be described.
[0144] 19 and 20, the information processing system 10 according to the ninth embodiment determines whether to correct a slip of the tongue included in the audio data based on the level of tension in the meeting. In this way, it is possible to appropriately correct a slip of the tongue while preventing a non-sleepless speech from being erroneously corrected.
[0145] In the eighth to tenth embodiments, the information processing system 10 is described as having the speech recognition unit 300, but the configurations of these embodiments may be combined. That is, the configurations of the eighth to tenth embodiments may be combined to realize a speech recognition unit 300 that performs speech recognition operations.
[0146] The scope of each embodiment also includes a processing method in which a program that operates the configuration of each embodiment to realize the functions of the above-described embodiments is recorded on a recording medium, the program recorded on the recording medium is read as code, and the program is executed on a computer. In other words, a computer-readable recording medium is also included in the scope of each embodiment. Furthermore, each embodiment includes not only a recording medium on which the above-described program is recorded, but also the program itself.
[0147] Examples of recording media that can be used include floppy disks, hard disks, optical disks, magneto-optical disks, CD-ROMs, magnetic tapes, non-volatile memory cards, and ROMs. Furthermore, the scope of each embodiment is not limited to programs that execute processing by themselves, but also includes programs that execute processing by operating on an OS in cooperation with other software or functions of an expansion board. Furthermore, the program itself may be stored on a server, and part or all of the program may be downloadable from the server to a user terminal.
[0148] <Additional Notes> The above-described embodiment may be further described as follows, but is not limited to the following.
[0149] (Appendix 1) The information processing system described in Appendix 1 is an information processing system comprising: first text data acquisition means for acquiring first text data; text data conversion means for converting the first text data to generate converted text data; converted voice data generation means for generating converted voice data corresponding to the converted text data; and learning means for training speech recognition means that receives the first text data and the converted voice data as input and generates text data corresponding to the voice data from the voice data. is.
[0150] (Appendix 2) The information processing system described in Appendix 2 is the information processing system described in Appendix 1, further comprising a first voice data generation means for generating first voice data corresponding to the first text data, and the learning means trains the voice recognition means using the first text data, the converted voice data, and the first voice data as inputs.
[0151] (Appendix 3) The information processing system described in Supplementary Note 3 is the information processing system described in Supplementary Note 1 or 2, wherein the text data conversion means stores at least one conversion rule and generates the converted text data based on the conversion rule.
[0152] (Appendix 4) The information processing system described in Appendix 4 is the information processing system described in any one of Appendixes 1 to 3, further comprising a second text data acquisition means for acquiring second text data, and a conversion learning means for learning the text data conversion means using the second text data.
[0153] (Appendix 5) The information processing system described in Appendix 5 is the information processing system described in Appendix 4, wherein, when a first word and a second word that are similar to each other are contained within a predetermined range in the second text data, the conversion learning means determines that one of the first word and the second word is a misspelling of the other, and trains the text data conversion means.
[0154] (Appendix 6) The information processing system described in Appendix 6 is the information processing system described in Appendix 4 or 5, further comprising a presentation means for presenting the second text data to a user, and a third text data acquisition means for acquiring third text data corresponding to the second text data in response to the user's operation after receiving the presentation by the presentation means, wherein the conversion learning means trains the text data conversion means using the second text data and the third text data.
[0155] (Appendix 7) The information processing system described in Appendix 7 is an information processing system described in any one of Appendixes 4 to 6, further comprising a minutes text data acquisition means for acquiring multiple minutes text data that are text versions of speeches made during a meeting, and a tension level acquisition means for acquiring the tension level of the meeting, wherein the second text data acquisition means acquires, from the multiple minutes text data, one whose tension level is higher than a predetermined value as the second text data.
[0156] (Appendix 8) The information processing system described in Appendix 8 is the information processing system described in any one of Appendixes 1 to 7, further comprising the speech recognition means, wherein the speech recognition means outputs the text data in which a mistake in speech in the speech data has been corrected based on a learning result by the learning means.
[0157] (Appendix 9) The information processing system described in Appendix 9 is the information processing system described in Appendix 8, wherein the speech recognition means calculates a score indicating the possibility that the speech data contains a slip-up, and determines whether to correct the slip-up in the speech data based on the score.
[0158] (Appendix 10) The information processing system described in Appendix 10 is the information processing system described in Appendix 8 or 9, wherein the voice data is recorded conference proceedings voice data including the content of speech in the conference, and the voice recognition means determines the level of tension in the conference and decides whether or not to correct a slip of the tongue in the voice data based on the level of tension.
[0159] (Appendix 11) The information processing device described in Appendix 11 is an information processing device comprising: first text data acquisition means for acquiring first text data; text data conversion means for converting the first text data to generate converted text data; converted voice data generation means for generating converted voice data corresponding to the converted text data; and learning means for training voice recognition means that receives the first text data and the converted voice data as input and generates text data corresponding to the voice data from the voice data.
[0160] (Appendix 12) The information processing method described in Appendix 12 is an information processing method executed by at least one computer, which acquires first text data, converts the first text data to generate converted text data, generates converted voice data corresponding to the converted text data, and trains a voice recognition means that uses the first text data and the converted voice data as inputs and generates text data corresponding to the voice data from the voice data.
[0161] (Appendix 13) The recording medium described in Appendix 13 is a recording medium having recorded thereon a computer program for causing at least one computer to execute an information processing method of acquiring first text data, converting the first text data to generate converted text data, generating converted voice data corresponding to the converted text data, and training a voice recognition means that uses the first text data and the converted voice data as inputs and generates text data corresponding to the voice data from the voice data.
[0162] (Appendix 14) The computer program described in Appendix 14 is a computer program that causes at least one computer to execute an information processing method that acquires first text data, converts the first text data to generate converted text data, generates converted voice data corresponding to the converted text data, and trains a voice recognition means that uses the first text data and the converted voice data as inputs and generates text data corresponding to the voice data from the voice data.
[0163] This disclosure may be modified as appropriate within the scope that does not contradict the gist or idea of the invention that can be read from the claims and the entire specification, and information processing systems, information processing devices, information processing methods, and recording media that involve such modifications are also included in the technical idea of this disclosure. [Explanation of symbols]
[0164] 10 Information Processing Systems 11 processors 14 Storage device 50 Speech Recognizer 110 First text data acquisition unit 120 Text data conversion unit 121 Conversion rule memory 130 Converted voice data generation unit 140 Learning Department 150 First audio data generation unit 200 Second text data acquisition unit 210 Conversion Learning Unit 211 Similar Word Detection Unit 220 Second text data presentation unit 230 Third Text Data Acquisition Unit 240 Minutes Text Data Acquisition Unit 250 Tension level acquisition part 300 Voice Recognition Unit 301 Mispronunciation Correction Department 302 Score Calculation Unit 303 Tension level determination section
Claims
1. a first text data acquisition means for acquiring first text data; a text data conversion means for converting at least a part of the first text data into characters with different pronunciations to generate converted text data; a converted voice data generating means for generating converted voice data corresponding to the converted text data; a learning means for learning a speech recognition means that receives the first text data and the converted speech data as input and generates text data corresponding to the speech data from the speech data; An information processing system comprising:
2. a first voice data generating means for generating first voice data corresponding to the first text data; the learning means uses the first text data, the converted voice data, and the first voice data as inputs to learn the voice recognition means; The information processing system according to claim 1 .
3. the text data conversion means stores at least one conversion rule and generates the converted text data based on the conversion rule; 3. The information processing system according to claim 1 or 2.
4. A second text data acquisition means for acquiring second text data different from the first text data; conversion learning means for learning the text data conversion means using the second text data; The information processing system according to claim 1 , further comprising:
5. the conversion learning means, when a first word and a second word that are similar to each other are included within a predetermined range in the second text data, determines that one of the first word and the second word is a misspelling of the other, and performs learning of the text data conversion means. The information processing system according to claim 4 .
6. a presentation means for presenting the second text data to a user; a third text data acquisition means for acquiring third text data corresponding to the second text data in response to an operation of the user who has received the presentation by the presentation means; Further provided with the conversion learning means uses the second text data and the third text data to perform learning of the text data conversion means; 6. The information processing system according to claim 4 or 5.
7. a minutes text data acquisition means for acquiring a plurality of minutes text data obtained by converting the contents of speeches made in a meeting into text; a tension level acquisition means for acquiring a tension level of the meeting; Further provided with the second text data acquisition means acquires, from the plurality of minutes text data, the minutes text data whose degree of tension is higher than a predetermined value as the second text data; The information processing system according to any one of claims 4 to 6.
8. The voice recognition means is further provided, the speech recognition means outputs the text data in which mistakes in speech in the speech data have been corrected based on the learning result of the learning means. The information processing system according to any one of claims 1 to 7.
9. 1. An information processing method executed by at least one computer, comprising: Obtaining first text data; converting at least a portion of the first text data into characters with different pronunciations to generate converted text data; generating converted voice data corresponding to the converted text data; training a speech recognition means that receives the first text data and the converted speech data as inputs and generates text data corresponding to the speech data from the speech data; Information processing methods.
10. At least one computer Obtaining first text data; converting at least a portion of the first text data into characters with different pronunciations to generate converted text data; generating converted voice data corresponding to the converted text data; training a speech recognition means that receives the first text data and the converted speech data as inputs and generates text data corresponding to the speech data from the speech data; A computer program that executes an information processing method.
Citation Information
Patent Citations
Voice recognition device, error correction model learning method and program
JP2014074732A
Intention estimation device and model learning method
JP2015230384A
Dialogue method, dialogue system, dialogue device, and program
JP2017208003A
Natural language processing method and device, and method and device of learning natural language processing model
JP2018081298A
Acoustic model training using corrected terms
JP2019528470A