Speech recognition system, speech recognition method, and recording medium

The speech recognition system addresses the challenge of converting real utterance data into synthesized voice for effective voice recognition by using a system that acquires, converts, synthesizes, and generates a conversion model for the data, achieving high accuracy and cost-effectiveness.

JP7691027B2Active Publication Date: 2025-06-11NEC CORP
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2024504041
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-03-01
Publication Date
2025-06-11
Estimated Expiration
2042-03-01

AI Technical Summary

Technical Problem

Existing speech recognition systems face challenges in efficiently converting real utterance data into synthesized voice for effective voice recognition, while also being cost-effective.

Method used

The proposed system includes an utterance data acquisition means, text conversion means, voice synthesis means, conversion model generation means, and voice recognition means to convert real utterance data into text, then synthesize it into voice, and generate a conversion model for voice recognition.

Benefits of technology

This approach allows for high recognition accuracy at a lower cost by generating a conversion model using real utterance data and synthesized voice, eliminating the need for separate preparation of both types of data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007691027000001
    Figure 0007691027000001
  • Figure 0007691027000002
    Figure 0007691027000002
  • Figure 0007691027000003
    Figure 0007691027000003
Patent Text Reader

Abstract

A speech recognition system (10) comprises: an utterance data acquisition means (110) that acquires real utterance data obtained through utterance by a speaker; a text conversion means (120) that converts the real utterance data into text data; a speech synthesis means (130) that generates a corresponding synthesized speech corresponding to the real utterance data by performing speech synthesis using the text data; a conversion model generation means (140) that generates a conversion model for converting, into synthesized speech, input speech by using the real utterance data and the corresponding synthesized speech; and a speech recognition means (220) that performs speech recognition of the synthesized speech converted by using the conversion model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the technical fields of speech recognition systems, speech recognition methods, and recording media.

Background Art

[0002] As this type of system, those that generate synthetic speech are known. For example, in Patent Document 1, it is disclosed that synthetic speech is generated by converting feature amounts representing the timbre of speech using a pre-trained conversion model. In Patent Document 2, it is disclosed that a sentence in the target language is generated from text data obtained as a speech recognition result, and synthetic speech is generated from the sentence in the target language.

[0003] As other related technologies, for example, in Patent Document 3, it is disclosed that a speech conversion model is trained using a training corpus.

Prior Art Documents

Patent Documents

[0004]

Patent Document 1

Patent Document 2

Patent Document 3

Summary of the Invention

Problems to be Solved by the Invention

[0005] This disclosure aims to improve the technologies disclosed in the prior art documents.

Means for Solving the Problems

[0006] One aspect of the voice recognition system of this disclosure includes an utterance data acquisition means for acquiring real utterance data uttered by a speaker, a text conversion means for converting the real utterance data into text data, a voice synthesis means for generating a corresponding synthesized voice corresponding to the real utterance data by voice synthesis using the text data, a conversion model generation means for generating a conversion model for converting an input voice into a synthesized voice using the real utterance data and the corresponding synthesized voice, and a voice recognition means for performing voice recognition on the synthesized voice converted using the conversion model.

[0007] One aspect of the voice recognition system of this disclosure includes a sign language data acquisition means for acquiring sign language data, a text conversion means for converting the sign language data into text data, a voice synthesis means for generating a corresponding synthesized voice corresponding to the sign language data by voice synthesis using the text data, a conversion model generation means for generating a conversion model for converting an input sign language into a synthesized voice using the sign language data and the corresponding synthesized voice, and a voice recognition means for performing voice recognition on the synthesized voice converted using the conversion model.

[0008] One aspect of the voice recognition method of this disclosure is that at least one computer acquires real utterance data uttered by a speaker, converts the real utterance data into text data, generates a corresponding synthesized voice corresponding to the real utterance data by voice synthesis using the text data, generates a conversion model for converting an input voice into a synthesized voice using the real utterance data and the corresponding synthesized voice, and performs voice recognition on the synthesized voice converted using the conversion model.

[0009] One aspect of the recording medium of this disclosure is a computer program recorded on at least one computer, which causes the computer to obtain real speech data spoken by a speaker, convert the real speech data into text data, generate a corresponding synthesized speech corresponding to the real speech data by voice synthesis using the text data, generate a conversion model for converting an input speech into a synthesized speech using the real speech data and the corresponding synthesized speech, and perform speech recognition on the synthesized speech converted using the conversion model.

Brief Description of the Drawings

[0010]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Figure 18

Figure 19

Figure 20

Figure 21

Embodiments for Carrying Out the Invention

[0011] Hereinafter, embodiments of the speech recognition system, the speech recognition method, and the recording medium will be described with reference to the drawings.

[0012] <First Embodiment> The speech recognition system according to the first embodiment will be described with reference to FIGS. 1 to 4.

[0013] (Hardware Configuration) First, with reference to FIG. 1, the hardware configuration of the speech recognition system according to the first embodiment will be described. FIG. 1 is a block diagram showing the hardware configuration of the speech recognition system according to the first embodiment.

[0014] As shown in FIG. 1, the speech recognition system 10 according to the first embodiment includes a processor 11, a RAM (Random Access Memory) 12, a ROM (Read Only Memory) 13, and a storage device 14. The speech recognition system 10 may further include an input device 15 and an output device 16. The above-described processor 11, RAM 12, ROM 13, storage device 14, input device 15, and output device 16 are connected via a data bus 17.

[0015] The processor 11 reads a computer program. For example, the processor 11 is configured to read a computer program stored in at least one of the RAM 12, ROM 13, and storage device 14. Alternatively, the processor 11 may read a computer program stored in a computer-readable recording medium using a recording medium reader (not shown). The processor 11 may obtain (i.e., read) a computer program from a device (not shown) disposed outside the speech recognition system 10 via a network interface. By executing the read computer program, the processor 11 controls the RAM 12, storage device 14, input device 15, and output device 16. In particular, in this embodiment, when the processor 11 executes the read computer program, function blocks for performing speech recognition are realized within the processor 11. That is, the processor 11 may function as a controller that executes each control in the speech recognition system 10.

[0016] Processor 11 may be configured as, for example, a CPU (Central Processing Unit), GPU (Graphics Processing Unit), FPGA (field-programmable gate array), DSP (Demand-Side Platform), or ASIC (Application Specific Integrated Circuit). Processor 11 may be configured with one of these, or may be configured to use a plurality in parallel.

[0017] RAM 12 temporarily stores the computer programs executed by processor 11. RAM 12 temporarily stores the data temporarily used by processor 11 when processor 11 is executing a computer program. RAM 12 may be, for example, D-RAM (Dynamic Random Access Memory) or SRAM (Static Random Access Memory). Also, instead of RAM 12, other types of volatile memory may be used.

[0018] ROM 13 stores the computer programs executed by processor 11. ROM 13 may also store other fixed data. ROM 13 may be, for example, P-ROM (Programmable Read Only Memory) or EPROM (Erasable Read Only Memory). Also, instead of ROM 13, other types of non-volatile memory may be used.

[0019] Storage device 14 stores the data that voice recognition system 10 stores long-term. Storage device 14 may operate as a temporary storage device for processor 11. Storage device 14 may include, for example, at least one of a hard disk device, a magneto-optical disk device, an SSD (Solid State Drive), and a disk array device.

[0020] The input device 15 is a device that receives input instructions from the user of the voice recognition system 10. The input device 15 may include, for example, at least one of a keyboard, a mouse, and a touch panel. The input device 15 may be configured as a portable terminal such as a smartphone or a tablet. The input device 15 may be a device capable of voice input including, for example, a microphone.

[0021] The output device 16 is a device that outputs information regarding the voice recognition system 10 to the outside. For example, the output device 16 may be a display device (e.g., a display) capable of displaying information regarding the voice recognition system 10. Further, the output device 16 may be a speaker or the like capable of outputting information regarding the voice recognition system 10 as voice. The output device 16 may be configured as a portable terminal such as a smartphone or a tablet. Also, the output device 16 may be a device that outputs information in a format other than an image. For example, the output device 16 may be a speaker that outputs information regarding the voice recognition system 10 as voice.

[0022] Note that, in FIG. 1, an example of the voice recognition system 10 configured to include a plurality of devices has been given, but all or some of these functions may be realized as one device (voice recognition device). In that case, the voice recognition device may be configured to include, for example, only the above-described processor 11, RAM 12, and ROM 13, and for the other components (i.e., the storage device 14, the input device 15, and the output device 16), an external device connected to the voice recognition device may be provided with them. Also, the voice recognition device may be one in which some arithmetic functions are realized by an external device (e.g., an external server or the cloud).

[0023] (Functional configuration) Next, with reference to FIG. 2, the functional configuration of the voice recognition system 10 according to the first embodiment will be described. FIG. 2 is a block diagram showing the functional configuration of the voice recognition system according to the first embodiment.

[0024] As shown in FIG. 2, the speech recognition system 10 according to the first embodiment includes, as components for realizing its functions, a speech data acquisition unit 110, a text conversion unit 120, a speech synthesis unit 130, a conversion model generation unit 140, a voice conversion unit 210, and a speech recognition unit 220. Each of the speech data acquisition unit 110, the text conversion unit 120, the speech synthesis unit 130, the conversion model generation unit 140, the voice conversion unit 210, and the speech recognition unit 220 may be a processing block realized by, for example, the above-described processor 11 (see FIG. 1).

[0025] The speech data acquisition unit 110 is configured to be able to acquire real speech data uttered by a speaker. The real speech data may be audio data (e.g., waveform data). The real speech data may be acquired, for example, from a database (real speech audio corpus) that accumulates a plurality of real speech data. The real speech data acquired by the speech data acquisition unit 110 is configured to be output to the text conversion unit 120 and the conversion model generation unit 140.

[0026] The text conversion unit 120 is configured to be able to convert the real speech data acquired by the speech data acquisition unit 110 into text data. That is, the text conversion unit 120 is configured to be able to execute a process of converting audio data into text. Note that existing techniques may be appropriately adopted for the specific method of text conversion. The text data converted by the text conversion unit 120 (i.e., the text data corresponding to the real speech data) is configured to be output to the speech synthesis unit 130.

[0027] The voice synthesis unit 130 is configured to be able to generate a corresponding synthesized voice corresponding to the real speech data by synthesizing the text data changed by the text conversion unit 120. Regarding the specific method of voice synthesis, existing technologies can be appropriately adopted. The corresponding synthesized voice generated by the voice synthesis unit 130 is configured to be output to the conversion model generation unit 140. Note that the corresponding synthesized voice may be output to the conversion model generation unit 140 after being stored in a database (synthesized voice corpus) capable of storing a plurality of corresponding syntheses.

[0028] The conversion model generation unit 140 is configured to be able to generate a conversion model that converts an input voice into a synthesized voice using the real speech data acquired by the speech data acquisition unit 110 and the corresponding synthesized voice synthesized by the voice synthesis unit 130. The conversion model converts, for example, an input voice (i.e., a human voice) spoken by a speaker so as to approach a synthesized voice (i.e., a mechanical voice). The conversion model generation unit 140 may be configured to generate a conversion model using, for example, GAN (Generative Adversarial Network). The conversion model generated by the conversion model generation unit 140 is configured to be output to the voice conversion unit 210.

[0029] The voice conversion unit 210 is configured to be able to convert an input voice into a synthesized voice using the conversion model generated by the conversion model generation unit 140. The input voice input to the voice conversion unit 210 may be, for example, a voice input using a microphone or the like. The synthesized voice converted by the voice conversion unit 210 is configured to be output to the voice recognition unit 220.

[0030] The voice recognition unit 220 is configured to be able to recognize the synthesized voice converted by the voice conversion unit 210. That is, the voice recognition unit 220 is configured to be able to execute a process of converting the synthesized voice into text. The voice recognition unit 220 may be configured to output the voice recognition result of the synthesized voice. Note that the usage method of the voice recognition result is not particularly limited.

[0031] (Conversion model generation operation) Next, with reference to FIG. 3, the flow of the operation (hereinafter, appropriately referred to as the "conversion model generation operation") when generating a conversion model by the speech recognition system 10 according to the first embodiment will be described. FIG. 3 is a flowchart showing the flow of the conversion model generation operation by the speech recognition system according to the first embodiment.

[0032] As shown in FIG. 3, when the conversion model generation operation by the speech recognition system 10 according to the first embodiment is started, first, the utterance data acquisition unit 110 acquires real utterance data (step S101). Then, the text conversion unit 120 converts the real utterance data acquired by the utterance data acquisition unit 110 into text data (step S102).

[0033] Subsequently, the speech synthesis unit 130 synthesizes the text data converted by the text conversion unit 120 to generate corresponding synthesized speech corresponding to the real utterance data (step S103). Then, the conversion model generation unit 140 generates a conversion model based on the real utterance data acquired by the utterance data acquisition unit 110 and the corresponding synthesized speech generated by the speech synthesis unit 130 (step S104). After that, the conversion model generation unit 140 outputs the generated conversion model to the speech conversion unit 210 (step S105).

[0034] (Conversion recognition operation) Next, with reference to FIG. 4, the flow of the operation (hereinafter, appropriately referred to as the "speech recognition operation") when performing speech recognition by the speech recognition system 10 according to the first embodiment will be described. FIG. 3 is a flowchart showing the flow of the speech recognition operation by the speech recognition system according to the first embodiment.

[0035] As shown in FIG. 4, when the voice recognition operation by the voice recognition system 10 according to the first embodiment is started, first, the voice conversion unit 210 acquires the input voice (step S151). Then, the voice conversion unit 210 reads the conversion model generated by the conversion model generation unit 140 (step S152). After that, the voice conversion unit 210 performs voice conversion using the read conversion model, and converts the input voice into a synthesized voice (step S153).

[0036] Subsequently, the voice recognition unit 220 reads a voice recognition model (that is, a model for performing voice recognition) (step S154). Then, the voice recognition unit 220 performs voice recognition on the synthesized voice synthesized by the voice conversion unit 210 using the read voice recognition model (step S155). After that, the voice recognition unit 220 outputs the voice recognition result (step S156).

[0037] (Technical effect) Next, the technical effect obtained by the voice recognition system 10 according to the first embodiment will be described.

[0038] As described with reference to FIGS. 1 to 4, in the voice recognition system 10 according to the first embodiment, when generating the conversion model, real utterance data and corresponding synthesized voices corresponding to the real utterance data are used. In particular, the corresponding synthesized voice corresponding to the real utterance data is generated by converting the real utterance data into text and synthesizing the text data into voice. In this way, it is not necessary to prepare both the real utterance data and the corresponding synthesized voice (that is, if only the real utterance data is prepared, the corresponding synthesized voice can be generated), so the cost required to generate the conversion model can be suppressed. As a result, it is possible to realize voice recognition with high recognition accuracy at low cost.

[0039] <Second Embodiment> The voice recognition system 10 according to the second embodiment will be described with reference to FIGS. 5 and 6. Note that the second embodiment is only different from the above-described first embodiment in some configurations and operations, and other parts may be the same as those of the first embodiment. Therefore, hereinafter, the parts different from the first embodiment already described will be described in detail, and the description of other overlapping parts will be omitted as appropriate.

[0040] (Functional configuration) First, with reference to FIG. 5, the functional configuration of the voice recognition system 10 according to the second embodiment will be described. FIG. 5 is a block diagram showing the functional configuration of the voice recognition system according to the second embodiment. In FIG. 5, the same reference numerals are given to the same elements as those shown in FIG. 2.

[0041] As shown in FIG. 5, the voice recognition system 10 according to the second embodiment includes, as components for realizing its functions, a speech data acquisition unit 110, a text conversion unit 120, a speech synthesis unit 130, a conversion model generation unit 140, a voice conversion unit 210, and a voice recognition unit 220. And in the second embodiment in particular, the conversion model generation unit 140 is configured such that the input speech input to the voice conversion unit 210 and the recognition result by the voice recognition unit 220 are input thereto. The conversion model generation unit 140 according to the second embodiment is configured to be able to execute learning of the conversion model based on the input speech input to the voice conversion unit 210 and the recognition result by the voice recognition unit 220.

[0042] (Conversion model learning operation) Next, with reference to FIG. 6, the flow of the operation (hereinafter, appropriately referred to as "conversion model learning operation") when learning the conversion model by the voice recognition system 10 according to the second embodiment will be described. FIG. 6 is a flowchart showing the flow of the conversion model generation operation by the voice recognition system according to the second embodiment.

[0043] As shown in FIG. 6, when the conversion model learning operation by the speech recognition system 10 according to the second embodiment is started, first, the conversion model generation unit 140 acquires the input speech input to the speech conversion unit 210 (step S201). Then, the conversion model generation unit 140 further acquires the speech recognition result when the input speech is input (that is, the speech recognition result output in step S156 shown in FIG. 4) (step S202).

[0044] Subsequently, the conversion model generation unit 140 learns the conversion model based on the acquired input speech and speech recognition result (step S203). At this time, the conversion model generation unit 140 may adjust the parameters of the already generated conversion model. After that, the conversion model generation unit 140 outputs the learned conversion model to the speech conversion unit 210 (step S204).

[0045] (Technical effect) Next, the technical effect obtained by the speech recognition system 10 according to the second embodiment will be described.

[0046] As described with reference to FIGS. 5 and 6, in the speech recognition system 10 according to the second embodiment, the conversion model is learned based on the input speech and the speech recognition result. In this way, learning is performed in consideration of how the input speech is actually speech recognized, so that the conversion model can be learned so that more appropriate speech conversion can be performed. Specifically, the conversion model can be learned so that the accuracy of speech recognition performed using the synthesized speech obtained by speech conversion is improved.

[0047] <Third Embodiment> The speech recognition system 10 according to the third embodiment will be described with reference to FIGS. 7 and 8. Note that the third embodiment is only different from the above-described first and second embodiments in some configurations and operations, and the other parts may be the same as those of the first and second embodiments. Therefore, hereinafter, the parts different from the already described embodiments will be described in detail, and the description of the other overlapping parts will be omitted as appropriate.

[0048] (Functional configuration) First, with reference to FIG. 7, the functional configuration of the speech recognition system 10 according to the third embodiment will be described. FIG. 7 is a block diagram showing the functional configuration of the speech recognition system according to the third embodiment. In FIG. 7, the same reference numerals are given to the same elements as those shown in FIG. 2.

[0049] As shown in FIG. 7, the speech recognition system 10 according to the third embodiment includes, as components for realizing its functions, a speech data acquisition unit 110, a text conversion unit 120, a speech synthesis unit 130, a conversion model generation unit 140, a speech conversion unit 210, a speech recognition unit 220, and a speech recognition model generation unit 310. That is, the speech recognition system 10 according to the third embodiment further includes a speech recognition model generation unit 310 in addition to the configuration of the first embodiment (see FIG. 2). Note that the speech recognition model generation unit 310 may be a processing block realized by, for example, the above-described processor 11 (see FIG. 1).

[0050] The speech recognition model generation unit 310 is configured to be able to generate a speech recognition model that converts input speech into synthesized speech. Specifically, the speech recognition model generation unit 310 is configured to be able to generate a speech recognition model using the corresponding synthesized speech generated by the speech synthesis means. Note that the speech recognition model may generate the speech recognition model using the corresponding synthesized speech and other synthesized speech. The speech recognition model generation unit 310 may be configured to directly acquire the corresponding synthesized speech from the speech synthesis unit 130, or may be configured to acquire the corresponding synthesized speech from a synthesized speech corpus that stores a plurality of corresponding synthesized speeches generated by the speech synthesis means. The speech recognition model generated by the speech recognition model generation unit 310 is configured to be output to the speech recognition unit 220.

[0051] (Speech Recognition Model Generation Operation) Next, with reference to FIG. 8, the flow of the operation for generating the speech recognition model by the speech recognition system 10 according to the third embodiment (hereinafter, appropriately referred to as the "speech recognition model generation operation") will be described. FIG. 8 is a flowchart showing the flow of the speech recognition model generation operation by the speech recognition system according to the third embodiment.

[0052] As shown in FIG. 8, when the speech recognition model generation operation by the speech recognition system 10 according to the third embodiment is started, first, the speech recognition model generation unit 310 acquires the corresponding synthesized speech generated by the speech synthesis unit 130 (step S301).

[0053] Subsequently, the speech recognition model generation unit 310 generates a speech recognition model using the acquired corresponding synthesized speech (step S302). Then, the speech recognition model generation unit 310 outputs the generated speech recognition model to the speech recognition unit 220 (step S303).

[0054] (Technical effect) Next, the technical effect obtained by the speech recognition system 10 according to the third embodiment will be described.

[0055] As described with reference to FIGS. 7 and 8, in the speech recognition system 10 according to the third embodiment, a speech recognition model is generated using the corresponding synthesized speech. In this way, since it is not necessary to separately prepare the synthesized speech for generating the speech recognition model (that is, the corresponding synthesized speech used for generating the speech conversion model can be utilized), it is possible to efficiently generate the speech recognition model.

[0056] <Fourth Embodiment> The speech recognition system 10 according to the fourth embodiment will be described with reference to FIGS. 9 and 10. Note that the fourth embodiment is only different from the above-described third embodiment in some configurations and operations, and the other parts may be the same as those of the first to third embodiments. For this reason, hereinafter, the parts different from the already described embodiments will be described in detail, and the description of the other overlapping parts will be omitted as appropriate.

[0057] (Functional Configuration) First, with reference to FIG. 9, the functional configuration of the speech recognition system 10 according to the fourth embodiment will be described. FIG. 9 is a block diagram showing the functional configuration of the speech recognition system according to the fourth embodiment. In FIG. 9, the same elements as those shown in FIG. 7 are denoted by the same reference numerals.

[0058] As shown in FIG. 9, the speech recognition system 10 according to the fourth embodiment includes, as components for realizing its functions, a speech data acquisition unit 110, a text conversion unit 120, a speech synthesis unit 130, a conversion model generation unit 140, a speech conversion unit 210, a speech recognition unit 220, and a speech recognition model generation unit 310. In particular, in the fourth embodiment, the speech recognition model generation unit 310 is configured to receive the synthesized speech converted by the speech conversion unit 210 and the recognition result by the speech recognition unit 220. The speech recognition model generation unit 310 according to the fourth embodiment is configured to be able to execute learning of the speech recognition model based on the synthesized speech converted by the speech conversion unit 210 and the recognition result by the speech recognition unit 220.

[0059] (Speech Recognition Model Learning Operation) Next, with reference to FIG. 10, the flow of the operation (hereinafter, appropriately referred to as "speech recognition model learning operation") for learning the speech recognition model by the speech recognition system 10 according to the fourth embodiment will be described. FIG. 10 is a flowchart showing the flow of the speech recognition model learning operation by the speech recognition system according to the third embodiment.

[0060] As shown in FIG. 10, when the speech recognition model learning operation by the speech recognition system 10 according to the fourth embodiment is started, first, the speech recognition model generation unit 310 acquires the synthesized speech converted by the speech conversion unit 210 (that is, the synthesized speech input to the speech recognition unit 220) (step S401). Then, the speech recognition model generation unit 310 further acquires the speech recognition result of the synthesized speech (that is, the speech recognition result output in step S156 shown in FIG. 4) (step S402).

[0061] Subsequently, based on the acquired synthesized speech and the speech recognition result, the speech recognition model generation unit 310 learns the speech recognition model (step S403). At this time, the speech recognition model generation unit 310 may adjust the parameters of the conversion model that has already been generated. After that, the speech recognition model generation unit 310 outputs the learned speech recognition model to the speech conversion unit 210 (step S404).

[0062] (Technical Effect) Next, the technical effect obtained by the speech recognition system 10 according to the fourth embodiment will be described.

[0063] As described with reference to FIGS. 9 and 10, in the speech recognition system 10 according to the fourth embodiment, the conversion model is learned based on the synthesized speech and the speech recognition result. In this way, since the learning is performed in consideration of how the synthesized speech is actually recognized by speech recognition, the speech recognition model can be learned so that more appropriate speech recognition can be performed. Specifically, the speech recognition model can be learned so that the accuracy of speech recognition is improved.

[0064] <Fifth Embodiment> The speech recognition system 10 according to the fifth embodiment will be described with reference to FIGS. 11 and 12. Note that the fifth embodiment is only different from the first to fourth embodiments described above in some configurations and operations, and the other parts may be the same as those of the first to fourth embodiments. Therefore, hereinafter, the parts different from the embodiments already described will be described in detail, and the description of the other overlapping parts will be omitted as appropriate.

[0065] (Functional Configuration) First, with reference to FIG. 11, the functional configuration of the speech recognition system 10 according to the fifth embodiment will be described. FIG. 11 is a block diagram showing the functional configuration of the speech recognition system according to the fifth embodiment. In FIG. 11, the same reference numerals are given to the elements similar to those shown in FIG. 2.

[0066] As shown in FIG. 11, the speech recognition system 10 according to the fifth embodiment includes, as components for realizing its functions, a speech data acquisition unit 110, a text conversion unit 120, a speech synthesis unit 130, a conversion model generation unit 140, an attribute information acquisition unit 150, a voice conversion unit 210, and a speech recognition unit 220. That is, the speech recognition system 10 according to the fifth embodiment further includes an attribute information acquisition unit 150 in addition to the configuration of the first embodiment (see FIG. 2). Note that the attribute information acquisition unit 150 may be a processing block realized by, for example, the above-described processor 11 (see FIG. 1).

[0067] The attribute information acquisition unit 150 is configured to be able to acquire attribute information about the speaker of the real speech data. The attribute information may include, for example, information about the gender, age, occupation, etc. of the speaker. The attribute information acquisition unit 150 may be configured to be able to acquire attribute information from, for example, a terminal or ID card held by the speaker. Alternatively, the attribute information acquisition unit 150 may be configured to acquire attribute information input by the speaker. The attribute information acquired by the attribute information acquisition unit 150 is output to the speech synthesis unit 130. The attribute information may be stored in the real speech corpus in a state associated with the real speech data. In this case, the attribute information may be configured to be output from the real speech corpus to the speech synthesis unit 130.

[0068] (Conversion Model Generation Operation) Next, with reference to FIG. 12, the flow of the conversion model generation operation by the speech recognition system 10 according to the fifth embodiment will be described. FIG. 12 is a flowchart showing the flow of the conversion model generation operation by the speech recognition system according to the fifth embodiment. In FIG. 12, the same reference numerals are given to the same processes as those shown in FIG. 3.

[0069] As shown in FIG. 12, when the conversion model generation operation by the speech recognition system 10 according to the fifth embodiment is started, first, the utterance data acquisition unit 110 acquires real utterance data (step S101). Then, the attribute information acquisition unit 150 acquires attribute information regarding the speaker of the real utterance data (step S501). Note that the processes of steps S101 and S102 may be executed successively or simultaneously in parallel.

[0070] Subsequently, the text conversion unit 120 converts the real utterance data acquired by the utterance data acquisition unit 110 into text data (step S102). Thereafter, the speech synthesis unit 130 synthesizes the text data converted by the text conversion unit 120 to generate a corresponding synthesized speech corresponding to the real utterance data. In particular, in this embodiment, speech synthesis is also performed using the attribute information (step S502). For example, the speech synthesis unit 130 may perform speech synthesis considering the gender, age, occupation, etc. of the speaker of the real utterance data.

[0071] Subsequently, the conversion model generation unit 140 generates a conversion model based on the real utterance data acquired by the utterance data acquisition unit 110 and the corresponding synthesized speech generated by the speech synthesis unit 130 (here, the synthesized speech synthesized based on the attribute information) (step S104). Note that the set of the real utterance data and the corresponding synthesized speech input to the conversion model generation unit 140 may be provided with attribute information. In that case, the conversion model generation unit 140 may generate a conversion model in consideration of the attribute information. Thereafter, the conversion model generation unit 140 outputs the generated conversion model to the speech conversion unit 210 (step S105).

[0072] (Technical Effect) Next, the technical effect obtained by the speech recognition system 10 according to the fifth embodiment will be described.

[0073] As described with reference to FIGS. 11 and 12, in the speech recognition system 10 according to the fifth embodiment, a corresponding synthesized speech is generated using the speaker's attribute information. In this way, since the corresponding synthesized speech is generated in consideration of the speaker's attributes, it becomes possible to generate a more appropriate speech conversion model. Also, when generating a speech recognition model using the corresponding synthesized speech as in the third embodiment described above (see FIGS. 7 and 8), since the corresponding synthesized speech in which the attributes are considered is used, it becomes possible to generate a more appropriate speech recognition model.

[0074] <Sixth Embodiment> The speech recognition system 10 according to the sixth embodiment will be described with reference to FIGS. 13 and 14. Note that the sixth embodiment is only different from the first to fifth embodiments described above in some configurations and operations, and the other parts may be the same as those of the first to fifth embodiments. For this reason, in the following, the parts different from the embodiments already described will be described in detail, and the description of the other overlapping parts will be omitted as appropriate.

[0075] (Functional Configuration) First, with reference to FIG. 13, the functional configuration of the speech recognition system 10 according to the sixth embodiment will be described. FIG. 13 is a block diagram showing the functional configuration of the speech recognition system according to the sixth embodiment. Note that in FIG. 13, the same reference numerals are given to the same elements as those shown in FIG. 11.

[0076] As shown in FIG. 13, the speech recognition system 10 according to the sixth embodiment includes, as components for realizing its functions, a plurality of real uttered speech corpora 105a, 105b, and 105c (hereinafter, collectively referred to as "real uttered speech corpus 105" as appropriate), an utterance data acquisition unit 110, a text conversion unit 120, a speech synthesis unit 130, a conversion model generation unit 140, a speech conversion unit 210, and a speech recognition unit 220. That is, the speech recognition system 10 according to the sixth embodiment further includes a plurality of real uttered speech corpora 105 in addition to the configuration of the first embodiment (see FIG. 2). Note that the plurality of real uttered speech corpora 105 may be configured by, for example, the above-described storage device 14 (see FIG. 1).

[0077] The plurality of real uttered speech corpora 105 store real uttered data for each predetermined condition. The "predetermined condition" here is, for example, a condition set for classifying real uttered data. For example, each of the plurality of real uttered speech corpora 105 may store real uttered data by field. In this case, the real uttered speech corpus 105a may store real uttered data related to the field of law, the real uttered speech corpus 105b may store real uttered data related to the field of science, and the real uttered speech corpus 105c may be configured to store real uttered data related to the field of medicine. Note that, for convenience of explanation, three real uttered speech corpora 105 are illustrated here, but the number of real uttered speech corpora 105 is not particularly limited.

[0078] The utterance data acquisition unit 110 according to the sixth embodiment is configured to be able to select one from the plurality of real utterance voice corpora 105 described above and acquire real utterance data. Note that information regarding the real utterance voice corpus 105 selected here (specifically, information regarding a predetermined condition) may be output to the conversion model generation unit 140 together with the real utterance data. Then, the conversion model generation unit 140 may use information regarding the real utterance voice corpus 105 selected when generating the conversion model. Also, in the configuration for generating an acoustic model as in the third embodiment described above, information regarding the selected real utterance voice corpus 105 may be output to the acoustic model generation unit 310. Then, the acoustic model generation unit 310 may use information regarding the real utterance voice corpus 105 selected when generating the acoustic model.

[0079] (Conversion model generation operation) Next, with reference to FIG. 14, the flow of the conversion model generation operation by the speech recognition system 10 according to the sixth embodiment will be described. FIG. 14 is a flowchart showing the flow of the conversion model generation operation by the speech recognition system according to the sixth embodiment. Note that in FIG. 14, the same reference numerals are assigned to the same processes as those shown in FIG. 12.

[0080] As shown in FIG. 14, when the conversion model generation operation by the speech recognition system 10 according to the sixth embodiment is started, first, the utterance data acquisition unit 110 selects a corpus from which to acquire utterance data from among the plurality of real utterance voice corpora 105 (step S601). Then, the utterance data acquisition unit 110 acquires real utterance data from the selected real utterance voice corpus (step S602).

[0081] Subsequently, the text conversion unit 120 converts the real utterance data acquired by the utterance data acquisition unit 110 into text data (step S102). Then, the speech synthesis unit 130 synthesizes the text data converted by the text conversion unit 120 to generate a corresponding synthesized voice corresponding to the real utterance data (step S103).

[0082] Subsequently, the conversion model generation unit 140 generates a conversion model based on the real speech data acquired by the speech data acquisition unit 110 and the corresponding synthesized speech generated by the speech synthesis unit 130. In this embodiment, in particular, information regarding the selected real speech corpus is also used (step S606). Thereafter, the conversion model generation unit 140 outputs the generated conversion model to the speech conversion unit 210 (step S105).

[0083] (Technical Effect) Next, the technical effect obtained by the speech recognition system 10 according to the sixth embodiment will be described.

[0084] As described with reference to FIGS. 13 and 14, in the speech recognition system 10 according to the sixth embodiment, information regarding the real speech corpus 105 selected when acquiring real speech data is used when generating the conversion model. By doing so, since predetermined conditions (for example, fields) used for classifying the real speech data are taken into consideration, it becomes possible to generate a more appropriate conversion model.

[0085] <Seventh Embodiment> The speech recognition system 10 according to the seventh embodiment will be described with reference to FIGS. 15 and 16. Note that the seventh embodiment is different only in some configurations and operations from the first to sixth embodiments described above, and the other parts may be the same as those of the first to sixth embodiments. Therefore, hereinafter, the parts different from the embodiments already described will be described in detail, and the description of the other overlapping parts will be omitted as appropriate.

[0086] (Functional Configuration) First, with reference to FIG. 15, the functional configuration of the speech recognition system 10 according to the seventh embodiment will be described. FIG. 15 is a block diagram showing the functional configuration of the speech recognition system according to the seventh embodiment. In FIG. 15, the same reference numerals are given to the same elements as those shown in FIG. 2.

[0087] As shown in FIG. 15, the speech recognition system 10 according to the seventh embodiment includes, as components for realizing its functions, a speech data acquisition unit 110, a text conversion unit 120, a speech synthesis unit 130, a conversion model generation unit 140, a noise addition unit 160, a speech conversion unit 210, and a speech recognition unit 220. That is, the speech recognition system 10 according to the seventh embodiment further includes a noise addition unit 160 in addition to the configuration of the first embodiment (see FIG. 2). Note that the noise addition unit 160 may be a processing block realized by, for example, the above-described processor 11 (see FIG. 1).

[0088] The noise addition unit 160 is configured to be able to add noise to the text data generated by the text conversion unit 120. The noise addition unit 160 may, for example, add noise to the real speech data before text conversion so that the text data is provided with noise, or may add noise to the text data after text conversion. Alternatively, the noise addition unit 160 may add noise when the text conversion unit 120 performs text conversion on the real speech data. The noise addition unit 160 may add preset noise or may add randomly set noise.

[0089] (Conversion Model Generation Operation) Next, with reference to FIG. 16, the flow of the conversion model generation operation by the speech recognition system 10 according to the seventh embodiment will be described. FIG. 16 is a flowchart showing the flow of the conversion model generation operation by the speech recognition system according to the seventh embodiment. In FIG. 16, the same reference numerals are given to the same processes as those shown in FIG. 3.

[0090] As shown in FIG. 16, when the conversion model generation operation by the speech recognition system 10 according to the seventh embodiment is started, first, the utterance data acquisition unit 110 acquires real utterance data (step S101). Here, in particular in this embodiment, the noise addition unit 160 outputs noise information to the text conversion unit 120 (step S701). Then, the text conversion unit 120 converts the real utterance data acquired by the utterance data acquisition unit 110 into text data with noise added (step S702).

[0091] Subsequently, the speech synthesis unit 130 synthesizes the text data converted by the text conversion unit 120 (here, text data with noise added) into speech, and generates a corresponding synthesized speech corresponding to the real utterance data (step S103). Then, the conversion model generation unit 140 generates a conversion model based on the real utterance data acquired by the utterance data acquisition unit 110 and the corresponding synthesized speech generated by the speech synthesis unit 130 (step S104). After that, the conversion model generation unit 140 outputs the generated conversion model to the speech conversion unit 210 (step S105).

[0092] (Technical Effect) Next, the technical effect obtained by the speech recognition system 10 according to the seventh embodiment will be described.

[0093] As described with reference to FIGS. 15 and 16, in the speech recognition system 10 according to the seventh embodiment, the real utterance data is converted into text data with noise added. In this way, since the conversion model is generated using data including noise, it is possible to generate a conversion model that is robust to noise (for example, a conversion model that can appropriately perform speech conversion even when the input speech includes noise).

[0094] <Modification Example of the Seventh Embodiment> The voice recognition system 10 according to a modification of the seventh embodiment will be described with reference to FIGS. 17 and 18. Note that the modification of the seventh embodiment is only different from the above-described seventh embodiment in some configurations and operations, and other parts may be the same as those of the first to seventh embodiments. Therefore, hereinafter, the parts different from the respective embodiments already described will be described in detail, and the description of other overlapping parts will be omitted as appropriate.

[0095] (Functional configuration) First, with reference to FIG. 17, the functional configuration of the voice recognition system 10 according to a modification of the seventh embodiment will be described. FIG. 17 is a block diagram showing the functional configuration of the voice recognition system according to a modification of the seventh embodiment. In FIG. 17, the same reference numerals are given to the same elements as those shown in FIG. 15.

[0096] As shown in FIG. 17, the voice recognition system 10 according to a modification of the seventh embodiment includes, as components for realizing its functions, a speech data acquisition unit 110, a text conversion unit 120, a speech synthesis unit 130, a conversion model generation unit 140, a noise addition unit 160, a voice conversion unit 210, and a voice recognition unit 220. However, in the voice recognition system 10 according to a modification of the seventh embodiment, the noise addition unit 160 is configured to be able to output noise information to the speech synthesis unit 130. That is, in the modification of the seventh embodiment, noise is added during speech synthesis by the speech synthesis unit 130.

[0097] (Conversion model generation operation) Next, with reference to FIG. 18, the flow of the conversion model generation operation by the voice recognition system 10 according to a modification of the seventh embodiment will be described. FIG. 18 is a flowchart showing the flow of the conversion model generation operation by the voice recognition system according to a modification of the seventh embodiment. In FIG. 18, the same reference numerals are given to the same processes as those shown in FIG. 16.

[0098] As shown in FIG. 18, when the conversion model generation operation by the speech recognition system 10 according to the modification of the seventh embodiment is started, first, the utterance data acquisition unit 110 acquires real utterance data (step S101). Then, the text conversion unit 120 converts the real utterance data acquired by the utterance data acquisition unit 110 into text data (step S102).

[0099] Subsequently, in this embodiment, in particular, the noise addition unit 160 outputs noise information to the speech synthesis unit 130 (step S751). Then, the speech synthesis unit 130 synthesizes the text data converted by the text conversion unit 120 to generate a corresponding synthesized speech with noise added (step S752).

[0100] Subsequently, the conversion model generation unit 140 generates a conversion model based on the real utterance data acquired by the utterance data acquisition unit 110 and the corresponding synthesized speech generated by the speech synthesis unit 130 (here, the corresponding synthesized speech with noise added) (step S104). Then, the conversion model generation unit 140 outputs the generated conversion model to the speech conversion unit 210 (step S105).

[0101] (Technical Effect) Next, the technical effect obtained by the speech recognition system 10 according to the modification of the seventh embodiment will be described.

[0102] As described with reference to FIGS. 17 and 18, in the speech recognition system 10 according to the modification of the seventh embodiment, a corresponding synthesized speech with noise added is generated. In this way, since the conversion model is generated using data including noise, it is possible to generate a conversion model that is robust to noise (for example, a conversion model that can appropriately perform speech conversion even if the input speech includes noise).

[0103] <Eighth Embodiment> The speech recognition system 10 according to the eighth embodiment will be described with reference to FIGS. 19 to 21. Note that the eighth embodiment is only different from the first to seventh embodiments described above in some configurations and operations, and the other parts may be the same as those of the first to seventh embodiments. For this reason, in the following, the parts different from the embodiments already described will be described in detail, and the description of the other overlapping parts will be omitted as appropriate.

[0104] (Functional Configuration) First, with reference to FIG. 19, the functional configuration of the speech recognition system 10 according to the eighth embodiment will be described. FIG. 19 is a block diagram showing the functional configuration of the speech recognition system according to the eighth embodiment.

[0105] As shown in FIG. 19, the speech recognition system 10 according to the eighth embodiment includes, as components for realizing its functions, a sign language data acquisition unit 410, a text conversion unit 420, a speech synthesis unit 430, a conversion model generation unit 440, a voice conversion unit 510, and a speech recognition unit 520. Each of the sign language data acquisition unit 410, the text conversion unit 420, the speech synthesis unit 430, the conversion model generation unit 440, the voice conversion unit 510, and the speech recognition unit 520 may be a processing block realized by, for example, the above-described processor 11 (see FIG. 1).

[0106] The sign language data acquisition unit 410 is configured to be able to acquire sign language utterance data. The sign language data may be, for example, video data of sign language. The sign language data may be acquired from, for example, a database (sign language corpus) that stores a plurality of sign language data. The sign language data acquired by the sign language data acquisition unit 410 is output to the text conversion unit 120 and the conversion model generation unit 140.

[0107] The text conversion unit 420 is configured to be able to convert the sign language data acquired by the sign language data acquisition unit 410 into text data. That is, the text conversion unit 420 is configured to be able to execute a process of text-converting the content of the sign language included in the sign language data. Regarding the specific method of text conversion, existing technologies may be appropriately adopted. The text data converted by the text conversion unit 420 (that is, the text data corresponding to the sign language data) is configured to be output to the speech synthesis unit 430.

[0108] The speech synthesis unit 430 is configured to be able to generate a corresponding synthesized speech corresponding to the sign language data by synthesizing the text data changed by the text conversion unit 420. Regarding the specific method of speech synthesis, existing technologies can be appropriately adopted. The corresponding synthesized speech generated by the speech synthesis unit 430 is configured to be output to the conversion model generation unit 440. Note that the corresponding synthesized speech may be accumulated in a database (synthesized speech corpus) capable of accumulating a plurality of corresponding syntheses and then output to the conversion model generation unit 440.

[0109] The conversion model generation unit 440 is configured to be able to generate a conversion model that converts an input sign language into synthesized speech using the sign language data acquired by the sign language data acquisition unit 410 and the corresponding synthesized speech synthesized by the speech synthesis unit 430. The conversion model converts, for example, an input sign language (for example, a video of sign language) into synthesized speech (that is, mechanical speech). The conversion model generation unit 440 may be configured to generate a conversion model using, for example, a GAN. The conversion model generated by the conversion model generation unit 440 is configured to be output to the speech conversion unit 510.

[0110] The speech conversion unit 510 is configured to be able to convert an input sign language into synthesized speech using the conversion model generated by the conversion model generation unit 440. The input sign language input to the speech conversion unit 510 may be, for example, a video input using a camera or the like. The synthesized speech converted by the speech conversion unit 510 is configured to be output to the speech recognition unit 520.

[0111] The voice recognition unit 520 is configured to be able to recognize the synthesized voice converted by the voice conversion unit 510. That is, the voice recognition unit 520 is configured to be able to execute a process of converting the synthesized voice into text. The voice recognition unit 520 may be configured to be able to output the voice recognition result of the synthesized voice. Note that the usage method of the voice recognition result is not particularly limited.

[0112] (Conversion model generation operation) Next, with reference to FIG. 20, the flow of the conversion model generation operation by the voice recognition system 10 according to the eighth embodiment will be described. FIG. 20 is a flowchart showing the flow of the conversion model generation operation by the voice recognition system according to the eighth embodiment.

[0113] As shown in FIG. 20, when the conversion model generation operation by the voice recognition system 10 according to the eighth embodiment is started, first, the sign language data acquisition unit 410 acquires sign language data (step S801). Then, the text conversion unit 420 converts the sign language data acquired by the sign language data acquisition unit 410 into text data (step S802).

[0114] Subsequently, the voice synthesis unit 430 synthesizes the text data converted by the text conversion unit 420 to generate a corresponding synthesized voice corresponding to the sign language data (step S403). Then, the conversion model generation unit 140 generates a conversion model based on the sign language data acquired by the sign language data acquisition unit 410 and the corresponding synthesized voice generated by the voice synthesis unit 430 (step S804). Thereafter, the conversion model generation unit 440 outputs the generated conversion model to the voice conversion unit 510 (step S805).

[0115] (Conversion recognition operation) Next, with reference to FIG. 21, the flow of the voice recognition operation by the voice recognition system 10 according to the eighth embodiment will be described. FIG. 21 is a flowchart showing the flow of the voice recognition operation by the voice recognition system according to the eighth embodiment.

[0116] As shown in FIG. 21, when the voice recognition operation by the voice recognition system 10 according to the first embodiment is started, first, the voice conversion unit 510 acquires the input sign language (step S851). Then, the voice conversion unit 510 reads the conversion model generated by the conversion model generation unit 440 (step S852). After that, the voice conversion unit 210 performs voice conversion using the read conversion model, and converts the input sign language into synthesized voice (step S853).

[0117] Subsequently, the voice recognition unit 520 reads the voice recognition model (step S854). Then, the voice recognition unit 520 performs voice recognition on the synthesized voice synthesized by the voice conversion unit 510 using the read voice recognition model (step S855). After that, the voice recognition unit 520 outputs the voice recognition result (step S856).

[0118] (Technical Effect) Next, the technical effect obtained by the voice recognition system 10 according to the eighth embodiment will be described.

[0119] As described with reference to FIGS. 19 to 21, in the voice recognition system 10 according to the eighth embodiment, when generating the conversion model, sign language data and corresponding synthesized voice corresponding to the sign language data are used. In particular, the corresponding synthesized voice corresponding to the sign language data is generated by text-converting the sign language data and synthesizing the text data into voice. In this way, it is not necessary to prepare both the sign language data and the corresponding synthesized voice (that is, if only the sign language data is prepared, the corresponding synthesized voice can be generated), so the cost required to generate the conversion model can be suppressed. As a result, it is possible to realize voice recognition with high recognition accuracy at low cost.

[0120] A program that operates the configuration of the above-described embodiments to realize the functions of the respective embodiments is recorded on a recording medium, the program recorded on the recording medium is read as code, and a processing method executed on a computer is also included in the scope of each embodiment. That is, a computer-readable recording medium is also included in the scope of each embodiment. Further, not only the recording medium on which the above-described program is recorded, but also the program itself is included in each embodiment.

[0121] As the recording medium, for example, a floppy (registered trademark) disk, a hard disk, an optical disk, a magneto-optical disk, a CD-ROM, a magnetic tape, a non-volatile memory card, or a ROM can be used. Further, not only those that execute processing with the program recorded on the recording medium alone, but also those that operate on an OS and execute processing in cooperation with the functions of other software and expansion boards are included in the scope of each embodiment. Furthermore, the program itself may be stored in a server so that a part or all of the program can be downloaded from the server to the user terminal.

[0122] <Supplementary Note> Regarding the embodiments described above, they can be further described as follows in the supplementary notes below, but are not limited thereto.

[0123] (Supplementary Note 1) The speech recognition system described in Supplementary Note 1 includes a speech data acquisition means for acquiring real speech data spoken by a speaker, a text conversion means for converting the real speech data into text data, a speech synthesis means for generating a corresponding synthesized speech corresponding to the real speech data by speech synthesis using the text data, a conversion model generation means for generating a conversion model for converting an input speech into a synthesized speech using the real speech data and the corresponding synthesized speech, and a speech recognition means for performing speech recognition on the synthesized speech converted using the conversion model.

[0124] (Supplementary Note 2) The voice recognition system described in Supplementary Note 2 is the voice recognition system described in Supplementary Note 1, in which the conversion model generation means adjusts the parameters of the conversion model using the input voice and the recognition result of the voice recognition means.

[0125] (Supplementary Note 3) The voice recognition system described in Supplementary Note 3 further includes a voice recognition model generation means for generating a voice recognition model using the data including the corresponding synthesized voice, and the voice recognition means performs voice recognition using the voice recognition model, which is the voice recognition system described in Supplementary Note 1 or 2.

[0126] (Supplementary Note 4) The voice recognition system described in Supplementary Note 4 is the voice recognition system described in Supplementary Note 3, in which the voice recognition model generation means adjusts the parameters of the voice recognition model using the synthesized voice converted using the conversion model and the recognition result of the voice recognition means.

[0127] (Supplementary Note 5) The voice recognition system described in Supplementary Note 5 further includes an attribute acquisition means for acquiring attribute information indicating the attributes of the speaker, and the voice synthesis means generates the corresponding synthesized voice by performing voice synthesis using the attribute information, which is the voice recognition system described in any one of Supplementary Notes 1 to 4.

[0128] (Supplementary Note 6) The voice recognition system described in Supplementary Note 6 further includes a plurality of real utterance voice corpora for storing the real utterance data for each predetermined condition, and the utterance data acquisition means acquires the real utterance data by selecting one from the plurality of real utterance voice corpora, which is the voice recognition system described in any one of Supplementary Notes 1 to 5.

[0129] (Supplementary Note 7) The voice recognition system described in Supplementary Note 7 further includes a noise addition means for adding noise to at least one of the text data and the corresponding synthesized voice, which is the voice recognition system described in any one of Supplementary Notes 1 to 6.

[0130] (Appendix 8) The speech recognition system described in Appendix 8 includes a sign language data acquisition means for acquiring sign language data, a text conversion means for converting the sign language data into text data, a speech synthesis means for generating a corresponding synthesized speech corresponding to the sign language data by speech synthesis using the text data, a conversion model generation means for generating a conversion model for converting an input sign language into a synthesized speech using the sign language data and the corresponding synthesized speech, and a speech recognition means for performing speech recognition on the synthesized speech converted using the conversion model.

[0131] (Appendix 9) The speech recognition method described in Appendix 9 is a speech recognition method in which at least one computer acquires real speech data spoken by a speaker, converts the real speech data into text data, generates a corresponding synthesized speech corresponding to the real speech data by speech synthesis using the text data, generates a conversion model for converting an input speech into a synthesized speech using the real speech data and the corresponding synthesized speech, and performs speech recognition on the synthesized speech converted using the conversion model.

[0132] (Appendix 10) The recording medium described in Appendix 10 is a recording medium on which a computer program for causing at least one computer to execute a speech recognition method is recorded, the speech recognition method including acquiring real speech data spoken by a speaker, converting the real speech data into text data, generating a corresponding synthesized speech corresponding to the real speech data by speech synthesis using the text data, generating a conversion model for converting an input speech into a synthesized speech using the real speech data and the corresponding synthesized speech, and performing speech recognition on the synthesized speech converted using the conversion model.

[0133] (Appendix 11) The computer program described in Supplementary Note 11 causes at least one computer to execute a voice recognition method, which includes obtaining real speech data spoken by a speaker, converting the real speech data into text data, generating corresponding synthesized speech corresponding to the real speech data by voice synthesis using the text data, generating a conversion model for converting input speech into synthesized speech using the real speech data and the corresponding synthesized speech, and performing voice recognition on the synthesized speech converted using the conversion model.

[0134] (Supplementary Note 12) The voice recognition device described in Supplementary Note 12 includes a speech data acquisition unit that acquires real speech data spoken by a speaker, a text conversion unit that converts the real speech data into text data, a voice synthesis unit that generates corresponding synthesized speech corresponding to the real speech data by voice synthesis using the text data, a conversion model generation unit that generates a conversion model for converting input speech into synthesized speech using the real speech data and the corresponding synthesized speech, and a voice recognition unit that performs voice recognition on the synthesized speech converted using the conversion model.

[0135] (Supplementary Note 13) The voice recognition method described in Supplementary Note 13 includes obtaining sign language data by at least one computer, converting the sign language data into text data, generating corresponding synthesized speech corresponding to the sign language data by voice synthesis using the text data, generating a conversion model for converting input sign language into synthesized speech using the sign language data and the corresponding synthesized speech, and performing voice recognition on the synthesized speech converted using the conversion model.

[0136] (Supplementary Note 14) The recording medium described in Supplementary Note 14 causes at least one computer to acquire sign language data, convert the sign language data into text data, generate corresponding synthesized speech corresponding to the sign language data by voice synthesis using the text data, generate a conversion model for converting input sign language into synthesized speech using the sign language data and the corresponding synthesized speech, and perform speech recognition on the synthesized speech converted using the conversion model, and is a recording medium on which a computer program for executing a speech recognition method is recorded.

[0137] (Supplementary Note 15) The computer program described in Supplementary Note 15 causes at least one computer to acquire sign language data, convert the sign language data into text data, generate corresponding synthesized speech corresponding to the sign language data by voice synthesis using the text data, generate a conversion model for converting input sign language into synthesized speech using the sign language data and the corresponding synthesized speech, and perform speech recognition on the synthesized speech converted using the conversion model, and is a computer program for executing a speech recognition method.

[0138] (Supplementary Note 16) The speech recognition device described in Supplementary Note 16 includes sign language data acquisition means for acquiring sign language data, text conversion means for converting the sign language data into text data, voice synthesis means for generating corresponding synthesized speech corresponding to the sign language data by voice synthesis using the text data, conversion model generation means for generating a conversion model for converting input sign language into synthesized speech using the sign language data and the corresponding synthesized speech, and speech recognition means for performing speech recognition on the synthesized speech converted using the conversion model, and is a speech recognition device.

[0139] This disclosure can be appropriately changed within a range not contrary to the gist or idea of the invention that can be read from the claims and the entire specification, and a speech recognition system, a speech recognition method, and a recording medium with such changes are also included in the technical idea of this disclosure.

Explanation of Reference Numerals

[0140] 10 Voice Recognition System 11 Processor 14 Memory Device 105 Real Spoken Voice Corpus 110 Speech Data Acquisition Unit 120 Text Conversion Unit 130 Speech Synthesis Unit 140 Conversion Model Generation Unit 150 Attribute Information Acquisition Unit 160 Noise Addition Unit 210 Voice Conversion Unit 220 Voice Recognition Unit 310 Voice Recognition Model Generation Unit 410 Sign Language Data Acquisition Unit 420 Text Conversion Unit 430 Speech Synthesis Unit 440 Conversion Model Generation Unit 510 Voice Conversion Unit 520 Voice Recognition Unit

Claims

1. An utterance data acquisition means for acquiring real utterance data uttered by a speaker; A text conversion means for converting the real utterance data into text data; An audio synthesis means for generating a corresponding synthesized audio corresponding to the real utterance data by audio synthesis using the text data; A conversion model generation means for generating a conversion model that converts input audio into synthesized audio using the real utterance data and the corresponding synthesized audio; An audio recognition means for audio-recognizing the synthesized audio converted using the conversion model; An audio recognition system comprising the above.

2. The conversion model generation means adjusts the parameters of the conversion model using the input audio and the recognition result of the audio recognition means. The audio recognition system according to Claim 1.

3. Further comprising an audio recognition model generation means for generating an audio recognition model using data including the corresponding synthesized audio, The audio recognition means performs audio recognition using the audio recognition model. The audio recognition system according to Claim 1 or 2.

4. The audio recognition model generation means adjusts the parameters of the audio recognition model using the synthesized audio converted using the conversion model and the recognition result of the audio recognition means. The audio recognition system according to Claim 3.

5. Further comprising an attribute acquisition means for acquiring attribute information indicating the attributes of the speaker, The audio synthesis means generates the corresponding synthesized audio by performing audio synthesis using the attribute information. The audio recognition system according to any one of Claims 1 to 4.

6. Further comprising a plurality of real utterance audio corpora for storing the real utterance data for each predetermined condition, The utterance data acquisition means selects one from the plurality of real utterance audio corpora to acquire the real utterance data. The audio recognition system according to any one of Claims 1 to 5.

7. Further comprising a noise addition means for adding noise to at least one of the text data and the corresponding synthesized audio. The audio recognition system according to any one of Claims 1 to 6.

8. A sign language data acquisition means for acquiring sign language data; A text conversion means for converting the sign language data into text data; An audio synthesis means for generating a corresponding synthesized audio corresponding to the sign language data by audio synthesis using the text data; Conversion model generation means for generating a conversion model that converts input sign language into synthesized speech using the sign language data and the corresponding synthesized speech; Speech recognition means for speech-recognizing the synthesized speech converted using the conversion model; A speech recognition system comprising the above.

9. By at least one computer, Obtain real speech data spoken by a speaker, Convert the real speech data into text data, Generate corresponding synthesized speech corresponding to the real speech data by speech synthesis using the text data, Generate a conversion model that converts input speech into synthesized speech using the real speech data and the corresponding synthesized speech, Speech-recognize the synthesized speech converted using the conversion model, Speech recognition method.

10. On at least one computer, Obtain real speech data spoken by a speaker, Convert the real speech data into text data, Generate corresponding synthesized speech corresponding to the real speech data by speech synthesis using the text data, Generate a conversion model that converts input speech into synthesized speech using the real speech data and the corresponding synthesized speech, Speech-recognize the synthesized speech converted using the conversion model, A computer program for executing a speech recognition method.

Citation Information

Patent Citations

  • Method and apparatus for converting sign language into speech

    JP2003522978A

  • Voice quality conversion system, voice quality conversion method and voice quality conversion program

    JP2019008120A

  • System for improving dysarthria speech intelligibility and method thereof

    JP2020166224A

  • Speech processing system and terminal device

    WO2014010450A1

  • Voice conversion device, voice conversion method, and voice conversion program

    WO2021033685A1