Training data generation device, speech recognition model generation device, training data generation method, and program

The training data generation device and method address the limitations of synthesized speech and user-specific training by using actual speech and differential analysis to enhance speech recognition accuracy.

JP7764964B2Active Publication Date: 2025-11-06NEC CORP
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
JP2024532085
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2022-07-04
Filing Date
2023-06-29
Publication Date
2025-11-06
Estimated Expiration
2043-06-29

AI Technical Summary

Technical Problem

Existing speech recognition systems face limitations in improving recognition accuracy due to the use of synthesized speech data and user-specific training data, which hinders generalization across different users.

Method used

A training data generation device and method that utilizes a first speech recognition model trained with synthetic speech and standard text information to generate training data, and a second model to determine differences in output results, enabling the creation of high-accuracy training data for improving the second model's recognition accuracy.

Benefits of technology

Enhances speech recognition accuracy by using actual speech data and differential analysis to refine the training process, resulting in improved recognition models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007764964000001
    Figure 0007764964000001
  • Figure 0007764964000002
    Figure 0007764964000002
  • Figure 0007764964000003
    Figure 0007764964000003
Patent Text Reader

Abstract

A training data generation device (10) comprises: a voice recognition unit (140); and a generation unit (160). The voice recognition unit (140) generates text information by inputting voice data into a first voice recognition model which has already been trained. The generation unit (160) generates training data that includes the voice data and the text information. The first voice recognition model is a model that has been trained, using a synthetic sound which was generated by using input information relating to a predetermined item and previously prepared formatted text information.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a training data generation device, a speech recognition model generation device, a training data generation method, a speech recognition model generation method, and a recording medium. [Background technology]

[0002] In order to obtain a trained model for speech recognition, a large amount of training data must be prepared.

[0003] Patent Document 1 describes generating synthetic speech of optimal sentence examples for added words as speech data to be used for training a speech recognition system. Patent Document 1 also describes generating optimal sentence examples using sentence example templates.

[0004] Patent Document 2 describes a method of recognizing a user's speech using a recognition engine trained with training data for each user, and generating training data including the speech and the recognition results. [Prior art documents] [Patent documents]

[0005] [Patent Document 1] International Publication No. 2021 / 215352 [Patent Document 2] International Publication No. 2021 / 059968 Summary of the Invention [Problem to be solved by the invention]

[0006] In the above-mentioned Patent Document 1, the synthesized speech generated by the program is used as speech data for training a system that recognizes human voices. Therefore, there is a problem that there is a limit to the improvement of recognition accuracy by training using that speech data. Furthermore, the above-mentioned Patent Document 2 generates training data for each user, which makes it difficult to improve the recognition accuracy of speech recognition regardless of the user.

[0007] In view of the above-mentioned problems, an example of an object of the present invention is to provide a training data generation device, a speech recognition model generation device, a training data generation method, a speech recognition model generation method, and a recording medium that improve the recognition accuracy of a speech recognition model. [Means for solving the problem]

[0008] According to one aspect of the present invention, a speech recognition means for generating text information by inputting speech data into a trained first speech recognition model; generating means for generating training data including the voice data and the text information; The first speech recognition model is a model that has been trained using synthetic speech generated using input information related to a predetermined item and standard text information prepared in advance. A training data generation device is provided.

[0009] According to one aspect of the present invention, a speech recognition means for inputting speech data into a trained first speech recognition model and a trained second speech recognition model, and generating output results of the first speech recognition model and the second speech recognition model, respectively; a determination means for determining whether there is a difference between the output result of the first speech recognition model and the output result of the second speech recognition model; a generation means for generating training data including the speech data when the determination means determines that there is a difference between the output result of the first speech recognition model and the output result of the second speech recognition model. A training data generation device is provided.

[0010] According to one aspect of the present invention, The second speech recognition model is trained using the training data generated by the training data generation device. A speech recognition model generation device is provided.

[0011] According to one aspect of the present invention, One or more computers generating text information by inputting speech data into a trained first speech recognition model; generating training data including the voice data and the text information; The first speech recognition model is a model that has been trained using synthetic speech generated using input information related to a predetermined item and standard text information prepared in advance. A training data generation method is provided.

[0012] According to one aspect of the present invention, One or more computers inputting speech data into each of a trained first speech recognition model and a trained second speech recognition model, thereby generating output results of the first speech recognition model and the second speech recognition model; determining whether there is a difference between the output result of the first speech recognition model and the output result of the second speech recognition model; generating training data including the speech data when it is determined that there is a difference between the output result of the first speech recognition model and the output result of the second speech recognition model; A training data generation method is provided.

[0013] According to one aspect of the present invention, One or more computers use the training data generated by the training data generation method to train the second speech recognition model. A method for generating a speech recognition model is provided.

[0014] According to one aspect of the present invention, A computer-readable recording medium having a program recorded thereon, the program causing a computer to function as a training data generation device, the learning data generation device, a speech recognition means for generating text information by inputting speech data into a trained first speech recognition model; generating means for generating training data including the voice data and the text information; The first speech recognition model is a model that has been trained using synthetic speech generated using input information related to a predetermined item and standard text information prepared in advance. A recording medium is provided.

[0015] According to one aspect of the present invention, A computer-readable recording medium having a program recorded thereon, the program causing a computer to function as a training data generation device, the learning data generation device, a speech recognition means for inputting speech data into a trained first speech recognition model and a trained second speech recognition model, and generating output results of the first speech recognition model and the second speech recognition model, respectively; a determination means for determining whether there is a difference between the output result of the first speech recognition model and the output result of the second speech recognition model; a generation means for generating training data including the speech data when the determination means determines that there is a difference between the output result of the first speech recognition model and the output result of the second speech recognition model. A recording medium is provided.

[0016] According to one aspect of the present invention, A computer-readable recording medium having a program recorded thereon, the program causing a computer to function as a speech recognition model generation device, The speech recognition model generation device is realized by a program recorded on the recording medium, and performs training on the second speech recognition model using the training data generated by the training data generation device. A recording medium is provided. [Effects of the Invention]

[0017] According to one aspect of the present invention, a training data generation device, a speech recognition model generation device, a training data generation method, a speech recognition model generation method, and a recording medium are provided that improve the recognition accuracy of a speech recognition model. [Brief explanation of the drawings]

[0018] [Figure 1] 1 is a diagram illustrating an overview of a training data generation device according to a first embodiment. [Figure 2] FIG. 1 is a diagram illustrating an overview of a first speech recognition model. [Figure 3] FIG. 2 is a diagram illustrating an example of an outline of a method for generating a first speech recognition model. [Figure 4] FIG. 2 is a diagram illustrating a functional configuration of the training data generation device according to the first embodiment. [Figure 5] FIG. 2 is a diagram illustrating an example of an outline of a method in which a first model generation unit generates a first speech recognition model. [Figure 6] 1 is a diagram illustrating an example of a functional configuration of a speech recognition model generation device according to a first embodiment. [Figure 7] FIG. 10 is a diagram illustrating an example of a computer for realizing a training data generation device. [Figure 8] FIG. 1 is a diagram illustrating an overview of a training data generation method according to a first embodiment. [Figure 9] 4 is a flowchart illustrating the flow of a training data generation method according to the first embodiment. [Figure 10] FIG. 10 is a diagram illustrating an overview of a training data generation device according to a second embodiment. [Figure 11] FIG. 10 is a diagram illustrating a functional configuration of a training data generation device according to a second embodiment. [Figure 12]FIG. 10 is a diagram illustrating an overview of a training data generation method according to a second embodiment. [Figure 13] 10 is a flowchart illustrating the flow of a training data generation method according to a second embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0019] Hereinafter, embodiments of the present invention will be described with reference to the drawings. In all the drawings, like components are designated by like reference numerals, and their description will be omitted where appropriate.

[0020] (First embodiment) FIG. 1 is a diagram illustrating an overview of a training data generation device 10 according to a first embodiment. FIG. 2 is a diagram illustrating an overview of a first speech recognition model 51. The training data generation device 10 includes a speech recognition unit 140 and a generation unit 160. The speech recognition unit 140 generates text information by inputting speech data into a trained first speech recognition model 51. The generation unit 160 generates training data including speech data and text information. The first speech recognition model 51 is a model trained using synthetic speech generated using input information related to a predetermined item and standard text information prepared in advance.

[0021] According to this training data generation device 10, it is possible to improve the recognition accuracy of the voice recognition model.

[0022] A detailed example of the training data generation device 10 according to this embodiment will be described below.

[0023] In this embodiment, the voice data is data obtained by recording a person's speech. In other words, the voice data is not so-called synthetic sound data, which is artificially generated by a machine or the like. Furthermore, the voice data is data indicating a voice waveform or data indicating features of a voice waveform. The voice data is, for example, data obtained by recording a call voice in a voice call or a video call. As a specific example, the voice data is data obtained by recording a call voice for requesting the dispatch of an emergency vehicle (for example, police, fire engine, or ambulance). As another example, the voice data may be data obtained by recording the call voice of various call centers. One voice data may be generated for a series of calls, or multiple voice data may be generated by dividing a series of calls into multiple parts. However, the voice data is not limited to data obtained by recording a call.

[0024] FIG. 3 is a diagram illustrating an example of an outline of a method for generating the first speech recognition model 51. Both the first speech recognition model 51 and the second speech recognition model 52 are speech recognition models obtained by machine learning. As shown in FIG. 2, the first speech recognition model 51 is a trained model that can convert speech data into text information indicating the content corresponding to the speech data. In other words, the input data of the first speech recognition model 51 includes speech data, and the output data of the first speech recognition model 51 includes text information. The first speech recognition model 51 is a model generated by training the second speech recognition model 52 using synthetic speech.

[0025] Like the first speech recognition model 51, the second speech recognition model 52 is a model capable of converting speech data into text information indicating the content corresponding to the speech data. The second speech recognition model 52 is a model trained using a plurality of training data including speech data and text information indicating the content corresponding to the speech data. The second speech recognition model 52 is preferably a model that has not been trained using synthetic speech. However, the second speech recognition model 52 may be a model that has been trained using synthetic speech as part of its training.

[0026] That is, the first speech recognition model 51 may be a model trained using both one or more pieces of training data including speech data other than synthetic speech and one or more pieces of training data including synthetic speech as a whole of the training history. The first speech recognition model 51 may be a model trained using at least one piece of training data including synthetic speech. Of the multiple pieces of training data used to train the first speech recognition model 51, the number of pieces of training data including synthetic speech may be only one. Synthetic speech will be described in detail later.

[0027] By training using synthetic speech, the first speech recognition model 51 is expected to have higher speech recognition accuracy than the second speech recognition model 52. The speech recognition unit 140 of the training data generation device 10 inputs speech data to the first speech recognition model 51, causing it to output text information. The generation unit 160 then generates training data, which associates the speech data input to the first speech recognition model 51 with the text information output from the first speech recognition model 51. As described above, the speech data included in this training data is not synthetic speech but recording data of actual speech by a person. While synthetic speech differs from actual speech in terms of phrase features, prosodic features, frequency characteristics, and the like, the training data generation device 10 according to this embodiment generates training data including recording data of actual speech by a person. Therefore, the training data generated by the training data generation device 10 according to this embodiment enables training that reflects these features, thereby achieving a speech recognition model with higher recognition accuracy.

[0028] FIG. 4 is a diagram illustrating the functional configuration of a training data generation device 10 according to this embodiment. In the example shown in this figure, the training data generation device 10 further includes an acquisition unit 110, a first model generation unit 120, a template text storage unit 130, a model storage unit 150, and a training data storage unit 170. In addition, in the example shown in this figure, the first model generation unit 120 includes a synthetic speech text generation unit 121, a synthetic speech generation unit 122, and a first training unit 123. Note that one or more of the template text storage unit 130, the model storage unit 150, and the training data storage unit 170 may be storage devices provided external to the training data generation device 10. Each functional component of the training data generation device 10 is described in detail below.

[0029] The acquisition unit 110 acquires input information and voice data that are associated with each other. The input information corresponds to the content of the utterance in the voice data associated with the input information. For example, when the voice data is obtained by recording a phone call, the input information is generated as follows. For example, the call recipient (inputter) inputs the content of the call into a terminal while talking on the phone. For example, the terminal presents the inputter with multiple items to be input, and the input work is performed by filling in the input field for each item. Then, input information indicating the inputted content of the call is generated. As another example, after the call ends, the worker who inputs the information inputs the content of the call into the terminal, thereby generating input information.

[0030] For example, an identification ID is assigned to the voice data of each call, and the input information is associated with the voice data by associating the call's identification ID with the input information. The input information may also include the voice data's identification ID. As another example, the input information may include the identification ID of the receiving terminal of the call and the date and time of the call, thereby associating the input information with the voice data. In this case, the voice data is assigned information indicating the receiving terminal's identification ID and the recording date and time (i.e., the date and time of the call).

[0031] The items included in the input information correspond to the items to be entered into the terminal described above. The multiple items included in the input information may include, for example, one or more of the "name of the callee," "address," "telephone number," and "items related to the subject of the call." For example, if the call is a request for the dispatch of an emergency vehicle (e.g., police, fire engine, or ambulance), the "items related to the subject of the call" may include, for example, the "type of incident" (e.g., whether it is an incident, accident, fire, or sudden illness) and the "location of the request" (the location of the accident, etc.). Here, the "type of incident" may be indicated by, for example, a predetermined number or symbol for each of the incident, accident, fire, and sudden illness. For example, if the call is a request for the dispatch of an ambulance, the "items related to the subject of the call" may further include one or more of the "body part," "condition of injury, etc.," and "symptoms." The items included in the input information may vary depending on the "type of incident." For example, if the "type of incident" is an incident or an accident, the "items related to the subject of the call" may include the "scene situation" (e.g., a car rollover). For example, if the call is for placing a telephone shopping order, the "items related to the business" may include "information indicating the purchased product," "quantity purchased," "delivery address," etc. The content entered for each item may be text.

[0032] Such input work is usually performed within the scope of call reception work and does not need to be performed specially to have the training data generation device 10 generate training data. Therefore, by using the training data generation device 10, training data can be generated without requiring special effort, and the accuracy of the speech recognition model can be improved. However, the input information is not limited to the above example, and may be any information that can generate text for synthesis by applying the content of each item to standard text information, as described below. The input information does not necessarily have to be related to the speech data acquired by the acquisition unit 110. In this case, multiple pieces of input information may be used to generate multiple texts for synthesis and multiple synthetic sounds. The first speech recognition model 51 may be a model trained using multiple synthetic sounds.

[0033] The method by which the acquisition unit 110 acquires the input information and voice data is not particularly limited, but for example, the acquisition unit 110 can acquire the input information and voice data by reading them from a storage device in which the input information and voice data are stored. As another example, the acquisition unit 110 may acquire the input information directly from a terminal into which the contents of the call are input.

[0034] The acquisition unit 110 can acquire multiple pieces of voice data. The acquisition unit 110 may acquire each piece of voice data each time it is generated, or may acquire multiple pieces of voice data at once. It is preferable that the first model generation unit 120 generates a first voice recognition model 51 for each piece of voice data acquired by the acquisition unit 110.

[0035] FIG. 5 is a diagram illustrating an example of an outline of a method by which the first model generation unit 120 generates a first speech recognition model 51. The synthetic speech text generation unit 121 of the first model generation unit 120 acquires input information corresponding to, for example, certain speech data. The synthetic speech text generation unit 121 also acquires standard text information from the standard text storage unit 130. The standard text information is prepared in advance and stored in the standard text storage unit 130. The standard text information is information indicating standard text such as, for example, "This is xx. An accident occurred at yy." The synthetic speech text generation unit 121 generates text for synthetic speech by applying the input information to the standard text information. Specifically, for example, the "xx" part of "This is xx. An accident occurred at yy." is replaced with the name indicated in the input information, and the "yy" part is replaced with the requested location indicated in the input information. In this way, the synthetic speech text generation unit 121 can generate text for synthetic speech using the standard text information and the input information.

[0036] Here, the synthetic speech text generation unit 121 may select the fixed text information to be used from multiple fixed text information stored in the fixed text storage unit 130. For example, the fixed text storage unit 130 stores fixed text information for each type of incident. Each piece of associated text information is associated with one of the types of incident. The synthetic speech text generation unit 121 then selects the fixed text information corresponding to the type of incident indicated in the input information as the fixed text information to be used. For example, if the type of incident is an accident, the fixed text information "This is xx. An accident occurred at yy." is selected; if the type of incident is a fire, the fixed text information "This is xx. A fire occurred at yy." is selected; and if the type of incident is a sudden illness, the fixed text information "This is xx. There is a person who is suddenly ill at yy." is selected. The synthetic speech text generation unit 121 then generates the text for synthetic speech using the selected fixed text information in the same manner as described above.

[0037] The synthetic speech generation unit 122 acquires the text for synthetic speech generated by the text for synthetic speech generation unit 121 and converts it into synthetic speech. The synthetic speech corresponds to the content of the text for synthetic speech and corresponds to the speech of the text for synthetic speech being read aloud. Existing technology can be used as a method for converting the text for synthetic speech into synthetic speech. The text for synthetic speech generation unit 121 can convert the text for synthetic speech into synthetic speech, for example, using a trained model that takes text as input and outputs synthetic speech.

[0038] The first learning unit 123 generates training data for generating the first speech recognition model 51 by associating the synthetic speech text generated by the synthetic speech text generation unit 121 with the synthetic speech generated by the synthetic speech generation unit 122. The first learning unit 123 then uses the generated training data to train the second speech recognition model 52, thereby generating the first speech recognition model 51. The second speech recognition model 52 is stored in the model storage unit 150, and the first learning unit 123 can read the second speech recognition model 52 from the model storage unit 150 and use it to generate the first speech recognition model 51. The first speech recognition model 51 trained using training data including the synthetic speech text and synthetic speech is output to the speech recognition unit 140. In this way, the first model generation unit 120 can generate the first speech recognition model 51 with improved recognition accuracy by using input information without requiring much effort.

[0039] 4, the speech recognition unit 140 acquires the first speech recognition model 51 generated by the first model generation unit 120. Then, the speech data acquired by the acquisition unit 110 is input to the acquired first speech recognition model 51. Then, text information corresponding to the speech data is generated as an output of the first speech recognition model 51.

[0040] The generation unit 160 generates training data by associating the voice data acquired by the acquisition unit 110 with the text information generated by the speech recognition unit 140. The generation unit 160 stores the generated training data in the training data storage unit 170, for example. However, instead, the generation unit 160 may output the generated training data to an external device.

[0041] The training data generation device 10 does not have to include the acquisition unit 110 and the first model generation unit 120. In this case, the first speech recognition model 51, which has been trained in advance using synthetic speech, is stored in a storage device accessible by the speech recognition unit 140, and the speech recognition unit 140 can read and use the first speech recognition model 51.

[0042] The effect of generating the first speech recognition model 51 for each piece of speech data, as performed by the first model generation unit 120, is described below. The first speech recognition model 51 generated by the first model generation unit 120 as described above is considered to have particularly high recognition accuracy for the speech data acquired by the acquisition unit 110. In other words, the first speech recognition model 51 trained with synthetic speech based on input information associated with certain speech data k can be considered a model particularly suited to recognizing the speech data k. Such a first speech recognition model 51 is likely to be able to correctly recognize the speech data k. In other words, text information obtained by inputting speech data k into such a first speech recognition model 51 is likely to correctly represent the spoken content of the speech data k. Therefore, the text information can be suitably used as correct answer data for training data. Note that the first speech recognition model 51 may be deleted after generating correct answer data for the speech data k. A new first speech recognition model 51 may be generated for another piece of speech data k+1.

[0043] The training data generated by the training data generation device 10 is preferably used for training the second speech recognition model 52, but may also be used for training a speech recognition model other than the second speech recognition model 52.

[0044] FIG. 6 is a diagram illustrating the functional configuration of a speech recognition model generation device 20 according to this embodiment. The speech recognition model generation device 20 uses training data generated by the training data generation device 10 to train the second speech recognition model 52. As described above, the first speech recognition model 51 is expected to have higher recognition accuracy than the second speech recognition model 52. Therefore, by using training data in which text information generated by the first speech recognition model 51 is used as correct answer data, the speech recognition accuracy of the second speech recognition model 52 can be improved. Furthermore, since the second speech recognition model 52 can be trained using speech data that is not synthetic speech, it is possible to further improve the recognition accuracy of actual speech.

[0045] In the example shown in the figure, the speech recognition model generation device 20 includes a second learning unit 220. The second learning unit 220 acquires the second speech recognition model 52 from the model storage unit 150 and acquires the training data generated by the training data generation device 10 from the training data storage unit 170. The second learning unit 220 then performs training on the second speech recognition model 52 using the acquired training data, thereby generating a second speech recognition model 52 with improved recognition accuracy. The speech recognition model generation device 20 may acquire the training data every time it is generated by the training data generation device 10, or may acquire the training data collectively after multiple pieces of training data are generated by the training data generation device 10 and stored in the training data storage unit 170.

[0046] The speech recognition model generation device 20 may update the second speech recognition model 52 stored in the model storage unit 150 with the trained second speech recognition model 52. The updated second speech recognition model 52 can be used again in the training data generation device 10 to generate the first speech recognition model 51.

[0047] The speech recognition model generation device 20 may be integrated with the training data generation device 10 or may be a device separate from the training data generation device 10.

[0048] According to the speech recognition model generation device 20 of this embodiment, a speech recognition model generation method is executed in which one or more computers train the second speech recognition model 52 using training data generated by the training data generation device 10.

[0049] The hardware configuration of the training data generation device 10 is described below. Each functional component of the training data generation device 10 may be realized by hardware that realizes the functional component (e.g., a hardwired electronic circuit, etc.), or by a combination of hardware and software (e.g., a combination of an electronic circuit and a program that controls it). Below, a case where each functional component of the training data generation device 10 is realized by a combination of hardware and software will be further described.

[0050] FIG. 7 is a diagram illustrating a computer 1000 for implementing the training data generation device 10. The computer 1000 is any computer. For example, the computer 1000 is a system on chip (SoC), a personal computer (PC), a server machine, a tablet terminal, or a smartphone. The computer 1000 may be a dedicated computer designed to implement the training data generation device 10, or may be a general-purpose computer. The training data generation device 10 may be implemented by a single computer 1000 or a combination of multiple computers 1000.

[0051] The computer 1000 includes a bus 1020, a processor 1040, a memory 1060, a storage device 1080, an input / output interface 1100, and a network interface 1120. The bus 1020 is a data transmission path through which the processor 1040, the memory 1060, the storage device 1080, the input / output interface 1100, and the network interface 1120 transmit and receive data to and from each other. However, the method of interconnecting the processor 1040 and other components is not limited to bus connection. The processor 1040 may be any of various processors, such as a central processing unit (CPU), a graphics processing unit (GPU), or a field-programmable gate array (FPGA). The memory 1060 is a main storage device implemented using a random access memory (RAM) or the like. The storage device 1080 is an auxiliary storage device implemented using a hard disk, a solid state drive (SSD), a memory card, a read-only memory (ROM), or the like.

[0052] The input / output interface 1100 is an interface for connecting the computer 1000 to an input / output device. For example, an input device such as a keyboard and an output device such as a display are connected to the input / output interface 1100. The input / output interface 1100 may be connected to the input device or output device by wireless connection or by wired connection.

[0053] The network interface 1120 is an interface for connecting the computer 1000 to a network. This communication network is, for example, a LAN (Local Area Network) or a WAN (Wide Area Network). The network interface 1120 may be connected to the network wirelessly or by wire.

[0054] The storage device 1080 stores program modules that realize each functional component of the training data generation device 10. The processor 1040 reads each of these program modules into the memory 1060 and executes them to realize the function corresponding to each program module. Furthermore, when the fixed text storage unit 130, the model storage unit 150, and the training data storage unit 170 are each provided inside the training data generation device 10, the fixed text storage unit 130, the model storage unit 150, and the training data storage unit 170 are realized by the storage device 1080.

[0055] The hardware configuration of a computer that realizes the speech recognition model generation device 20 according to this embodiment is shown in, for example, Fig. 7 , similar to the training data generation device 10. However, a storage device 1080 of a computer 1000 that realizes the speech recognition model generation device 20 stores program modules that realize the functions of the speech recognition model generation device 20.

[0056] FIG. 8 is a diagram showing an overview of the training data generation method according to this embodiment. The training data generation method according to this embodiment is executed by one or more computers. The training data generation method according to this embodiment includes a speech recognition step S10 and a generation step S11. In the speech recognition step S10, text information is generated by inputting speech data into a trained first speech recognition model 51. In the generation step S11, training data including speech data and text information is generated. The first speech recognition model 51 is a model trained using synthetic speech generated using input information related to a predetermined item and standard text information prepared in advance.

[0057] FIG. 9 is a flowchart illustrating the flow of the training data generation method according to this embodiment. In the training data generation method according to this embodiment, the acquisition unit 110 acquires input information and speech data that are associated with each other (S100). Next, the synthetic speech text generation unit 121 generates text for synthetic speech using the input information and standard text information, and the synthetic speech generation unit 122 generates synthetic speech using the text for synthetic speech (S110). Next, the first learning unit 123 trains the second speech recognition model 52 using the text for synthetic speech and the synthetic speech, thereby generating the first speech recognition model 51 (S120). Next, the speech recognition unit 140 inputs speech data into the first speech recognition model 51 to generate text information (S130). Then, the generation unit 160 generates training data that includes speech data and text information in a state where the speech data and the text information are associated with each other (S140). The processes from S100 to S140 are performed, for example, for each piece of speech data.

[0058] As described above, according to this embodiment, the speech recognition unit 140 generates text information by inputting speech data into the trained first speech recognition model 51. The generation unit 160 generates training data including speech data and text information. The first speech recognition model 51 is a model trained using synthetic speech generated using input information and pre-prepared standard text information. Therefore, training data including speech data can be easily generated using the first speech recognition model 51, the accuracy of which has been improved using synthetic speech. Consequently, a speech recognition model with high recognition accuracy is realized.

[0059] (Second embodiment) FIG. 10 is a diagram illustrating an overview of a training data generation device 10 according to the second embodiment. The training data generation device 10 according to this embodiment includes a speech recognition unit 140, a determination unit 180, and a generation unit 160. The speech recognition unit 140 inputs speech data to a trained first speech recognition model and a trained second speech recognition model, respectively, to generate output results for the first speech recognition model and the second speech recognition model. The determination unit 180 determines whether there is a difference between the output result of the first speech recognition model and the output result of the second speech recognition model. If the determination unit 180 determines that there is a difference between the output result of the first speech recognition model and the output result of the second speech recognition model, the generation unit 160 generates training data including speech data.

[0060] According to the training data generation device 10 of this embodiment, it is possible to improve the recognition accuracy of the voice recognition model.

[0061] A detailed example of the training data generation device 10 according to this embodiment will be described below. However, the training data generation device 10 according to this embodiment is not limited to the following example.

[0062] 11 is a diagram illustrating the functional configuration of a training data generation device 10 according to this embodiment. The training data generation device 10 according to this embodiment is the same as the training data generation device 10 according to the first embodiment, except for the points described below.

[0063] In this embodiment, when the first model generation unit 120 generates the first speech recognition model 51, the speech recognition unit 140 inputs the speech data acquired by the acquisition unit 110 to the generated first speech recognition model 51, as in the first embodiment. The speech recognition unit 140 also inputs the same speech data as that input to the first speech recognition model 51 to the second speech recognition model 52 read from the model storage unit 150. Then, text information, which is the output result, is obtained from each of the first speech recognition model 51 and the second speech recognition model 52.

[0064] However, the training data generation device 10 does not necessarily have to include the acquisition unit 110 and the first model generation unit 120. In this case, the speech recognition unit 140 reads and acquires the first speech recognition model 51 and the second speech recognition model 52 stored in advance in a storage device accessible from the speech recognition unit 140. However, the first speech recognition model 51 is a model with higher speech recognition accuracy than the second speech recognition model 52.

[0065] The determination unit 180 compares the text information that is the output result of the first speech recognition model 51 and the text information that is the output result of the second speech recognition model 52, both of which are generated by the speech recognition unit 140. For example, if the output result of the first speech recognition model 51 and the output result of the second speech recognition model 52 match, the determination unit 180 outputs determination result information to the generation unit 160 indicating that there is no need to generate training data. If the output result of the first speech recognition model 51 and the output result of the second speech recognition model 52 do not match, the determination unit 180 outputs determination result information to the generation unit 160 indicating that training data should be generated.

[0066] The generation unit 160 acquires determination result information from the determination unit 180. If the generation unit 160 acquires determination result information indicating that it is not necessary to generate training data, the generation unit 160 does not generate training data, and the training data generation device 10 ends processing of the speech data. If the generation unit 160 acquires determination result information indicating that training data should be generated, the generation unit 160 generates training data. Specifically, the generation unit 160 generates training data by associating the speech data acquired by the acquisition unit 110 with text information that is the output result of the first speech recognition model 51 generated by the speech recognition unit 140. The generation unit 160 stores the generated training data in the training data storage unit 170, for example. However, the generation unit 160 may instead output the generated training data to an external device.

[0067] According to the training data generation device 10 of this embodiment, training data including speech data is generated only when it is determined that there is a difference between the output result of the first speech recognition model 51 and the output result of the second speech recognition model 52. When the same speech data is input, if the output result of the first speech recognition model 51 and the output result of the second speech recognition model 52 are the same, it is highly likely that both speech recognition models have produced correct output. In this case, further training using that speech data in the second speech recognition model 52 is not very effective.

[0068] On the other hand, as described above in the first embodiment, the first speech recognition model 51 is expected to have higher recognition accuracy than the second speech recognition model 52. Therefore, when the same speech data is input, if the output result of the first speech recognition model 51 and the output result of the second speech recognition model 52 differ, the output result of the first speech recognition model 51 is likely to be more accurate than the output result of the second speech recognition model 52. Therefore, it is preferable to generate training data using the output result of the first speech recognition model 51. By training using the generated training data, the recognition accuracy of the second speech recognition model 52 can be improved.

[0069] In this way, the determination unit 180 determines whether there is a difference between the output result of the first speech recognition model and the output result of the second speech recognition model, thereby generating learning data that enables efficient learning.

[0070] The speech recognition model generation device 20 according to this embodiment is the same as the speech recognition model generation device 20 according to the first embodiment, except that training is performed on a second speech recognition model 52 using training data generated by the training data generation device 10 according to the second embodiment.

[0071] The hardware configuration of a computer that realizes the training data generation device 10 according to this embodiment is shown, for example, in Fig. 7, similar to the training data generation device 10 according to the first embodiment. The hardware configuration of a computer that realizes the speech recognition model generation device 20 according to this embodiment is shown, for example, in Fig. 7, similar to the speech recognition model generation device 20 according to the first embodiment. However, a storage device 1080 of a computer 1000 that realizes the training data generation device 10 according to this embodiment further stores a program module that realizes the determination unit 180 of the training data generation device 10 according to this embodiment.

[0072] FIG. 12 is a diagram illustrating an overview of a training data generation method according to this embodiment. The training data generation method according to this embodiment is executed by one or more computers. The training data generation method according to this embodiment includes a speech recognition step S20, a determination step S21, and a generation step S22. In the speech recognition step S20, speech data is input to a trained first speech recognition model and a trained second speech recognition model, respectively, to generate output results for the first speech recognition model and the second speech recognition model. In the determination step S21, it is determined whether there is a difference between the output result of the first speech recognition model and the output result of the second speech recognition model. In the generation step S22, if it is determined that there is a difference between the output result of the first speech recognition model and the output result of the second speech recognition model, training data including speech data is generated.

[0073] FIG. 13 is a flowchart illustrating the flow of the training data generation method according to this embodiment. The processes from S200 to S220 are the same as the processes from S100 to S120 in the first embodiment. In the training data generation method according to this embodiment, after S220, the speech recognition unit 140 inputs speech data to each of the first speech recognition model 51 and the second speech recognition model 52, thereby generating text information as the output result of each speech recognition model (S230). Then, the determination unit 180 determines whether there is a difference between the output result of the first speech recognition model 51 and the output result of the second speech recognition model 52 (S240). If it is determined in S240 that there is a difference (Yes in S240), the generation unit 160 generates training data including the speech data (S250). Then, processing related to the speech data ends. On the other hand, if it is determined in S240 that there is no difference (No in S240), no training data is generated and processing related to the speech data ends. The processes from S200 to S250 are performed for each piece of audio data, for example.

[0074] A modified example of the method in which the determination unit 180 determines whether or not there is a difference between the output result of the first speech recognition model 51 and the output result of the second speech recognition model 52 will be described below.

[0075] Instead of determining whether the output result of the first speech recognition model 51 and the output result of the second speech recognition model 52 match, the determination unit 180 may determine whether there is a difference between the output result of the first speech recognition model 51 and the output result of the second speech recognition model 52 based on whether the target word is included in each of the output results of the first speech recognition model 51 and the second speech recognition model 52. The target word is, for example, one or more of the contents of multiple items included in the input information. Preferably, the target word is all of the contents of multiple items included in the input information. The determination unit 180 can identify the target word using predetermined information indicating the items to be the target word and the input information acquired by the acquisition unit 110.

[0076] In this modification, when there is a difference in the recognition results of the target words, the determination unit 180 determines that there is a difference between the output result of the first speech recognition model 51 and the output result of the second speech recognition model 52. Specifically, the determination unit 180 detects one or more target words included in the text information that is the output result of the first speech recognition model 51. The determination unit 180 also detects one or more target words included in the text information that is the output result of the second speech recognition model 52. Then, when the one or more target words detected in the output result of the first speech recognition model 51 and the one or more target words detected in the output result of the second speech recognition model 52 all match, the determination unit 180 determines that there is no difference between the output result of the first speech recognition model 51 and the output result of the second speech recognition model 52. Then, if one or more target words detected in the output result of the first speech recognition model 51 do not match one or more target words detected in the output result of the second speech recognition model 52, it is determined that there is a difference between the output result of the first speech recognition model 51 and the output result of the second speech recognition model 52.

[0077] For example, suppose the target words are word A, word B, and word C. Here, if word A, word B, and word C are detected in the output result of first speech recognition model 51, and only word A and word B are detected in the output result of second speech recognition model 52, determination unit 180 determines that there is a difference between these output results.

[0078] As another example, when there are multiple target words, the determination unit 180 may determine whether there is a difference between the two output results by comparing the number of detected target words. That is, when the number of target words detected in the output result of the first speech recognition model 51 matches the number of target words detected in the output result of the second speech recognition model 52, the determination unit 180 determines that there is no difference between the output result of the first speech recognition model 51 and the output result of the second speech recognition model 52. On the other hand, when the number of target words detected in the output result of the first speech recognition model 51 does not match the number of target words detected in the output result of the second speech recognition model 52, the determination unit 180 determines that there is a difference between the output result of the first speech recognition model 51 and the output result of the second speech recognition model 52.

[0079] If at least one of the following (1) to (3) is true, the determining unit 180 may output to the generating unit 160 determination result information indicating that it is not necessary to generate learning data. (1) At least one target word is detected only in the output result of the second speech recognition model 52. (2) The number of target words detected in the output result of the second speech recognition model 52 is greater than the number of target words detected in the output result of the first speech recognition model 51. (3) Neither the target word was detected in the output result of the first speech recognition model 51 nor in the output result of the second speech recognition model 52.

[0080] Next, the operation and effect of this embodiment will be described. In this embodiment, the same operation and effect as in the first embodiment can be obtained. In addition, the determination unit 180 determines whether there is a difference between the output result of the first speech recognition model and the output result of the second speech recognition model, thereby generating learning data that enables efficient learning.

[0081] Although the embodiments of the present invention have been described above with reference to the drawings, these are merely examples of the present invention, and various other configurations can also be adopted.

[0082] In addition, in the flowcharts used in the above description, multiple steps (processes) are described in order, but the order of execution of the steps performed in each embodiment is not limited to the order described. In each embodiment, the order of the steps shown in the drawings can be changed to the extent that the content is not affected. Furthermore, the above-mentioned embodiments can be combined to the extent that the content is not contradictory.

[0083] Some or all of the above-described embodiments can be described as, but are not limited to, the following supplementary notes. 1-1. A speech recognition means for generating text information by inputting speech data into a trained first speech recognition model; generating means for generating training data including the voice data and the text information; The first speech recognition model is a model that has been trained using synthetic speech generated using input information related to a predetermined item and standard text information prepared in advance. Training data generation device. 1-2. In the training data generation device described in 1-1, The apparatus further includes an acquisition unit for acquiring the input information and the voice data associated with each other. Training data generation device. 1-3. In the training data generation device described in 1-2, the acquiring means acquires a plurality of pieces of the voice data, The apparatus further includes a first model generation means for generating the first speech recognition model for each of the speech data. Training data generation device. 1-4. In the training data generation device according to 1-3, The first model generation means generates the first speech recognition model by training a second speech recognition model using the synthetic speech. Training data generation device. 2-1. A speech recognition means for inputting speech data into a trained first speech recognition model and a trained second speech recognition model, respectively, to generate output results of the first speech recognition model and the second speech recognition model; a determination means for determining whether there is a difference between the output result of the first speech recognition model and the output result of the second speech recognition model; a generation means for generating training data including the speech data when the determination means determines that there is a difference between the output result of the first speech recognition model and the output result of the second speech recognition model. Training data generation device. 2-2. In the training data generation device described in 2-1, The system further includes a first model generation means for generating the first speech recognition model by training the second speech recognition model using the synthesized speech. Training data generation device. 2-3. In the training data generation device described in 2-2, A training data generation device in which the synthetic sound is a sound generated using input information related to a predetermined item and standard text information prepared in advance. 2-4. In the training data generation device according to 2-2 or 2-3, The apparatus further includes an acquisition unit for acquiring the input information and the voice data associated with each other. Training data generation device. 2-5. In the training data generation device described in 2-4, the acquiring means acquires a plurality of pieces of audio data; The first model generation means generates the first speech recognition model for each of the speech data. Training data generation device. 2-6. The training data generation device according to any one of 2-2 to 2-5, The generation means generates training data including the output result of the first speech recognition model and the speech data when the determination means determines that there is a difference between the output result of the first speech recognition model and the output result of the second speech recognition model. Training data generation device. 3-1. Training the second speech recognition model using the training data generated by the training data generation device according to any one of 1-4 and 2-1 to 2-6. Speech recognition model generator. 4-1.1 or higher computers, generating text information by inputting speech data into a trained first speech recognition model; generating training data including the voice data and the text information; The first speech recognition model is a model that has been trained using synthetic speech generated using input information related to a predetermined item and standard text information prepared in advance. Training data generation method. 4-2. In the training data generation method described in 4-1, The one or more computers further acquire the input information and the voice data associated with each other. Training data generation method. 4-3. In the training data generation method described in 4-2, The one or more computers acquiring a plurality of said voice data; Furthermore, the first speech recognition model is generated for each of the speech data. Training data generation method. 4-4. In the training data generation method described in 4-3, The one or more computers generate the first speech recognition model by training a second speech recognition model using the synthesized speech. Training data generation method. 5-1. One or more computers inputting speech data into each of a trained first speech recognition model and a trained second speech recognition model, thereby generating output results of the first speech recognition model and the second speech recognition model; determining whether there is a difference between the output result of the first speech recognition model and the output result of the second speech recognition model; generating training data including the speech data when it is determined that there is a difference between the output result of the first speech recognition model and the output result of the second speech recognition model; Training data generation method. 5-2. In the training data generation method described in 5-1, The one or more computers further generate the first speech recognition model by training the second speech recognition model using the synthesized speech. Training data generation method. 5-3. In the training data generation method described in 5-2, A learning data generation method in which the synthetic sound is a sound generated using input information related to a predetermined item and standard text information prepared in advance. 5-4. In the training data generation method described in 5-2. or 5-3., The one or more computers further acquire the input information and the voice data associated with each other. Training data generation method. 5-5. In the training data generation method described in 5-4, The one or more computers acquiring a plurality of said voice data; The first speech recognition model is generated for each of the speech data. Training data generation method. 5-6. The training data generation method according to any one of 5-2 to 5-5, When it is determined that there is a difference between the output result of the first speech recognition model and the output result of the second speech recognition model, the one or more computers generate training data including the output result of the first speech recognition model and the speech data. Training data generation method. 6-1. One or more computers use the training data generated by the training data generation method according to any one of 4-4 and 5-1 to 5-6 to train the second speech recognition model. A method for generating a speech recognition model. 7-1. A program that causes a computer to function as a training data generation device, the learning data generation device, a speech recognition means for generating text information by inputting speech data into a trained first speech recognition model; generating means for generating training data including the voice data and the text information; The first speech recognition model is a model that has been trained using synthetic speech generated using input information related to a predetermined item and standard text information prepared in advance. program. 7-2. In the program described in 7-1., The training data generation device further includes an acquisition unit for acquiring the input information and the voice data associated with each other. program. 7-3. In the program described in 7-2., the acquiring means acquires a plurality of pieces of the voice data, The training data generation device further includes a first model generation means for generating the first speech recognition model for each of the speech data. program. 7-4. In the program described in 7-3., The first model generation means generates the first speech recognition model by training a second speech recognition model using the synthetic speech. program. 8-1. A program that causes a computer to function as a training data generation device, the learning data generation device, a speech recognition means for inputting speech data into a trained first speech recognition model and a trained second speech recognition model, and generating output results of the first speech recognition model and the second speech recognition model, respectively; a determination means for determining whether there is a difference between the output result of the first speech recognition model and the output result of the second speech recognition model; a generation means for generating training data including the speech data when the determination means determines that there is a difference between the output result of the first speech recognition model and the output result of the second speech recognition model. program. 8-2. In the program described in 8-1., The training data generation device further includes a first model generation means for generating the first speech recognition model by training the second speech recognition model using synthetic speech. program. 8-3. In the program described in 8-2., The synthesized sound is generated using input information relating to a predetermined item and pre-prepared standard text information. 8-4. In the program described in 8-2. or 8-3., The training data generation device further includes an acquisition unit for acquiring the input information and the voice data associated with each other. program. 8-5. In the program described in 8-4., the acquiring means acquires a plurality of pieces of audio data; The first model generation means generates the first speech recognition model for each of the speech data. program. 8-6. In the program according to any one of 8-2. to 8-5., The generation means generates training data including the output result of the first speech recognition model and the speech data when the determination means determines that there is a difference between the output result of the first speech recognition model and the output result of the second speech recognition model. program. 9-1. A program that causes a computer to function as a speech recognition model generation device, The speech recognition model generation device performs training on the second speech recognition model using the training data generated by a training data generation device realized by any one of the programs described in 7-4 and 8-1 to 8-6. program. 10-1. A computer-readable recording medium having a program recorded thereon, the program causing a computer to function as a training data generation device, the learning data generation device, a speech recognition means for generating text information by inputting speech data into a trained first speech recognition model; generating means for generating training data including the voice data and the text information; The first speech recognition model is a model that has been trained using synthetic speech generated using input information related to a predetermined item and standard text information prepared in advance. Recording medium. 10-2. In the recording medium described in 10-1, The training data generation device further includes an acquisition unit for acquiring the input information and the voice data associated with each other. Recording medium. 10-3. In the recording medium described in 10-2, the acquiring means acquires a plurality of pieces of the voice data, The training data generation device further includes a first model generation means for generating the first speech recognition model for each of the speech data. Recording medium. 10-4. In the recording medium according to 10-3, The first model generation means generates the first speech recognition model by training a second speech recognition model using the synthetic speech. Recording medium. 11-1. A computer-readable recording medium having a program recorded thereon, the program causing a computer to function as a training data generation device, the learning data generation device, a speech recognition means for inputting speech data into a trained first speech recognition model and a trained second speech recognition model, and generating output results of the first speech recognition model and the second speech recognition model, respectively; a determination means for determining whether there is a difference between the output result of the first speech recognition model and the output result of the second speech recognition model; a generation means for generating training data including the speech data when the determination means determines that there is a difference between the output result of the first speech recognition model and the output result of the second speech recognition model. Recording medium. 11-2. In the recording medium described in 11-1, The training data generation device further includes a first model generation means for generating the first speech recognition model by training the second speech recognition model using synthetic speech. Recording medium. 11-3. In the recording medium described in 11-2, The recording medium is a recording medium in which the synthesized sound is generated using input information relating to a predetermined item and standard text information prepared in advance. 11-4. In the recording medium according to 11-2. or 11-3., The training data generation device further includes an acquisition unit for acquiring the input information and the voice data associated with each other. Recording medium. 11-5. In the recording medium described in 11-4, the acquiring means acquires a plurality of pieces of audio data; The first model generation means generates the first speech recognition model for each of the speech data. Recording medium. 11-6. The recording medium according to any one of 11-2 to 11-5, The generation means generates training data including the output result of the first speech recognition model and the speech data when the determination means determines that there is a difference between the output result of the first speech recognition model and the output result of the second speech recognition model. Recording medium. 12-1. A computer-readable recording medium having a program recorded thereon, the program causing a computer to function as a speech recognition model generation device, The speech recognition model generation device is realized by a program recorded on a recording medium described in any one of 10-4. and 11-1. to 11-6., and performs training on the second speech recognition model using the training data generated by the training data generation device. Recording medium.

[0084] This application claims priority based on Japanese Patent Application No. 2022-107582, filed on July 4, 2022, the disclosure of which is incorporated herein by reference in its entirety. [Explanation of symbols]

[0085] 10. Training data generation device 20. Speech recognition model generation device 51 First Speech Recognition Model 52 Second Speech Recognition Model 110 Acquisition Department 120 First model generation unit 121 Text generation unit for synthesized speech 122 Synthetic sound generation unit 123 First Learning Section 130 Fixed text storage section 140 Voice Recognition Unit 150 Model Memory Unit 160 Generation part 170 Learning data storage unit 180 Judgment section 220 Second Learning Section 1000 calculator 1020 Bus 1040 processor 1060 memory 1080 storage device 1100 Input / Output Interface 1120 Network Interface

Claims

1. a speech recognition means for generating text information by inputting speech data into a trained first speech recognition model; generating means for generating training data including the voice data and the text information; The first speech recognition model is a model that has been trained using synthetic speech generated using input information related to a predetermined item and standard text information prepared in advance. Training data generation device.

2. 2. The training data generation device according to claim 1, The apparatus further includes an acquisition unit for acquiring the input information and the voice data associated with each other. Training data generation device.

3. 3. The training data generation device according to claim 2, the acquiring means acquires a plurality of pieces of the voice data, The voice recognition system further includes a first model generating means for generating the first voice recognition model for each of the voice data. Training data generation device.

4. 4. The training data generation device according to claim 3, The first model generation means generates the first speech recognition model by training a second speech recognition model using the synthetic speech. Training data generation device.

5. The second speech recognition model is trained using the training data generated by the training data generation device according to claim 4. Speech recognition model generator.

6. One or more computers generating text information by inputting speech data into a trained first speech recognition model; generating training data including the voice data and the text information; The first speech recognition model is a model that has been trained using synthetic speech generated using input information related to a predetermined item and standard text information prepared in advance. Training data generation method.

7. Making a computer function as a training data generation device, the learning data generation device, a speech recognition means for generating text information by inputting speech data into a trained first speech recognition model; generating means for generating training data including the voice data and the text information; The first speech recognition model is a model that has been trained using synthetic speech generated using input information related to a predetermined item and standard text information prepared in advance. program.

Citation Information

Patent Citations

  • Voice recognition device

    JP2003029776A

  • Device and program for speech recognition, and method and device for language model generation

    JP2005208483A

  • Speech chain apparatus, computer program, and DNN speech recognition / synthesis cross-learning method

    JP2019120841A

  • Data generation device, data generation method, and program

    JP2021131514A

  • Synthesized Data Augmentation Using Voice Conversion and Speech Recognition Models

    US20220068257A1