Method, apparatus and system for constructing data set for automatic speech recognition
By generating and verifying audio datasets and assigning pseudo-labels, the automatic speech recognition module is trained to deal with uncommon languages, solving the problem of insufficient data and improving the accuracy and consistency of recognition.
Patent Information
- Application Number
- CN202380083622.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-12-08
- Filing Date
- 2023-11-10
- Publication Date
- 2025-08-15
AI Technical Summary
Existing automatic speech recognition technologies are difficult to effectively handle relatively less common languages and dialects, and publicly available data is not enough to build robust data sets, resulting in inconsistent recognition results and limited accuracy.
By generating the input dataset, obtaining the associated audio dataset, verifying and assigning pseudo-labels using the audio verification module, repeating this process until a predetermined number is reached, and then training the automatic speech recognition module based on the pseudo-labels, in particular introducing relatively infrequently used vocabulary and languages.
The automatic speech recognition system's recognition accuracy and consistency of uncommon languages and dialects is improved, and the robustness of the data set is enhanced.
Smart Images

Figure CN120500720A_ABST
Abstract
Description
Technical Field
[0001] Various aspects of the present disclosure relate to methods, apparatus, and systems for constructing a dataset for automatic speech recognition. Background Art
[0002] Automatic speech recognition (ASR) technology is used to convert spoken words (typically captured as audio signals or data) into written text. ASR technology can be used to transcribe audio files / datasets and perform post-call reviews, among other things.
[0003] While adoption of ASR technology is increasing, current ASR technology may not be sufficient for less common languages, such as Southeast Asian languages and dialects. Additionally, while there is publicly available data for some less common languages, this publicly available data may not be sufficient to construct a sizable dataset for robust ASR. ASR-based recognition results may be further affected by how the audio data is collected, i.e., via scripted or spontaneous readings, which may make ASR results less consistent or different, and any uncommon vocabulary used may affect ASR accuracy.
[0004] Therefore, there is a need for an improved automatic speech recognition system, method and / or apparatus. Summary of the Invention
[0005] The technical solution is directed to providing a method, apparatus, and / or system for constructing one or more datasets for automatic speech recognition. In some aspects, the datasets constitute at least a portion of a training dataset used to train an artificial intelligence-based automatic speech recognition (ASR) module to recognize relatively uncommon languages. In some aspects, the technical solution may include means for generating an input dataset for speaker verification and obtaining an audio dataset from one or more selected users. The obtained audio datasets may be assigned pseudo-labels to facilitate semi-supervised training of the ASR module.
[0006] In one aspect of the present disclosure, a method for constructing a dataset for automatic speech recognition is provided, the method comprising: generating an input dataset; obtaining an audio dataset associated with, corresponding to, or based on the input dataset; verifying the audio dataset using an audio verification module; assigning at least one pseudo-label to the verified audio dataset; storing at least one pseudo-labeled audio dataset in a training database; repeating the above steps until a predetermined number of pseudo-labeled audio datasets are stored in the training database; and training the automatic speech recognition module based on the predetermined number of pseudo-labeled audio datasets.
[0007] In some implementations, the input dataset may include a text dataset, and generating the text dataset includes introducing at least one of relatively uncommon vocabulary and relatively uncommon language into the text dataset.
[0008] In some embodiments, the method further includes comparing the text dataset for speaker verification with the output transcription file of the automatic speech recognition module and identifying at least one unsatisfactory result from the comparison. In some embodiments, identifying at least one unsatisfactory training result may include determining a precision and / or recall associated with one or more words in the text dataset.
[0009] In some implementations, the method further comprises the steps of selecting a speaker from a group of users historically associated with accurate text-dependent speaker verification, and assigning at least one pseudo label to an audio dataset generated by the speaker without verification.
[0010] In some implementations, the method further includes generating an output transcription file via an automatic speech recognition module.
[0011] In some embodiments, the input data set includes a first text data set and a second text data set, generating corresponding first output transcription files and second output transcription files, wherein the first text data set and the second output transcription file are configured to be input to a first reasoning module, and the second text data set and the first output transcription file are configured to be input to a second reasoning module.
[0012] In some implementations, the output of the first reasoning module is compared to the output of the second reasoning module.
[0013] In another aspect of the present disclosure, an apparatus for constructing a dataset for automatic speech recognition includes a processor configured to repeatedly: generate an input dataset for speaker verification; obtain an audio dataset based on the generated input dataset; verify the audio dataset using an audio verification module; assign at least one pseudo-label to the verified audio dataset; and store the at least one pseudo-labeled audio dataset in a training database until a predetermined number of pseudo-labeled audio datasets are stored in the training database.
[0014] In some implementations, the input dataset includes a text dataset, and the processor is configured to introduce at least one of relatively uncommon vocabulary and relatively uncommon language in generating the text dataset.
[0015] In some implementations, the processor is configured to compare a text dataset for speaker verification with an output transcription file of an automatic speech recognition module and to identify at least one unsatisfactory result from the comparison.
[0016] In some implementations, the processor is configured to identify at least one unsatisfactory training result based on a precision and / or recall associated with one or more words in the output transcription file.
[0017] In some implementations, the processor is configured to select a speaker from a group of users historically associated with accurate text-dependent speaker verification, and the processor is further configured to assign at least one pseudo label to an audio dataset generated by the speaker without verification.
[0018] In some implementations, the processor is configured to generate the text dataset based on feedback of at least one unsatisfactory result.
[0019] In some embodiments, the input data set includes a first text data set and a second text data set, generating a first output transcription file and a second output transcription file, respectively, wherein the first text data set and the second output transcription file are configured to be input to a first reasoning module, and the second text data set and the first output transcription file are configured to be input to a second reasoning module.
[0020] In some implementations, the output of the first reasoning module is compared to the output of the second reasoning module.
[0021] In another aspect of the present disclosure, a non-transitory computer-readable storage medium comprising instructions is provided, which, when executed by one or more processors, causes the execution of a method for constructing a dataset for automatic speech recognition according to any of the method embodiments described above.
[0022] In another aspect of the present disclosure, a data processing apparatus is provided, which is configured to perform a method according to any one of the method embodiments described above.
[0023] In another aspect of the present disclosure, a computer executable code is provided, comprising instructions for performing a method according to any one of the method embodiments described above.
[0024] In another aspect of the present disclosure, an automatic speech recognition module or system is provided, which is trained by any one of the method embodiments described above. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] The present disclosure will be better understood with reference to the detailed description when considered in conjunction with the non-limiting examples and accompanying drawings, in which:
[0026] - Figure 1 is a flowchart of a method for constructing data for automatic speech recognition according to various embodiments;
[0027] - Figure 2 is a block diagram of a system for constructing a dataset for training an automatic speech recognition module according to various embodiments;
[0028] - Figure 3 is a block diagram of a system for incorporating cross-validation to improve the quality of pseudo labels assigned to a dataset used to train an automatic speech recognition module;
[0029] - Figure 4 The application of a trained ASR module in the form of a post-call analysis system is shown;
[0030] - Figure 5 shows the application of the trained ASR module in the form of a transcription service for audio recordings of vehicle driving; and
[0031] - Figure 6 A server computer according to an embodiment is shown. DETAILED DESCRIPTION
[0032] The following specific embodiments are with reference to the accompanying drawings, which illustrate the specific details and embodiments of the present disclosure in a diagrammatic manner. These embodiments are described in sufficient detail to enable those skilled in the art to practice the present disclosure. Other embodiments may be utilized without departing from the scope of the present disclosure, and structural changes and logical changes may be made. The various embodiments are not necessarily mutually exclusive, as some embodiments may be combined with one or more other embodiments to form new embodiments.
[0033] Embodiments described in the context of one of a housing system, device, or method are similarly valid for the other system, device, or method. Similarly, embodiments described in the context of a system are similarly valid for the device or method, and vice versa.
[0034] Features described in the context of one embodiment may be applied accordingly to the same or similar features in other embodiments. Features described in the context of one embodiment may be applied accordingly to other embodiments, even if not explicitly described in these other embodiments. In addition, additions, combinations, and / or substitutions described for features in the context of one embodiment may be applied accordingly to the same or similar features in other embodiments.
[0035] In the context of various embodiments, the articles “a,” “an,” and “the” as used with respect to features or elements include reference to one or more of the features or elements.
[0036] As used herein, the term "and / or" includes any and all combinations of one or more of the associated listed items.
[0037] As used herein, the term "dataset" may be understood to include information in any suitable analog or digital form, for example, provided as a file, a portion of a file, a collection of files, a signal or stream, a portion of a signal or stream, a collection of signals or streams, a waveform, etc. However, the term dataset is not limited to the above examples and may take various forms and represent any information as understood in the art.
[0038] As used herein, the term "speaker verification" broadly encompasses methods or systems that accept or reject identity claims by comparing at least two audio samples (one used as a reference for the identity and the other collected from the person making the claim during testing / authentication).
[0039] As used herein, the term "pseudo-labeling" broadly encompasses the process or system of using a labeled data model to predict labels for unlabeled data. First, an initial set of pseudo-labeled data can be generated, which is assumed to be accurate based on historical records. This initial set of pseudo-labeled data can be used to generate subsequent pseudo-labels for unlabeled data sets.
[0040] As used herein, the term "module" refers to, forms part of, or includes an application-specific integrated circuit (ASIC); an electronic circuit; a combinational logic circuit; a field-programmable gate array (FPGA); a processor (shared, dedicated, or group) that executes code; other suitable hardware components that provide the described functionality; or a combination of some or all of the above (such as in a system-on-chip). The term module may include memory (shared, dedicated, or group) that stores code executed by the processor. A single module or a combination of modules may be considered a device.
[0041] As used herein, the terms "association," "associated," and "associated" indicate a defined relationship (or cross-reference) between two items. For example, an audio file may be associated with a text file, indicating that the audio file may be derived, calculated, and / or generated using the text file as a source or reference.
[0042] As used herein, "memory" may be understood as a non-transitory computer-readable medium in which data or information can be stored for retrieval. Therefore, references to "memory" herein may be understood to refer to volatile or non-volatile memory, including random access memory ("RAM"), read-only memory ("ROM"), flash memory, solid-state storage devices, magnetic tape, hard disk drives, optical disk drives, and the like, or any combination thereof. Furthermore, it should be understood that registers, shift registers, processor registers, data buffers, and the like are also encompassed by the term "memory" herein. It should be understood that a single component referred to as "memory" or "a memory" may be comprised of more than one different type of memory and, therefore, may refer to a collective component comprising one or more types of memory. It will be readily understood that any single memory component may be divided into multiple, mutually equivalent memory components, and vice versa. Furthermore, while memory may be depicted as separate from one or more other components (such as in the accompanying figures), it should be understood that memory may be integrated within another component (such as on a common integrated chip).
[0043] According to one aspect of the present disclosure and with reference to Figure 1 , a method 100 for constructing a dataset for automatic speech recognition. The method 100 can be implemented by a computer and includes the following steps: generating an input dataset (step 102); obtaining an audio dataset associated with the input dataset (step 104); verifying the audio dataset using an audio verification module (step 106); assigning at least one pseudo-label to the verified audio dataset (step 108); storing at least one pseudo-labeled audio dataset in a training database (step 110); repeating steps 102 to 110 until a predetermined number of pseudo-labeled audio datasets are stored in the training database (step 112); and training an automatic speech recognition module based on the predetermined number of pseudo-labeled audio datasets (step 114).
[0044] The input dataset in step 102 may be a text file or a text dataset containing multiple words in a language to be verified by a speaker. The multiple words may be arranged into the sentence "The quick brown fox jumps over the lazy dog," or arranged as discrete words representing consecutive numbers (e.g., "one," "two," "three," etc.), or arranged as discrete words representing random numbers (e.g., "nine," "six," "one," "seven," etc.). In some embodiments, generating the text dataset for speaker verification further includes introducing at least one of relatively uncommon vocabulary and relatively uncommon language into the text dataset. For example, the text file may include common words such as "the," "in," and "and." For another example, the text file may include uncommon words such as "unverified," "abstain," and "suicide." These uncommon words may also include informal words in slang or dialects, or may include a mixture of different languages, such as Singlish. In some embodiments, the text file may include Singlish words such as "alamak," "boleh," "cannot-lah," etc. In some embodiments, the input dataset may include a non-text dataset and a text dataset. In some embodiments, the input dataset may include multimedia content, and the text dataset may be extracted from the multimedia content.
[0045] In some implementations, introducing at least one of relatively uncommon vocabulary and relatively uncommon language into the text dataset can be performed after a predetermined number of iterations associated with training the ASR module.
[0046] In some embodiments, introducing at least one of relatively uncommon vocabulary and relatively uncommon language into the text dataset can be performed after receiving feedback that the ASR module has not been properly trained to transcribe relatively uncommon vocabulary and / or relatively uncommon language. The feedback can be generated based on identifying words with low precision or recall after training the ASR module for a period of time.
[0047] In some embodiments, speakers who generate audio files or datasets associated with text files for verification are selected from a group of registered users of an audio verification system. In some embodiments, the audio verification system can be a text-dependent speaker verification system, and the group of registered users has previously been verified using a speaker verification module. Registered users can register their audio files with the speaker verification system in advance, and one or more subgroups of registered users who have become accustomed to this process and have consistently provided accurate text during verifications over a period of time can be identified. The audio files previously deemed accurate over a period of time will form a historical database that is assumed to be accurate / correct for the training of the automatic speech recognition system of the present disclosure. These one or more subgroups of registered users will be referred to as or referred to as a high-quality user subgroup.
[0048] In step 104 , speakers are selected from the high-quality user subgroup to obtain audio files corresponding to the text dataset.
[0049] In step 106 , verification may be performed by an audio verification module in the form of a text-dependent speaker verification module for speaker registration.
[0050] In step 108, at least one pseudo-label is assigned to each of the verified audio files provided by the speakers selected from the high-quality user subgroup. The pseudo-label can be in the form of a mark to indicate that the audio file is accurate and will be used for subsequent training of the automatic speech recognition system of the present disclosure. In some embodiments, the pseudo-label can include metadata associated with the speaker and / or the text file.
[0051] The pseudo-tagged audio files may be stored as entries in a training database in step 110. Steps 102 to 108 may then be repeated until a predetermined number of entries have been entered into the training database.
[0052] In step 112, the pseudo-labeled audio files can be fed into an automatic speech recognition (ASR) model for training. The ASR model can include a machine learning (ML) algorithm or an artificial intelligence (AI) algorithm to process human speech into readable text. In some embodiments, the ML / AI algorithm can be trained using unsupervised learning, supervised learning, and / or hybrid training methods. In some embodiments, the ASR module includes an artificial intelligence module, which can include one or more neural network algorithms.
[0053] In some implementations, the input dataset to be fed into the ASR may include an audio file, and the output produced by the ASR may be a text output dataset based on the audio file.
[0054] In some embodiments, the text output dataset can be compared to a test / test set to identify relatively poorly performing words that may require more data collection to retrain the ASR. The test set can be a separate set from the pseudo-labeled training set and can be pseudo-labeled or properly labeled. Additionally or alternatively, the text output dataset can be reviewed by one or more users to identify relatively poorly performing words that may require retraining.
[0055] In another embodiment, steps 102 to 106 can be replaced by selecting a set of audio files from high-quality users. The assignment of pseudo labels is based on the assumption that high-quality users produce accurate audio files without further verification.
[0056] Figure 2 A system 200 for constructing a dataset for automatic speech recognition is shown. System 200 can form part of an apparatus for constructing a dataset for automatic speech recognition. System 200 includes a selector module 202 configured to generate input data, such as text data for speaker verification; a collector module 204 configured to obtain an audio file based on the generated text data; and an ASR module 206 arranged in data communication with selector module 202 and collector module 204. System 200 can include a review module 208.
[0057] The selector module 202 is configured to generate a text dataset (e.g., words) for speaker verification and subsequent training of the ASR module 206. The selector module 202 can be configured to include uncommon or rare words for each training or retraining of the ASR module 206, so that over time, the ASR module 206 can be trained to recognize more words, including uncommon words and words from uncommon languages. In some embodiments, the selector module 202 can be configured to include words from different languages, such as Southeast Asian languages. Non-limiting examples include Thai, Malay, Indonesian, Tagalog, etc.
[0058] The collector module 204 can be configured to collect a plurality of pseudo-labeled audio files and store the plurality of pseudo-labeled audio files in a training database. The plurality of pseudo-labeled audio files can be collected until a predetermined number (a target number) of audio files has been collected. The collector module 204 can include a verification module configured to verify the audio files based on the text and assign at least one pseudo-label to the verified audio files.
[0059] In some embodiments, the collected audio files may form a training set to train the ASR module 206. In some embodiments, the collected audio files may form a test set to test the trained ASR module 206.
[0060] The ASR module 206 can be trained and / or retrained based on the inspection module 208. In some embodiments, the ASR module 206 is configured to automatically identify patterns in an input audio file that include one or more speech waveforms. Patterns that can be detected from speech can include the identity of the speaker (for voice verification), language, emotion, and text transcription of spoken words.
[0061] In some embodiments, the ASR module 206 may include a machine learning (ML) algorithm or an artificial intelligence (AI) algorithm to process human speech into readable text. In some embodiments, the ML / AI algorithm may be trained using unsupervised learning, supervised learning, and / or hybrid training methods. In some embodiments, the ASR module 206 may include one or more neural network algorithms. In some embodiments, the ASR module 206 may be modeled as a pattern classifier and configured or programmed to process any input audio file using one or more of the following known techniques: feature extraction, segmentation, embedding, clustering, mapping / classification, and refinement. An example of an AI / ML algorithm would be a deep learning convolutional neural network.
[0062] The check module 208 is configured to check the output text data (transcription data) of the ASR module 206. In some embodiments, the check module 208 includes verifying the output text data and identifying words with relatively low precision / recall. In some embodiments, the check module 208 may include obtaining / receiving human intervention and feedback, such that if a user feels that some words are not well recognized, feedback can be sent to the selector module 202 to include more poorly recognized words. The selector module 202 can then be configured to customize the generation of the text dataset so that poorly recognized words can be selected for retraining.
[0063] In some embodiments, one or more of modules 202, 204, 206, and 208 may be combined or integrated into a single module. For example, the functionality associated with modules 202, 204, 206, and 208 may be implemented by a single computer processor or a group of computer processors.
[0064] Figure 3Another embodiment of a system 300 for constructing a dataset for automatic speech recognition is shown. The system 300 implements a cross-validation model to improve the quality and robustness of pseudo-labels that can be assigned to audio files. Two sets of collected data 302a and 302b in the form of audio files can be input into trained automatic speech recognition modules 306a and 306b, respectively. The output data 308a of the speech recognition module 306a can then be fed into an inference model 310a along with the input data 302b to perform inference on the output data 308a based on the input data 302b. The output data 308b of the speech recognition module 306b can then be fed into an inference model 310b along with the input data 302a to perform inference on the output data 308b based on the input data 302a.
[0065] The outputs from the inference models 310a and 310b can be compared, and any erroneous labels can be corrected and / or removed. For example, if the outputs from 310a and 310b are different or incorrectly identified based on comparison with a selected set of pseudo labels that are identified as always true or close to true, the erroneous label can be corrected and / or removed.
[0066] Figure 4 and Figure 5 Two applications of the trained ASR module of the present disclosure are shown.
[0067] Figure 4 The application of a trained ASR module in the form of a post-call analysis system 400 between a customer service representative (also known as an agent) and a customer is shown. The trained ASR module 450 can be deployed or implemented as a supplementary service to known conversation recording systems. The ASR module 450 can be configured to transcribe the conversation into text in real time or near real time. Figure 4 , two scenarios are shown: one scenario without using the ASR module 450 (labeled as A), and the other scenario with using the ASR module 450 (labeled as B).
[0068] In scenario A, a supervisor (e.g., a manager) will need to listen to the conversation recording system to identify the context and issues associated with the call. If necessary, the manager will need to call the customer for a solution. If the customer mentions other things that happened during the call with the agent, the manager will need to listen to the conversation recording system to confirm if there is anything missed.
[0069] In Scenario B, ASR module 450 transcribes the conversation based on an audio file obtained from a known conversation recording system. The manager can read the transcript and refer to it when calling the customer seeking a solution. If the customer mentions something else that occurred during the call with the agent, the manager only needs to quickly glance at the transcript. Compared to Scenario A, Scenario B enables the manager to call the customer back concurrently and identify key issues by reviewing the transcript.
[0070] Figure 5 Another application 500 is shown for transcribing a conversation between a service provider (e.g., a vehicle on-demand driver) and a user (e.g., a passenger) in an audio recording while a vehicle is in motion. The trained ASR module 550 can be deployed or implemented as a supplementary service to an in-vehicle conversation recording system. The ASR module 550 can be configured to transcribe the conversation into text in real time or near real time. Figure 5 , two scenarios are shown: one scenario without using the ASR module 550 (labeled A), and the other scenario with using the ASR module 550 (labeled B). If the trip ends in a conflict, the service provider can submit the trip recording for further investigation by the authorized party.
[0071] In Scenario A, a person (e.g., an administrator) may need to listen to the conversation recording system to identify conflicts associated with the recording. The administrator will need to transcribe the conversation and generate a report to be submitted to the next responsible person (e.g., a manager). At least one round of clarification between the administrator and the manager may be required before the report is finalized.
[0072] In Scenario B, ASR module 550 transcribes the conversation based on the audio file captured by the onboard recording system. The administrator can read the transcript while preparing the report and refer to the audio recording as needed. Again, the transcript is sent to the manager along with the audio recording for completeness. Compared to Scenario A, Scenario B provides more efficient time management associated with generating conflict reports.
[0073] Figure 1 The method is, for example, Figure 6 The server computer shown executes.
[0074] Figure 6 A server computer 600 is shown according to an embodiment.
[0075] The server computer 600 includes a communication interface 601 (e.g., configured to receive interaction data, i.e., information about the interaction). The server computer 600 also includes a processing unit 602 and a memory 603. The processing unit 602 can use the memory 603 to store, for example, data to be processed, such as generated text files and audio files. The server computer is configured to execute Figure 1method.
[0076] The methods described herein may be performed, and the various processing or computing units and devices and computing entities described herein may be implemented by one or more circuits. In an embodiment, a "circuit" may be understood as any kind of logic implementation entity, which may be hardware, software, firmware, or any combination thereof. Thus, in an embodiment, a "circuit" may be a hardwired logic circuit or a programmable logic circuit, such as a programmable processor, for example, a microprocessor. A "circuit" may also be software implemented or executed by a processor, for example, any kind of computer program, for example, a computer program using virtual machine code. Any other kind of implementation of the corresponding functions described herein may also be understood as a "circuit" according to an alternative embodiment.
[0077] Although the present disclosure has been particularly shown and described with reference to certain embodiments, it will be understood by those skilled in the art that various changes in form and details may be made therein without departing from the spirit and scope of the present disclosure as defined by the appended claims. The scope of the present disclosure is therefore indicated by the appended claims and all changes that come within the meaning and range of equivalents of the claims are therefore intended to be embraced.
Claims
1. A computer-implemented method for constructing at least one dataset for automatic speech recognition, the method comprising: Generate input dataset; Obtaining an audio data set associated with the input data set; Verifying the audio dataset using an audio verification module; assigning at least one pseudo-label to the verified audio dataset; storing at least one pseudo-labeled audio dataset in a training database; Repeat the above steps until a predetermined number of pseudo-labeled audio data sets have been stored in the training database; as well as An automatic speech recognition module is trained based on the predetermined number of pseudo-labeled audio data sets.
2. The computer-implemented method of claim 1 , wherein: The input dataset includes a text dataset, and generating the text dataset includes introducing at least one of relatively uncommon vocabulary and relatively uncommon language into the text dataset.
3. The computer-implemented method of claim 2 , further comprising: The text dataset for speaker verification is compared to an output transcription file of the automatic speech recognition module, and at least one unsatisfactory result is identified from the comparison.
4. The computer-implemented method of claim 3, wherein: Identifying at least one unsatisfactory training result includes determining a precision and / or a recall associated with one or more words in the text dataset.
5. The computer-implemented method of claim 3 or 4, further comprising: A speaker is selected from a group of users historically associated with accurate text-dependent speaker verification, and the at least one pseudo label is assigned to an audio dataset generated by the speaker without verification.
6. The computer-implemented method of any preceding claim, further comprising: An output transcription file is generated by the automatic speech recognition module.
7. The computer-implemented method of claim 1 , wherein: The input data set includes a first text data set and a second text data set, generating corresponding first output transcription files and second output transcription files, wherein the first text data set and the second output transcription file are configured to be input to a first reasoning module, and the second text data set and the first output transcription file are configured to be input to a second reasoning module.
8. The computer-implemented method of claim 7, wherein: The output of the first reasoning module is compared to the output of the second reasoning module.
9. An apparatus for constructing a dataset for automatic speech recognition, the apparatus comprising a processor configured to repeatedly: Generate input dataset for audio verification; obtaining an audio dataset based on the generated input dataset; Verifying the audio dataset using an audio verification module; assigning at least one pseudo-label to the verified audio dataset; storing at least one pseudo-labeled audio dataset in a training database; Until a predetermined number of audio data sets with pseudo labels are stored in the training database.
10. The device according to claim 9, wherein The input dataset is a text dataset, and the processor is configured to introduce at least one of relatively uncommon vocabulary and relatively uncommon language in generating the text dataset.
11. The device according to claim 10, wherein The processor is configured to compare the text data set for speaker verification with an output transcription file of the automatic speech recognition module and to identify at least one unsatisfactory result from the comparison.
12. The device according to claim 11, wherein The processor is configured to identify at least one unsatisfactory training result based on precision and / or recall associated with one or more words in the output transcription file.
13. The device according to claim 11 or 12, wherein: The processor is configured to select a speaker from a group of users historically associated with accurate speaker verification, and the processor is further configured to assign the at least one pseudo label to an audio dataset generated by the speaker without verification.
14. The device according to claim 11, wherein The processor is configured to generate the text dataset based on feedback of at least one unsatisfactory result.
15. The device according to claim 9, wherein The input data set includes a first text data set and a second text data set, generating a first output transcription file and a second output transcription file respectively, wherein the first text data set and the second output transcription file are configured to be input to a first reasoning module, and the second text data set and the first output transcription file are configured to be input to a second reasoning module.
16. The device according to claim 15, wherein The output of the first reasoning module is compared to the output of the second reasoning module.
17. A non-transitory computer-readable storage medium comprising instructions, which, when executed by one or more processors, cause the method for constructing a dataset for automatic speech recognition according to any one of claims 1 to 8 to be performed.
18. A data processing device configured to perform the method according to any one of claims 1 to 8.
19. A computer executable code comprising instructions for constructing a data set for automatic speech recognition according to any one of claims 1 to 8.
20. An automatic speech recognition module trained by the method according to any one of claims 1 to 8.